Jehanne Dussert – govllm
Responses
In your opinion, what outcomes would make the first Global Dialogue on AI Governance a success?
The Global Digital Compact frames AI governance as requiring cooperation that is "agile and adaptable to the rapidly changing digital landscape." The first Dialogue will succeed if it moves beyond reaffirming this aspiration and begins operationalising it, specifically by establishing a shared baseline for what "accountable AI" means in practice, not only in principle. From my perspective as a technical practitioner working on LLM governance in a large public administration, success would require three concrete outputs: * First, agreement on a minimum set of observable indicators that public sector actors can use to demonstrate continuous AI compliance, not merely at deployment, but throughout the production lifecycle. Current frameworks, including the AI Act, remain largely snapshot-based. The Dialogue should initiate work toward a shared vocabulary of runtime metrics (output consistency, refusal rates, regulatory scoring) that can anchor accountability claims. * Second, a mechanism for cross-jurisdictional comparability of governance approaches, as called for in resolution 79/325. Interoperability of frameworks is unachievable without shared reference points. Producing even a preliminary taxonomy of evaluation criteria across major regulatory regimes (EU AI Act, GDPR, national security frameworks) would represent tangible progress. * Third, genuine integration of technical community contributions into the normative outputs. The Dialogue's legitimacy depends on closing the gap between the actors who write governance rules and those who implement them under real operational constraints. A success indicator would be that technical practitioners (particularly from public sector contexts) are reflected in the Dialogue's resulting recommendations, not merely consulted. The Dialogue should be measured not by the declarations it produces, but by whether it generates governance tools that practitioners can actually use.
From your perspective, which of the following thematic areas identified by the General Assembly Resolution 79/325 for the AI Dialogue reflect your priorities for urgent action and active engagement?
- Transparency, accountability, and human oversight
- Safe, secure and trustworthy AI
- Interoperability of governance approaches
- AI capacity-building
Please briefly explain your selection.
7
These four priorities reflect the operational reality of deploying LLMs at scale in a public administration context, and the governance gaps I have encountered directly as a Tech & Governance Lead. * Transparency, accountability, and human oversight is the foundational priority. The Global Digital Compact commits to "a responsible, accountable, transparent and human-centric approach to the life cycle of digital and emerging technologies." In practice, this lifecycle commitment remains unfulfilled: most governance frameworks assess AI systems at intake, not in production. Accountability requires continuous observability (trace logging, output scoring against regulatory criteria, anomaly detection...) not only documentation at deployment. * Safe, secure and trustworthy AI connects directly to this: in public sector deployments, safety is not a binary property established at release. It must be monitored against evolving threat vectors (prompt injection, data leakage, model drift) and regulatory standards such. An AI system that was compliant at launch may not remain so. * Interoperability of governance approaches is technically urgent. Practitioners operating across regulatory jurisdictions must simultaneously satisfy requirements from multiple frameworks (AI Act risk classification, GDPR data minimisation, sector-specific constraints). Without interoperable evaluation criteria, compliance becomes fragmented and unverifiable. The Dialogue should prioritise producing a common evaluation layer that maps across existing frameworks rather than adding a new one. * AI capacity-building is essential but needs reframing: the gap is not primarily one of awareness but of tooling. Non-technical public decision-makers need accessible, interpretable governance interfaces. Capacity-building efforts should invest in open-source tools that translate technical monitoring outputs into governance-legible evidence.
In your opinion, are there any cross-cutting or emerging issues not captured by the listed themes above? If so, please explain.
2
Two issues are structurally absent from the thematic areas identified in resolution 79/325, yet they condition the effectiveness of all others. * The governance gap between static compliance and dynamic behaviour All current frameworks (including the AI Act) operate on a point-in-time logic: a system is classified, documented, and assessed at deployment. But large language models are not static artefacts. Their outputs shift with prompt variation, user context, and model updates. This creates a compliance fiction: a system can be fully documented and still exhibit harmful or non-compliant behaviour at runtime. The Dialogue should initiate work on a framework for continuous governance, defining what ongoing monitoring obligations look like, what metrics constitute evidence of compliance over time, and how to link production observability to institutional accountability cycles. This is distinct from the existing "transparency" theme, which addresses disclosure rather than runtime verification. * Evaluative sovereignty States and public institutions deploying AI systems increasingly depend on evaluation infrastructure (benchmarks, red-teaming services, LLM-based judges...) that is itself proprietary or foreign-controlled. The capacity to assess one's own AI systems independently is a prerequisite for meaningful governance, yet it is treated as a technical detail rather than a governance priority. Resolution 79/325 references open-source software and open models, but does not address the specific challenge of evaluation tooling sovereignty. Public administrations in particular need access to self-hostable, configurable evaluation frameworks that do not route sensitive data through third-party infrastructure. Both gaps share a common structure: governance frameworks assume a level of institutional technical capacity that most public actors do not yet have. Addressing this is a precondition for the Dialogue's outputs to have real-world effect.
How are the governance gaps and related developments/advances in the thematic areas you selected above affecting your country, region, or sector? Please highlight the most significant challenges.
The French public sector illustrates both the ambition and the limits of current AI governance approaches. Several genuinely advanced initiatives exist: - the DINUM's EvalAP platform (https://evalap.etalab.gouv.fr/) provides a sovereign, reproducible model evaluation infrastructure for public administrations - Albert API offers a self-hosted LLM stack that avoids routing sensitive data through commercial providers (https://albert.sites.beta.gouv.fr/) - and the ALLiaNCE incubator has begun mapping AI governance documents across ministries (https://alliance.numerique.gouv.fr/cartographie/portail-des-chartes-ia-dans-ladministration/). Yet the dominant governance instrument remains terms of use, a declaratory document, typically under ten pages, establishing usage principles for public agents. The ALLiaNCE portal now aggregates these across ministries, from the Élysée's Paris Charter to sectoral guides in education and defence. Their proliferation signals institutional awareness. Their content, however, largely addresses surface-level questions (prohibiting use of consumer-facing tools, reminding agents of confidentiality obligations) rather than the governance problems that arise once AI systems are actually deployed at scale: output consistency, compliance drift, accountability for automated decisions, and evidence of ongoing regulatory conformity. The gap between these two layers is significant. EvalAP evaluates models at selection time, charters frame acceptable use in principle. Neither addresses what happens in production: how output quality is monitored, how regulatory compliance is verified continuously, how governance decisions are triggered by observed system behaviour rather than periodic audits. This is not a criticism of these initiatives: they represent genuine public investment. The challenge is structural: French public sector AI governance currently has strong foundations at intake (evaluation, procurement) and weak infrastructure at runtime (monitoring, continuous compliance, governance-legible evidence). Closing this gap requires neither more chartes nor new regulatory obligations, but operational tooling that connects production observability to institutional accountability and international frameworks that recognise this as a governance requirement, not a technical detail.
What role can the AI Dialogue play in advancing international cooperation on AI governance?
The AI Dialogue occupies a structural position that no regional or bilateral mechanism can replicate: it is the only forum where technically divergent governance regimes must find common ground within a single multilateral process. Its comparative advantage is not standard-setting, but convergence-enabling: creating shared reference points that allow different frameworks to interoperate without requiring harmonisation. Concretely, the Dialogue can advance international cooperation in three ways that are currently unmet. * First, by establishing a common evaluation vocabulary. "Transparency," "accountability," and "human oversight" appear across every major governance framework, but with operationally incompatible definitions. A practitioner deploying an LLM under the EU AI Act, GDPR, and an OWASP-aligned security policy simultaneously cannot currently demonstrate compliance to all three without redundant, incommensurable documentation. The Dialogue should mandate a working group to produce a cross-framework mapping of evaluation criteria, not to unify the frameworks, but to make them mutually legible. * Second, by anchoring the Independent International Scientific Panel's outputs to implementable governance indicators. The Panel risks producing authoritative assessments that have no pathway into operational practice. The Dialogue should establish an explicit bridge: Panel findings should generate concrete monitoring recommendations that public sector actors can adopt. * Third, by creating accountability for implementation. The GDC's commitments are non-binding, but the Dialogue can introduce a light peer-review mechanism (voluntary, institutionally lightweight) through which member states and technical community actors share evidence of how they are operationalising governance commitments. This shifts the Dialogue from a forum of declarations to a forum of practice.
What are some of the existing initiatives, partnerships, or mechanisms that the AI Dialogue should build upon or connect with, and what added value could the AI Dialogue bring?
Several existing initiatives provide foundations the Dialogue should explicitly connect rather than duplicate. * The OECD AI Policy Observatory and its AI Principles offer the most widely adopted baseline for governance criteria. The Dialogue should treat this as a floor, focusing its multilateral added value on what OECD instruments cannot provide: normative weight for non-member states and a legitimate space for governance approaches developed outside the Euro-Atlantic axis. * The EU AI Act implementation infrastructure (including the European AI Office and the codes of practice under development) represents the most operationally advanced attempt to translate governance principles into institutional process. The Dialogue should learn from it while identifying what it does not address: runtime monitoring obligations, continuous compliance evidence, and applicability beyond the EU regulatory perimeter. * National AI safety institutes (including the UK AI Safety Institute, France's INESIA, and the US NIST AI Safety Institute) are developing technical evaluation capacity that intergovernmental processes cannot replicate. Their structural limitation is that their mandates are national and their outputs not mutually recognised. The Dialogue should acknowledge them as an emerging layer of technical governance infrastructure without attempting to subsume or coordinate them. * The Fiscalis EU working group on GenAI in tax administration offers a more directly relevant model: sector-specific, cross-border governance cooperation among public practitioners, producing outputs grounded in operational reality rather than regulatory aspiration. This format (practitioner-led, evidence-based, institutionally lightweight) is replicable across sectors and represents the kind of governance development the Dialogue should actively encourage and reference. The Dialogue's specific added value sits above all of these: creating the multilateral legitimacy and shared vocabulary that allows divergent national and regional approaches to become mutually legible, without requiring harmonisation that is neither feasible nor necessarily desirable at this stage.
How can different stakeholders contribute to the AI Dialogue? Please share recommendations for the format and structure of the AI Dialogue.
The Dialogue's legitimacy depends on closing the structural gap between those who write governance rules and those who implement them. This requires differentiating stakeholder contributions by function, not merely by category. Governments bring normative authority and implementation mandates. Their contribution should be structured around accountability: what governance commitments have they operationalised, and with what observable results. -> Format recommendation: member state reporting should include a standardised evidence component, not only policy descriptions. Technical community actors (including practitioners from public administrations, open-source developers, and applied researchers) bring operational knowledge that is currently underutilised in normative processes. Their contribution should be structured as problem specification: identifying where governance frameworks break down in practice, and what tooling or evaluation infrastructure is missing. -> Format recommendation: dedicated technical working sessions prior to plenary, with outputs feeding directly into Dialogue recommendations rather than being summarised into oblivion in side-event reports. Civil society brings rights-holder perspectives and accountability pressure. Their contribution is most effective when they can interrogate implementation claims rather than respond to polished presentations. -> Format recommendation: structured challenge sessions in which civil society actors can formally question member state and private sector governance claims, with responses on record. On structure: the Dialogue should resist the temptation to organise sessions thematically by the seven listed areas. The most productive sessions will be cross-cutting, for example, "what does transparency require technically, legally, and institutionally?" because governance problems do not respect thematic boundaries. A matrix format, pairing thematic domains with implementation stages (design, deployment, production, audit), would surface the gaps that purely thematic organisation obscures.
Which voices, communities, or perspectives are currently underrepresented in global discussions on AI governance? How could they be included?
From my point of view, one category of actors is structurally underrepresented in current global AI governance discussions: public sector technical practitioners. The governance conversation is dominated by policymakers, legal experts, and civil society organisations on one side, and large technology companies on the other. The practitioners who actually deploy, monitor, and maintain AI systems within public administrations are largely absent. They hold knowledge that neither regulators nor vendors possess: what governance frameworks look like when they meet real operational constraints, what breaks, and what the actual cost of compliance is. Including them requires creating submission pathways and participation formats that do not require institutional endorsement from ministries, many operate under constraints that make formal representation difficult.
What innovative engagement formats could most effectively foster meaningful and dynamic engagement during the AI Dialogue?
The formats that would most effectively foster meaningful engagement are those that foreground implementation evidence rather than position statements. Governance red-teaming sessions. Structured exercises in which practitioners attempt to apply existing governance frameworks to concrete deployment scenarios, and document where the frameworks fail, contradict each other, or produce unverifiable requirements. This format generates actionable gap analysis rather than additional declarations. It also creates a natural entry point for technical community actors who have direct operational experience but limited appetite for plenary debate. Open evidence submissions with structured peer response. Rather than written contributions disappearing into secretariat synthesis, submissions from technical community actors should be publicly accessible and subject to formal response from member states or other stakeholders. This introduces a light accountability mechanism without requiring treaty-level commitments, and models the kind of iterative, evidence-based process the Dialogue should itself embody. Cross-jurisdictional implementation comparisons. Sessions structured around a shared scenario, for example, "deploying a generative AI system for public-facing administrative services under your jurisdiction's regulatory regime" in which practitioners from different countries work through the same problem. The divergences that emerge are more informative than any gap analysis produced by secretariat staff, and the convergences identify where shared tooling or evaluation criteria are already possible. The Dialogue should also establish a persistent digital workspace, not a document repository but a structured, searchable record of governance problems raised, responses offered, and commitments made. The absence of institutional memory across sessions is one of the reasons multilateral governance processes accumulate declarations without accumulating knowledge.
Please share examples of policies, practices, platforms, or approaches that promote effective AI governance or offer concrete solutions to addressing its challenges.
5
Three approaches illustrate what effective AI governance looks like when normative frameworks connect to operational practice. * Sovereign evaluation infrastructure as a governance foundation. France's EvalAP platform, developed by Etalab/DINUM, demonstrates that public administrations can build reproducible, self-hosted model evaluation capacity independent of commercial providers. This addresses evaluative sovereignty directly: administrations select and assess models against their own criteria, without routing sensitive data through third-party infrastructure. The limitation is that EvalAP operates at intake and it informs procurement decisions but does not extend into production monitoring. This is the next necessary layer: connecting pre-deployment evaluation to runtime observability so that governance evidence is continuous rather than point-in-time. * Declaratory governance as a starting point, not an endpoint. The ALLiaNCE portal aggregating AI usage charters across French ministries represents a genuine coordination effort: shared vocabulary, mutual visibility, and a baseline of usage principles across administrations. These instruments are necessary but insufficient. A charter that instructs agents not to use consumer-facing tools does not address what happens when a sovereign LLM deployment produces inconsistent, non-compliant, or drifting outputs at scale. The gap between declaratory governance and operational governance (between stating principles and verifying them continuously) is where the most significant accountability failures occur. * Metrics-based governance closing that gap. The missing practice is one that connects these two layers: production observability feeding institutional accountability cycles. This means continuous scoring of LLM outputs against regulatory criteria, use-case-by-model evaluation matrices interpretable by non-technical governance actors, and routing decisions derived from monitored compliance indicators rather than static policy documents. Open-source, self-hostable tooling implementing this model exists and is replicable across public administrations facing equivalent AI Act implementation constraints. The Dialogue should recognise this architecture as a transferable public good, not a national technical detail. A working implementation of the monitoring architecture described above is available as an open-source project: github.com/JehanneDussert/govllm