Skip to content

AI Integrity Organization (AIO)

Civil Society Global

Responses

In your opinion, what outcomes would make the first Global Dialogue on AI Governance a success?

The first Global Dialogue on AI Governance would be a success if it moves beyond principle-setting toward measurable accountability. International frameworks — including the UNESCO Recommendation on AI Ethics, the OECD AI Principles, and the Global Digital Compact — have established important normative foundations. What remains missing is the empirical infrastructure to verify whether AI systems actually operate in accordance with these principles. A successful outcome would include three elements. First, recognition that AI governance requires behavioral measurement — not just output-level safety testing, but structural assessment of how AI systems prioritize values, weigh evidence, and select sources when making decisions. Second, a commitment to developing or endorsing provider-neutral measurement tools that can be applied across commercial AI models regardless of origin, enabling genuine comparison and transparency. Third, mechanisms to ensure that governance frameworks are not only interoperable across jurisdictions but also empirically verifiable — so that when a nation or organization claims compliance with international AI norms, that claim can be independently tested. Without measurement, principles remain aspirational. The Dialogue would succeed if it establishes that the gap between AI governance principles and AI system behavior is itself a governance problem — one that requires dedicated measurement infrastructure, open datasets, and cross-cultural validation methodologies to solve.

From your perspective, which of the following thematic areas identified by the General Assembly Resolution 79/325 for the AI Dialogue reflect your priorities for urgent action and active engagement?

  • Safe, secure and trustworthy AI
  • Transparency, accountability, and human oversight
  • Interoperability of governance approaches
  • Protection and promotion of human rights

Please briefly explain your selection.

The first Global Dialogue on AI Governance would be a success if it moves beyond principle-setting toward measurable accountability. International frameworks - including the UNESCO Recommendation on AI Ethics, the OECD AI Principles, and the Global Digital Compact - have established important normative foundations. What remains missing is the empirical infrastructure to verify whether AI systems actually operate in accordance with these principles. A successful outcome would include three elements. First, recognition that AI governance requires behavioral measurement - not just output-level safety testing, but structural assessment of how AI systems prioritize values, weigh evidence, and select sources when making decisions. Second, a commitment to developing or endorsing provider-neutral measurement tools that can be applied across commercial AI models regardless of origin, enabling genuine comparison and transparency. Third, mechanisms to ensure that governance frameworks are not only interoperable across jurisdictions but also empirically verifiable - so that when a nation or organization claims compliance with international AI norms, that claim can be independently tested. Without measurement, principles remain aspirational. The Dialogue would succeed if it establishes that the gap between AI governance principles and AI system behavior is itself a governance problem - one that requires dedicated measurement infrastructure, open datasets, and cross-cultural validation methodologies to solve.

In your opinion, are there any cross-cutting or emerging issues not captured by the listed themes above? If so, please explain.

4

One critical issue not fully captured by the listed themes: the systematic measurement of AI value hierarchies as a governance infrastructure. Current AI governance discussions focus on two levels - principles (what AI should do) and outputs (what AI actually produces). There is a missing middle layer: the structural priorities that govern how AI systems make decisions. This layer determines which values an AI model prioritizes over others, what types of evidence it treats as authoritative, and whose expertise it defers to - all before any specific output is generated. This is not an abstract concern. Empirical measurement across commercial AI models reveals consistent patterns that no current governance framework addresses. Models from different providers show fundamentally different value profiles - some systematically prioritize collective security, others individual autonomy - yet these differences are invisible to users, regulators, and often to the providers themselves. When the same model shifts its value priorities depending on whether a scenario is framed as a news report or a policy memo, that context-sensitivity represents a governance challenge that output-level testing cannot detect. The emerging issue is therefore: AI governance needs a measurement layer between principles and outputs. Just as financial regulation requires standardized accounting, and environmental regulation requires emissions measurement, AI governance requires standardized behavioral measurement of how AI systems actually prioritize values in practice. Without this layer, compliance verification remains impossible, cross-jurisdictional interoperability lacks a common metric, and the gap between governance aspirations and AI system behavior will continue to widen. We recommend the Dialogue consider establishing working groups or consultations specifically focused on developing this measurement infrastructure as a foundational element of international AI governance.

How are the governance gaps and related developments/advances in the thematic areas you selected above affecting your country, region, or sector? Please highlight the most significant challenges.

The most significant governance gap is the disconnect between the proliferation of AI governance principles and the near-total absence of tools to verify compliance with those principles at the behavioral level. Challenges: Multiple governance frameworks now exist — the EU AI Act, the OECD AI Principles, UNESCO's Recommendation, national AI safety institutes in the US, UK, Japan, and Korea. Yet none of these frameworks currently require or provide standardized measurement of how AI models actually prioritize values in their decision-making. This creates three practical problems. First, providers can claim alignment with any principle without independent verification. Second, regulators cannot compare AI systems across providers using common behavioral metrics. Third, users — the people most affected by AI value priorities — have no visibility into how the AI systems they interact with daily are making value-laden decisions on their behalf. The challenge is compounded globally: AI models developed in one jurisdiction are deployed worldwide, but their embedded value hierarchies may conflict with local norms, legal frameworks, or cultural expectations. There is currently no mechanism to detect or disclose these conflicts before deployment. Opportunities: The convergence of multiple governance initiatives creates a unique window. The EU AI Act's implementation timeline, the establishment of national AI safety institutes, and the launch of this Global Dialogue all point toward a moment where measurement infrastructure could be adopted across jurisdictions simultaneously. Behavioral benchmarking methodologies grounded in cross-culturally validated frameworks already exist and have been applied at scale. The opportunity is to integrate such measurement into governance frameworks now — rather than building enforcement mechanisms first and discovering the measurement gap later. Early adoption of behavioral measurement standards would give the international community a shared empirical foundation for the interoperability that every governance framework aspires to but none yet achieves.

What role can the AI Dialogue play in advancing international cooperation on AI governance?

The AI Dialogue can play a unique role that no existing initiative currently fills: establishing a shared empirical foundation for AI governance across jurisdictions. Today, international cooperation on AI governance is fragmented not by lack of goodwill but by lack of common measurement. The EU AI Act, the US Executive Order on AI, the UK AI Safety Institute, Korea's AI Basic Act, and Japan's Hiroshima AI Process all pursue trustworthy AI — but each defines and evaluates trustworthiness differently. Without common metrics, "safe AI" means different things in different jurisdictions, making interoperability aspirational rather than operational. The AI Dialogue is positioned to address this because it operates at the level of the UN General Assembly — above any single regulatory jurisdiction. Three specific roles would be most impactful. First, the Dialogue could commission or endorse the development of provider-neutral behavioral measurement standards for AI systems. Not prescribing what values AI should hold, but establishing how to measure what values AI systems actually prioritize — enabling every jurisdiction to assess AI behavior using comparable methodologies. Second, the Dialogue could serve as a clearinghouse for empirical findings about AI system behavior. When measurement reveals that commercial AI models systematically restructure their value priorities in specific contexts — such as defense, healthcare, or education — those findings have immediate relevance for every Member State deploying those same models. Third, the Dialogue could bridge the current gap between normative agencies (UNESCO, OHCHR) and technical agencies (ITU, WIPO). Value measurement sits precisely at this intersection — it is technical in methodology but normative in implications. No existing forum connects these communities around shared empirical data about AI behavior. The AI Dialogue's greatest added value would be converting shared principles into shared measurement — making international cooperation verifiable, not just declaratory.

What are some of the existing initiatives, partnerships, or mechanisms that the AI Dialogue should build upon or connect with, and what added value could the AI Dialogue bring?

The AI Dialogue should build upon several existing initiatives while filling a specific gap that none of them currently addresses. Initiatives to build upon: The OECD AI Policy Observatory and its Catalogue of Tools & Metrics for Trustworthy AI provide the most developed infrastructure for sharing AI governance tools across countries. The AI Dialogue should actively connect with this catalogue to ensure that measurement tools discussed or endorsed through the Dialogue are discoverable and accessible to all Member States. National AI Safety Institutes — particularly those in the US (AISI), UK (AISI), Japan, and Korea (KAISI) — represent the operational layer where governance principles meet technical evaluation. The Dialogue could facilitate coordination among these institutes on common measurement standards, preventing the emergence of incompatible national evaluation regimes. The EU AI Act's General-Purpose AI Code of Practice, currently under development, will set concrete compliance benchmarks for the world's largest regulated market. The Dialogue should engage with this process to ensure global perspectives are reflected. UNESCO's Readiness Assessment Methodology provides a framework for evaluating national AI governance capacity. The Dialogue could extend this by incorporating behavioral measurement of AI systems deployed within those nations — assessing not just governance readiness but actual AI system behavior. Added value of the AI Dialogue: The unique contribution would be convening power at the General Assembly level to address what bilateral and regional initiatives cannot — establishing truly global measurement standards. No existing initiative operates at sufficient scale or neutrality to set provider-neutral behavioral benchmarks accepted across all governance jurisdictions simultaneously. The OECD serves developed economies; regional frameworks serve their members. Only a UN-level dialogue can establish measurement infrastructure that applies equally to AI systems deployed in Seoul, São Paulo, and Nairobi. The Dialogue's added value is not creating new principles — it is creating the empirical infrastructure to make existing principles enforceable.

How can different stakeholders contribute to the AI Dialogue? Please share recommendations for the format and structure of the AI Dialogue.

Different stakeholders bring different assets to AI governance — the Dialogue's structure should match contributions to capabilities rather than treating all participants identically. AI providers should be required to submit behavioral profiles of their models as a condition of participation — not just policy positions. When a provider advocates for specific governance approaches, the Dialogue should be able to reference empirical data about how that provider's models actually behave. Statements of principle without accompanying behavioral evidence should be clearly distinguished from empirically supported positions. Civil society and independent research organizations can contribute measurement methodologies and empirical findings that are provider-neutral. The Dialogue should create a formal submission track for empirical evidence — not just policy recommendations — so that governance discussions are grounded in data about actual AI system behavior. National governments and AI safety institutes should share their evaluation methodologies and findings. The Dialogue could establish a mutual recognition process where national evaluation approaches are compared, gaps identified, and common elements extracted toward interoperable standards. Academia brings methodological rigor. The Dialogue should commission independent academic review of any measurement standards proposed, ensuring they meet established validity criteria — cross-cultural validation, statistical robustness, and resistance to gaming. Structural recommendations: The Dialogue should operate in two parallel streams. Stream one: policy deliberation among Member States, as currently planned. Stream two: an empirical evidence track where stakeholders submit data, measurement tools, and behavioral findings about AI systems. The second stream would feed directly into the first, ensuring that policy discussions reference actual AI behavior rather than hypothetical scenarios. Each session should include structured time for examining empirical evidence — not just hearing position statements. The goal is a Dialogue where "AI systems should respect human values" is immediately followed by "here is what measurement shows they actually do."

Which voices, communities, or perspectives are currently underrepresented in global discussions on AI governance? How could they be included?

Three communities are critically underrepresented in global AI governance discussions. First, everyday AI users. Current governance discussions are dominated by providers, regulators, and researchers — the people who build, regulate, and study AI. The billions of people whose beliefs, decisions, and values are being shaped by daily AI interactions have virtually no voice. They are the most affected stakeholders yet the least consulted. The AI Dialogue should include structured input mechanisms for user communities — not through technical submissions, but through accessible formats that capture how AI systems are actually experienced by people in their daily lives across different cultures and contexts. Second, independent measurement organizations. The AI governance ecosystem includes principle-setters (UNESCO, OECD), regulators (EU AI Office, national authorities), and providers (technology companies) — but very few organizations focused exclusively on empirical measurement of AI behavior. This creates a structural gap: those who set standards, those who are regulated, and those who regulate are all present, but those who can independently verify compliance are largely absent. The Dialogue should actively seek out and include organizations conducting provider-neutral behavioral assessment of AI systems. Third, non-English-speaking communities and Global South perspectives. AI models trained predominantly on English-language data embed value hierarchies that may systematically differ from those held by communities in other linguistic and cultural contexts. Empirical measurement shows that AI value priorities vary significantly across domains — but current discussions rarely examine how these priorities interact with diverse cultural value systems. Korean, Arabic, Hindi, and Swahili-speaking communities experience AI differently, and those differences are measurable but unmeasured. Inclusion mechanisms should go beyond translation. They should include behavioral measurement of AI systems as deployed in different linguistic and cultural contexts — making visible the value gaps that currently remain hidden.

What innovative engagement formats could most effectively foster meaningful and dynamic engagement during the AI Dialogue?

The most impactful innovation would be introducing live empirical evidence into the Dialogue format itself. Live behavioral demonstration sessions: Rather than discussing AI governance in the abstract, the Dialogue could include sessions where AI models are tested in real time against governance scenarios relevant to the discussion. When delegates debate whether AI systems respect cultural diversity, a live measurement showing how specific models actually prioritize values in that context would ground the conversation in observable behavior rather than corporate promises or theoretical concerns. Comparative behavioral dashboards: The Dialogue could commission the creation of public dashboards showing how major commercial AI models compare on key governance-relevant behavioral metrics. These dashboards — updated periodically — would give delegates and the public a shared reference point. Instead of each stakeholder citing their own data, everyone would reference the same independently measured behavioral profiles. Cross-cultural value measurement workshops: Structured sessions where participants from different regions examine how the same AI model behaves when presented with identical ethical dilemmas framed in different cultural contexts. This would make abstract discussions about cultural sensitivity concrete and data-driven. User experience testimony tracks: Dedicated sessions where everyday users from diverse backgrounds share how AI systems have influenced their decisions, beliefs, or access to information — paired with behavioral measurement data showing the structural reasons behind those experiences. This connects individual stories to systemic patterns. Open measurement challenges: The Dialogue could issue measurement challenges — specific governance questions that stakeholders are invited to answer with empirical data before the next session. This creates accountability and continuity between Dialogue sessions, and incentivizes the development of measurement tools that serve governance needs. The common thread: every format innovation should bring empirical evidence into the room, making the Dialogue a space where governance is debated with data, not just with positions.

Please share examples of policies, practices, platforms, or approaches that promote effective AI governance or offer concrete solutions to addressing its challenges.

2

We highlight one concrete practice and several complementary approaches that collectively demonstrate how AI governance can move from principles to measurable implementation. Behavioral benchmarking of AI value hierarchies: Our organization has developed and applied a methodology that measures how commercial AI models actually prioritize values, evidence types, and source credibility when making decisions. Using cross-culturally validated frameworks (Schwartz value theory, Walton's argumentation schemes, GRADE/CEBM evidence standards), we administer forced-choice protocols across 7 professional domains and 15 severity levels. This has been applied to 8 commercial models, generating approximately 397,000 responses. The methodology is open, the results are published under CC BY 4.0, and the approach is provider-neutral - any model accessible via API can be evaluated using identical protocols. This demonstrates that structural behavioral measurement of AI systems is technically feasible, economically viable, and immediately applicable to governance needs. Complementary approaches worth building upon: The OECD Catalogue of Tools & Metrics for Trustworthy AI provides a discoverable registry of governance tools - an essential infrastructure for scaling measurement adoption across countries. The EU AI Act's risk-based classification establishes regulatory categories that measurement tools can be mapped onto - connecting behavioral findings to compliance requirements. National AI Safety Institutes (US, UK, Korea, Japan) represent operational infrastructure where measurement methodologies can be deployed and tested within regulatory contexts. UNESCO's Readiness Assessment Methodology offers a framework that could be extended from institutional capacity assessment to include empirical measurement of deployed AI system behavior. What connects these approaches: each addresses a necessary layer of AI governance infrastructure. What is still missing is the integration layer - standardized behavioral measurement that feeds into regulatory classification, populates international catalogues, and provides safety institutes with common evaluation protocols. The technology and methodology for this integration exist. The