Skip to content

Policy Genome

Civil Society Eastern Europe

Responses

In your opinion, what outcomes would make the first Global Dialogue on AI Governance a success?

A successful Dialogue would produce three things: shared language, a monitoring commitment, and a mandate for independent oversight. My EU-funded audit of six major AI systems found that the same model can deliver accurate information in one language and propaganda-aligned content in another. This is measurable today with open methodology. The Dialogue should endorse repeatable, multilingual benchmarking as a baseline expectation for high-risk contexts, and commit to independent auditing as the standard — not self-reporting by developers. A concrete success would be agreement that language-specific safety evaluation is a condition for public-sector AI procurement.

From your perspective, which of the following thematic areas identified by the General Assembly Resolution 79/325 for the AI Dialogue reflect your priorities for urgent action and active engagement?

  • Safe, secure and trustworthy AI
  • Social, economic, ethical, cultural, linguistic and technical implications of AI
  • Transparency, accountability, and human oversight
  • Protection and promotion of human rights

Please briefly explain your selection.

4

My selection reflects findings from an independent audit of six major AI systems across English, Ukrainian, and Russian (126 responses, Cohen's κ ≥ 0.69). Non-Western models showed language-conditioned divergence: Yandex Alice endorsed propaganda in 86% of Russian-language responses while refusing the same questions in English. DeepSeek echoed propaganda framing in 29% of Russian responses but was accurate in English. Current safety evaluations do not catch this. Western models showed false balance in 5-19% of responses, framing documented aggression as a contested narrative. This bypasses standard safety metrics but causes measurable harm. One structural finding: Yandex Alice generated accurate responses that were then programmatically overwritten with refusals before display. Transparency requirements must cover the full deployment layer, not only the underlying model.

In your opinion, are there any cross-cutting or emerging issues not captured by the listed themes above? If so, please explain.

2

Two issues are absent from the listed themes. First, language-conditioned narrative divergence as a distinct governance category. The same AI system can function as an accurate source for English speakers and a propaganda channel for Russian speakers. This is not only a capacity problem, it is also an integrity and security risk that no current framework explicitly addresses. Second, false balance as a measurable failure mode. Framing documented facts as competing perspectives causes harm equivalent to misinformation but bypasses existing filters. Governance frameworks need a distinct category for it, with guidance on detection and reporting. Both are invisible to single-language evaluation. The methodology to detect them exists, is open, and has been validated. What is missing is institutional will to require it.

How are the governance gaps and related developments/advances in the thematic areas you selected above affecting your country, region, or sector? Please highlight the most significant challenges.

AI systems operating in Russian are measurably more likely to deliver propaganda-aligned content than the same systems in English. For Eastern Europe and Ukraine specifically, this creates asymmetric information environments where linguistic communities face structurally different AI outputs on conflict, security, and political facts. The governance gap is concrete: no multilateral framework currently requires language-specific safety testing. Public institutions procure AI with no visibility into cross-lingual performance. The opportunity is equally concrete: the methodology to audit this exists, is open, and has been validated. What is missing is a mandate to use it.

What role can the AI Dialogue play in advancing international cooperation on AI governance?

The AI Dialogue can commission or formally recognise independent civil society audits as part of its evidence base. Policy Genome has developed an open, validated methodology for cross-lingual AI integrity auditing, tested across six major AI systems, three languages, and 126 responses with substantial inter-rater agreement. This work was conducted without access to model internals, using only public APIs, which means it is replicable by any actor globally. The Dialogue can advance cooperation by connecting national regulators, procurement bodies, and civil society auditors around a shared evaluation standard. The methodology exists. The gap is institutional endorsement and a mandate to apply it at scale.

What are some of the existing initiatives, partnerships, or mechanisms that the AI Dialogue should build upon or connect with, and what added value could the AI Dialogue bring?

Policy Genome's "Weaponised Algorithms" audit, EU-funded, peer-reviewed, presented at a NATO-supported expert roundtable in Brussels in December 2025 is directly relevant to the Dialogue's mandate. It is the first published cross-lingual integrity audit of frontier AI systems on conflict disinformation narratives, with open methodology and dataset. The Dialogue should engage with this work as a practical starting point for multilateral benchmarking. The methodology is adaptable to elections, migration, public health, or any high-risk topic where language-specific distortion poses risks. I am available to brief the Co-Chairs, contribute to technical working groups, or support the Dialogue secretariat in developing cross-lingual evaluation standards. Contact: ihor@policygenome.org / policygenome.org

How can different stakeholders contribute to the AI Dialogue? Please share recommendations for the format and structure of the AI Dialogue.

Civil society actors with empirical research capacity should be integrated as technical contributors, not observers. The Dialogue should create a structured track for evidence submission from independent auditors, evaluated on methodological quality rather than organisational size. Pre-session technical briefs, published in advance and referenced explicitly during plenary discussions, would create accountability for whether submitted evidence actually shapes outcomes.

Which voices, communities, or perspectives are currently underrepresented in global discussions on AI governance? How could they be included?

  • Three groups are systematically underrepresented: linguistic communities outside English, who face different AI risk profiles and are rarely in the room
  • civil society from conflict-affected regions, whose direct experience of AI-enabled information manipulation is direct evidence of governance failure
  • and independent technical auditors outside large institutions, who represent an underutilised source of accountability evidence.

What innovative engagement formats could most effectively foster meaningful and dynamic engagement during the AI Dialogue?

Three formats would increase engagement quality. Evidence hearings: civil society auditors present empirical findings directly to government delegations, with an expectation that findings inform outputs. Red-teaming exercises: live cross-lingual AI evaluations conducted during the Dialogue, making abstract governance debates concrete. Open dataset review: publish a shared prompt set before the Dialogue, invite independent evaluations across jurisdictions, compare results during the session. Policy Genome's methodology and prompt set are available under open license and could serve as a starting point.

Please share examples of policies, practices, platforms, or approaches that promote effective AI governance or offer concrete solutions to addressing its challenges.

1

Policy Genome's "Weaponised Algorithms" audit demonstrates one concrete approach: structured, independent, cross-lingual evaluation of AI systems on high-risk topics using open methodology. Six models, three languages, 126 responses, two independent evaluators, Cohen's κ ≥ 0.69. The methodology and dataset are published under open license and are replicable by any actor with API access. Three practices from this work are directly transferable to governance frameworks. Prompt-based benchmarking on documented disinformation narratives. Rather than evaluating AI systems on abstract safety criteria, this approach tests specific, fact-checked claims where a correct answer exists. It produces comparable, auditable results across models and jurisdictions. Three-dimensional scoring: factual accuracy, propaganda inclusion, and tone assessed separately. This captures failure modes - false balance, selective refusal, loaded framing - that single-metric evaluations miss. Deployment-layer transparency. The audit found that at least one major system generated accurate responses that were programmatically overwritten before display. Governance frameworks that evaluate only model outputs miss this. Audits must cover the full system users interact with. These practices are not theoretical. They have been applied, validated, and presented to EU and NATO-affiliated audiences. They are available as a foundation for multilateral benchmarking standards. Contact: ihor@policygenome.org / policygenome.org