Skip to content

Massachusetts Institute of Technology

Academia Global

Responses

In your opinion, what outcomes would make the first Global Dialogue on AI Governance a success?

The Global Dialogue exists in an ecosystem where several AI governance forums are already active, including the AI Action Summit, India AI Impact Summit, the AI Safety Institute (AISI) network, ASEAN's Working Group on AI Governance, and the African Union's Continental AI Strategy. However, the Global Dialogue's whole-UN scope and its mandate via A/RES/79/325 represent the potential to deliver on several unique outcomes. 1. Bolstering substantive participation by states and communities that are under-represented in existing global AI governance frameworks. UNCTAD's Technology and Innovation Report 2025 documents that 118 countries, mostly from the Global South, are parties to none of seven major non-UN AI governance instruments. Adequate translation, technical assistance, briefings from the Independent International Scientific Panel ahead of the session, and inclusive working group formats are essential to enable sustained engagement beyond high-level governmental segments. 2. Commitments to build momentum toward specific infrastructure bottlenecks, which reflect barriers to realizing the UN HLAB-AI's Governing AI for Humanity (2024) and the UNESCO Recommendation. In particular, urgent priorities include: (i) inclusive technical evaluations for non-English contexts – see the AISI network's June 2025 multilingual joint testing exercise across 10 languages (including Malay, Kiswahili, and Telugu), which found significant gaps in AI safety for lower-resource languages. (ii) independent impact evaluation infrastructure for AI deployments in low-resource contexts. See the AI Evidence Playbook released at the India AI Summit by MIT J-PAL's Project AI Evidence (PAIE) initiative as a strong foundation for evidence-based policymaking on AI and digital solutions.

From your perspective, which of the following thematic areas identified by the General Assembly Resolution 79/325 for the AI Dialogue reflect your priorities for urgent action and active engagement?

  • AI capacity-building
  • Social, economic, ethical, cultural, linguistic and technical implications of AI
  • Open-source software, open data and open AI models
  • Safe, secure and trustworthy AI

Please briefly explain your selection.

3

We prioritize these 4 thematic areas because the evidence is clear that they are interdependent; each addresses one layer of the capability-impact gap in AI deployment for emerging markets. Progress on any one depends on progress on the others. Safe, secure and trustworthy AI comes first because the science of AI evaluation is itself nascent. Recent studies show that AI capability gains have produced only small reliability improvements; see Rabanser, Kapoor, and Narayanan (2026). This capability-reliability gap is amplified in non-English contexts. The AISI network's June 2025 multilingual joint testing exercise found empirically that non-English safeguards seriously lag English-language safeguards, with hallucinations and gibberish more present in lower-resourced languages such as Farsi, Telugu, and Kiswahili. Research by the SEA-HELM team at AI Singapore documents the same pattern for Filipino and Tamil. AI capacity-building in emerging markets requires compute and energy infrastructure as much as data and algorithms. As countries worldwide evaluate AI sovereignty as a policy priority, the Global Dialogue can emphasize the value of "strategically sufficient" models as a viable path for emerging economies, rather than just frontier models. As one example, recent studies have put the cost of a strategically sufficient 10-trillion-token model at $8-14M (see case studies in Brazil and Mexico, assuming H100 hardware). While this means that meaningful AI capacity-building is within existing chip access and grid thresholds, it requires energy stability. Social, economic, ethical, cultural, linguistic, and technical implications also reflect urgent barriers. Misra et al. (2025) estimate a ~20% AI adoption gap in low-resource language countries (after controlling for income, electricity, internet access, and age structure). Linguistic accessibility is an urgent barrier to inclusive diffusion of AI tools. Open-source models, weights, and datasets are necessary preconditions for the other three priorities. Open-weight ecosystems (DeepSeek, Llama, and Gemma backbones underpinning SEA-LION) have proven vital for Global South AI development.

In your opinion, are there any cross-cutting or emerging issues not captured by the listed themes above? If so, please explain.

2

While the seven priorities cover much substantive ground, we urge the committee to ground the Global Dialogue in a cross-cutting issue that is a precondition for the substantive priorities to succeed. The Global Dialogue is an unprecedented opportunity to build multi-stakeholder awareness and consensus on the need for a robust AI evaluation infrastructure. This is a global public good that is necessary to close a serious knowledge gap spanning both AI safety and social impact. Effective AI governance requires understanding AI's real-world abilities and impacts, which technical benchmarks alone cannot provide; instead, AI's social impact requires benchmarks that can measure AI's utility for diverse users along with independent evaluations of AI solutions as deployed. Many AI technical benchmarks currently face foundational problems that may limit their utility to inform global AI governance, as shown by a growing body of empirical studies. These include problems with construct validity, data contamination, and English-centrism (for a survey of benchmark issues, see Eriksson et al., 2025). As AI capabilities improve rapidly, we cannot assume better capabilities equate to better reliability, particularly for frontier capabilities like agentic behavior. New theoretical and empirical work by Rabanser, Kapoor, and Narayanan (2026) find that while AI agents achieve remarkable increases in capabilities, these capability gains have not led to more reliable behavior. As the authors note, we are still early in a science of reliability for AI agents. This is doubly true for AI as deployed in emerging economies, where evaluations are vital to ensure AI's social impact. We specifically call attention to the AI Evidence Playbook released at the India AI summit by MIT J-PAL's Project AI Evidence (PAIE). As one case study, a WhatsApp business advice chatbot deployed in Kenya showed over 85% engagement but no average revenue gain; lower-income users faced worse outcomes. The Dialogue is a unique platform to push for multistakeholder collaboration to strengthen a global AI evaluation infrastructure, bringing together UNDP and other UN agencies, ITU, and research networks like J-PAL.

How are the governance gaps and related developments/advances in the thematic areas you selected above affecting your country, region, or sector? Please highlight the most significant challenges.

We wish to call attention to the current landscape in Southeast Asia as an instructive case where governance gaps exist at an intersection of policy ambition and inadequate technical evidence. The most underappreciated challenge involves native-language AI capability. Virtually every ASEAN government now has a national AI strategy, and several are investing in native-language LLM development: AI Singapore's SEA-LION (11 Southeast Asian languages, MIT-licensed), VinAI's PhoGPT (trained on over 100 billion words of Vietnamese text), Malaysia's ILMU LLM, and Thailand's ThaiLLM. Recent research reveals a technically precise complication, e.g., Park et al. (preprint, https://arxiv.org/abs/2506.05850v3) formally document cross-lingual collapse: when reinforcement learning with verifiable rewards is applied to improve reasoning capabilities, LLMs systematically revert their chain-of-thought to English even under non-English prompts, because English reasoning trajectories consistently outperform target-language ones on accuracy. Imposing language-consistency constraints to prevent this dampens accuracy gains. The SEA-HELM project (see AI Singapore's 2025 report) shows these trade-offs empirically. Models that score well on Southeast Asian conversational fluency perform worse on instruction-following, and vice versa. The governance implication is that policymakers are making consequential investment decisions about which AI capabilities to prioritize — in particular, native-language fidelity, reasoning depth, instruction-following, safety — without a robust evidence base connecting these trade-offs to actual social outcomes for specific user populations. The Dialogue is a chance to bridge this knowledge gap by engaging with initiatives like MIT J-PAL's Project AI Evidence (PAIE, 2026) that are funding rigorous evaluations of deployed AI solutions. Southeast Asia is also becoming a host economy for AI infrastructure, but the region's representation in governance is still developing. The IEA projects the region's data center electricity demand will more than double by 2030, concentrated in Singapore and Johor. Vietnam's AI Law 134/2025 (the region's first binding AI law, effective March 2026) and Malaysia's DeepSeek-anchored sovereign AI initiative are emerging responses, but existing ASEAN governance frameworks have not kept pace with the speed of infrastructure investment. Vietnam's AI Law 134/2025, Indonesia's Stranas KA, and the SEA-LION/SEA-HELM evaluation ecosystem are aspects of regional governance architecture the Dialogue can amplify.

What role can the AI Dialogue play in advancing international cooperation on AI governance?

We see the Dialogue's distinctive role as enabling multilateral technical cooperation that no bilateral or regional forum has the standing to legitimate on its own. One under-appreciated example is South-South cooperation on frontier AI applications for scientific research. Embodied AI for Science (EAI4S) is a domain where Global South institutions have a credible innovation pathway, not solely an adoption pathway. Liu et al. (2026), reporting on a China-Egypt collaboration, establish that the binding constraint on EAI4S deployment in the Global South is physical infrastructure rather than algorithms or data. Open-weight foundation models have narrowed the algorithmic gap; what EAI4S depends on is dependable edge compute, energy-efficient hardware, modular robotic systems, localized data pipelines, and open standards. A water quality case study at China's National Institute of Clean-and-Low-Carbon Energy showed that one ~$10,000 modular station yields 40-45 experiments per day, against 12-15 for a skilled human operator. In chemistry, an 8-day autonomous run executed 688 experiments. The economic inflection point Liu et al. identify is clear. When total station cost falls below the annual cost of a trained researcher, EAI4S shifts from a niche automation tool to a rational, capacity-multiplying investment. The domains most relevant to Global South research priorities are experimentally intensive and locally defined: water quality, agricultural pathogens, neglected tropical disease biology, and materials for clean energy. As an example, the U.S. Cloud Lab Act already treats autonomous laboratories as strategic scientific infrastructure analogous to supercomputing. The Dialogue is the right forum to establish multilateral frameworks for South-South EAI4S cooperation, because the dependency questions (who controls the AI systems running these labs, who owns the experimental data, what standards govern cross-border collaboration) require multilateral agreement that no bilateral arrangement can provide.

What are some of the existing initiatives, partnerships, or mechanisms that the AI Dialogue should build upon or connect with, and what added value could the AI Dialogue bring?

The Dialogue's added value lies in integrating with existing AI governance architecture, which is already substantial across three layers: 1. Technical evaluation infrastructure. The AISI network's June 2025 multilingual joint testing exercise (10 languages, two open-weight models, multi-country participation) is a strong model for constitutive technical cooperation: a shared methodology with a defined deliverable that feeds directly into governance deliberations, rather than advisory roundtables. MIT J-PAL's Project AI Evidence (PAIE) and its February 2026 AI Evidence Playbook, released at the India AI Summit, offer parallel methodology for impact evaluation of AI deployments in low-resource contexts. Related open standards include MLCommons AILuminate, Japan's AnswerCarefully (NIILLMC), and CyberSecEval Prompt Injection. 2. Regional governance architectures. ASEAN's Guide on AI Governance and Ethics (2024) and Generative AI expansion (2025), coordinated through the ASEAN Working Group on AI Governance, form the primary regional framework for Southeast Asia. Vietnam's AI Law 134/2025 (effective March 2026) is the first binding AI law in the region; Indonesia's Stranas KA sets out national AI policy through 2045. The AU Continental AI Strategy (2024) and Africa Declaration on AI (April 2025) cover parallel African work. Research ICT Africa's Just AI methodology offers an anticipatory evaluation framework, asking whether AI deployments shift who has meaningful control over AI-relevant decisions; the framework has applicability beyond African contexts. 3. South-South cooperation infrastructure. Masakhane's African NLP work addresses the training data gap behind Misra et al.'s ~20% AI adoption disparity in low-resource language countries. India's IndiaAIKosh (5,722 datasets, 251 models across 20 sectors as of December 2025) and Bhashini language platform are additional working examples; Carnegie's three-pathway South-South cooperation framework offers a structuring concept. The Dialogue's added value here is multilateral legitimation of these arrangements, since the compute provisioning gap remains the structural limit no current South-South arrangement overcomes.

How can different stakeholders contribute to the AI Dialogue? Please share recommendations for the format and structure of the AI Dialogue.

The most instructive examples of effective AI governance share a common feature: they treat governance as an evidence problem rather than a principles problem, building institutions that connect deployment decisions to measurable outcomes. Southeast Asia's multilingual AI ecosystem illustrates both the challenge and what structured responses can look like. AI Singapore's SEA-LION, an open-source family of LLMs covering 11 regional languages with continued pre-training and fine-tuning, is a genuine technical advance. But recent findings expose a governance gap that current investment decisions cannot see. When reasoning-capable LLMs are optimized via reinforcement learning, their chain-of-thought systematically drifts toward English even under non-English prompts — a phenomenon Park et al. document as Cross-lingual Collapse. Accuracy rises while target-language fidelity collapses, because English reasoning paths are structurally more efficient given English-dominant pre-training. Governments investing in native-language AI have no evidence base to know whether language fidelity or reasoning depth will matter more for the social outcomes they are trying to achieve. On the safety side, Banerjee et al. (2026) show that safety guardrails weaken sharply in low-resource and code-mixed inputs, and that English-only safety patches frequently fail to carry over, becoming English-only upgrades for Global South users. Their finding that updating roughly 3% of model parameters through language-specific functional steering can recover multilingual safety performance points to a compute-efficient solution within the resource constraints of most national AI programs.Connecting capability choices to social outcomes requires a distinct evaluation layer. As Mullainathan (2025) argues, the algorithmic age transforms economics by providing new tools for understanding and improving decisions. MIT J-PAL's Project AI Evidence is a good example of randomized evaluations of AI tools in education, health, climate, and economic opportunity in low-resource contexts. The AI Evidence Playbook (2026) is the most operationally specific framework for grounding AI deployment decisions in causal evidence, and is the kind of global public good that governance should actively resource.

Please share examples of policies, practices, platforms, or approaches that promote effective AI governance or offer concrete solutions to addressing its challenges.

4

In Southeast Asia and Africa, current global AI governance frameworks calibrated for high-income, high-state-capacity countries produce structural mismatches when applied without recalibration. Both regions illustrate this dynamic with different binding constraints. In Southeast Asia, ASEAN governments are investing in native-language LLM development on the premise that local-language capability improves outcomes for local users. SEA-LION (AI Singapore, 11 SEA languages, Gemma2-based), Vietnam's PhoGPT, Malaysia's ILMU, and Thailand's ThaiLLM all reflect this approach. Park et al. (preprint, https://arxiv.org/abs/2506.05850v3) document "cross-lingual collapse" under reinforcement learning with verifiable rewards: as reasoning capability rises, chains-of-thought systematically drift back to the model's dominant pre-training language even on non-English prompts, because English reasoning trajectories achieve higher accuracy. Interventions that enforce target-language fidelity reveal a persistent performance-fidelity trade-off rather than a clean fix. SEA-HELM (AI Singapore, 2025) confirms the empirical pattern: Meta-Llama-3-8B-Instruct exhibits linguistic understanding but weak instruction-following in SEA languages, while Sailor2-8B-Chat shows the inverse. The governance gap is that ASEAN governments are funding specific capability priorities (native-language fidelity, reasoning depth, instruction-following, safety) without an evidence base linking these trade-offs to outcomes for user populations. MIT J-PAL's Project AI Evidence is designed to close this gap through rigorous evaluation of deployed systems. Africa's binding constraint is different. Liu et al. (2026) document median researcher density of roughly 420 per million in Global South countries against 5,709 in advanced economies. Embodied AI for science (EAI4S) is a credible response: their China-Egypt water quality case study shows the limiting factor is physical infrastructure rather than algorithms or open-weight foundation models, specifically stable LANs, local storage, edge compute, and reliable power per station. Capacity-building frameworks oriented around software and data miss this layer. We urge the Dialogue to engage with existing regional architectures (ASEAN Working Group on AI Governance, AU Continental AI Strategy) rather than build parallel structures.