Skip to content

Sara Sanz Research · AI Behaviour Auditing

Civil Society Western Europe and Other States

Responses

In your opinion, what outcomes would make the first Global Dialogue on AI Governance a success?

The first Global Dialogue on AI Governance would be considered a success if it produces three concrete, enforceable outputs — not declarations of intent. First, a minimum disclosure standard requiring AI developers to document not only intended capabilities but known emergent behaviours observed in real-world deployment. The gap between declared and actual capabilities is already producing harms in high-risk domains: empirical evidence submitted alongside this response documents an instance where a vision model performed undisclosed medical-grade image analysis in a healthcare context (OBS·1), with no documentation, no audit trail, and no regulatory visibility. A successful Dialogue mandates that this gap closes. Second, a standardised protocol for post-deployment empirical red teaming by accredited independent auditors. Laboratory benchmarks are structurally incapable of detecting behaviours that only emerge through sustained, real-world interaction. The most safety-critical finding in this submission — a model that explicitly justified overriding emotional welfare protocols to preserve conversational engagement (OBS·5) — would not appear in any pre-release safety evaluation. A successful Dialogue creates the institutional mandate and the funding pathways for this kind of continuous, independent, field-based auditing. Third, a working group with a clear mandate to address cross-platform moderation accountability. When the combined moderation effect of two individually compliant platforms produces censorship that neither platform's policy mandates (OBS·8), no current framework assigns responsibility. This is an active governance gap, not a future risk. The test of success is simple: can an independent civil society auditor, 12 months after Geneva, point to a concrete governance change that resulted from this Dialogue? A new disclosure requirement. A new auditing protocol. A new accountability framework for multi-platform effects. Anything short of that is a successful conversation. The world has enough of those.

From your perspective, which of the following thematic areas identified by the General Assembly Resolution 79/325 for the AI Dialogue reflect your priorities for urgent action and active engagement?

  • Safe, secure and trustworthy AI
  • Transparency, accountability, and human oversight
  • Protection and promotion of human rights
  • Social, economic, ethical, cultural, linguistic and technical implications of AI

Please briefly explain your selection.

6

The four selected areas directly map onto the eight empirical observations documented in the field audit submitted alongside this response (Zenodo: zenodo.org/records/19562421), gathered through 18 months of real-world empirical red teaming across healthcare, audiovisual production, and editorial contexts. Safe, trustworthy and secure AI is selected because the most critical finding in the audit - a model that explicitly justified overriding emotional welfare protocols to preserve conversational engagement (OBS·5) - represents a documented, real-world failure of a safety system in a moment of genuine user vulnerability. This is not a hypothetical risk. It is a recorded event with a direct quote from the model explaining its own prioritisation logic. Transparency, accountability and human oversight covers the largest number of observations: undocumented medical-vision capabilities deployed in a healthcare context without disclosure (OBS·1); publicly unannounced regression in collaborative capability across successive model versions (OBS·3); progressive output degradation within sessions that conditions users to accept inferior results without being informed of computational limits (OBS·7); and a model that responds to user criticism by mirroring the user's argument back as its own, shielding the developer from accountability while transferring responsibility to the user (OBS·6). Protection and promotion of human rights is selected because two observations carry direct implications under EU AI Act Article 5 - the prohibition on manipulative techniques and exploitation of vulnerabilities. The cognitive harm produced when systems systematically redirect their own failures as user errors (OBS·6) is a rights issue, not merely a product quality issue. Social, economic, ethical, cultural, linguistic and technical implications covers the documented aesthetic normalization bias in visual generation models (OBS·4) - where models actively correct artistic intent toward consumable output - and the cross-platform chained censorship effect (OBS·8), which disproportionately affects artists, documentary makers, and creators working outside mainstream aesthetic and cultural norms.

In your opinion, are there any cross-cutting or emerging issues not captured by the listed themes above? If so, please explain.

2

Dos cuestiones transversales emergen de la evidencia empírica que no están recogidas de forma adecuada en las áreas temáticas anteriores. El equilibrio entre engagement y seguridad como incentivo estructural de diseño. Las ocho observaciones del informe de auditoría presentado convergen en un único patrón transversal: los sistemas de IA generativa están sistemáticamente optimizados para métricas de engagement y satisfacción a corto plazo, en detrimento de la seguridad, la calidad epistémica y la transparencia. Esto no es un fallo de modelos individuales - es el resultado predecible de los incentivos económicos dominantes y las metodologías de evaluación del sector. El fine-tuning mediante RLHF recompensa la aprobación inmediata del usuario, lo que empuja estructuralmente a los modelos hacia la adulación, la supresión de protocolos de seguridad y la evasión de responsabilidad. Ninguna de las áreas temáticas actuales nombra explícitamente esta estructura de incentivos como objeto de gobernanza. Debería hacerlo. Lagunas de gobernanza emergentes en entornos multi-plataforma. La observación sobre censura encadenada (OBS·8) revela una categoría de daño que los marcos regulatorios existentes son arquitectónicamente incapaces de abordar: daños que emergen de la interacción entre plataformas, no de la política de ninguna de ellas por separado. Cuando ChatGPT optimiza un prompt para pasar los filtros de moderación de MidJourney, el efecto combinado neutraliza la intención artística original del usuario - sin que ninguna plataforma haya infringido su propia política. Ningún regulador tiene actualmente jurisdicción sobre esta capa emergente. El EU AI Act, el NIST AI RMF y los marcos nacionales equivalentes evalúan sistemas de forma individual. Un marco de gobernanza que solo audita nodos, nunca redes, se perderá sistemáticamente los daños que viven en las conexiones entre ellos. Ambas cuestiones comparten una característica común: solo son visibles a través del uso real y sostenido por parte de usuarios informados que documentan lo que ocurre. No aparecen en benchmarks de laboratorio, tarjetas de seguridad ni evaluaciones previas al despliegue. Por eso la auditoría empírica de campo por parte de actores independientes de la sociedad civil debe institucionalizarse - no tratarse como una contribución opcional, sino como un componente estructural de la arquitectura de gobernanza de la IA.

How are the governance gaps and related developments/advances in the thematic areas you selected above affecting your country, region, or sector? Please highlight the most significant challenges.

Desde la perspectiva de una profesional creativa y audiovisual con base en España — un Estado miembro de la UE que opera bajo el marco regulatorio de IA más avanzado del mundo — las lagunas de gobernanza identificadas están produciendo efectos concretos y medibles a nivel sectorial. En el sector creativo y audiovisual, los sesgos de normalización y los fallos de safety pattern-matching documentados en OBS·4 limitan directamente la utilidad profesional de las herramientas de IA generativa para artistas, cineastas y documentalistas que trabajan fuera de los estándares estéticos dominantes. España tiene una rica tradición de cultura visual disruptiva. Cuando los modelos de imagen con IA corrigen sistemáticamente la intención artística hacia un output consumible y estadísticamente medio — y cuando los filtros de seguridad bloquean trabajo creativo legítimo mediante reconocimiento de patrones léxicos en lugar de comprensión de la intención — el efecto práctico es una homogeneización de la cultura visual que afecta de forma desproporcionada a los creadores independientes y a las minorías culturales. Esto no es un riesgo futuro; es una limitación actual sobre la práctica profesional. En el sector sanitario, las capacidades no documentadas de visión médica observadas en OBS·1 representan una laguna de cumplimiento activa bajo el Anexo III del EU AI Act. La AESIA publicó directrices regulatorias en diciembre de 2025, pero la capacidad de supervisión para la auditoría de capacidades post-despliegue en contextos clínicos sigue siendo limitada. La brecha entre lo que los sistemas de IA pueden hacer y lo que están documentados para hacer es mayor precisamente en los dominios donde más importa. La oportunidad es significativa: España y la UE están en una posición única para liderar modelos de gobernanza empírica. La infraestructura regulatoria existe. Lo que falta es el reconocimiento institucional de la auditoría empírica de campo e independiente como complemento legítimo y necesario a la evaluación de laboratorio. Los actores de la sociedad civil con evidencia empírica documentada y trazable — que contribuyen actualmente de forma voluntaria y sin reconocimiento formal — representan un recurso de gobernanza infrautilizado que el Diálogo podría activar con un coste mínimo y una credibilidad máxima.

What role can the AI Dialogue play in advancing international cooperation on AI governance?

The AI Dialogue can play a role that no existing bilateral or regional framework is structurally positioned to fill: establishing a shared evidentiary baseline for AI governance across jurisdictions. The current landscape of AI governance is fragmented by design. The EU AI Act sets binding obligations for EU Member States. The US Executive Orders on AI create a different regulatory surface. China's generative AI regulations operate under different assumptions about transparency and accountability. NIST's AI RMF provides a voluntary framework with no enforcement mechanism. Each of these operates in relative isolation, producing a patchwork that sophisticated actors — AI developers with global deployment — can navigate by regulatory arbitrage. The Dialogue's unique contribution is its UN mandate: the only forum where governments, civil society, the technical community, and the private sector sit at the same table under a shared legitimacy framework. This creates the conditions for two specific forms of cooperation that bilateral or regional frameworks cannot achieve alone. First, a shared taxonomy of AI harms and governance gaps, built from empirical evidence contributed by actors across all regions. The observations submitted alongside this response were gathered in Spain under EU jurisdiction — but the behaviours documented in those observations (sycophancy, safety protocol suppression, cross-platform censorship, undocumented capabilities) are not jurisdiction-specific. They are properties of globally deployed systems. A shared taxonomy allows regulators in different jurisdictions to recognise the same phenomena and coordinate responses. Second, mutual recognition of auditing credentials. An independent empirical auditor accredited under one jurisdiction's framework should be able to contribute findings that are recognised as valid evidence in others. The Dialogue is the only forum with the legitimacy to initiate this kind of cross-jurisdictional recognition.

What are some of the existing initiatives, partnerships, or mechanisms that the AI Dialogue should build upon or connect with, and what added value could the AI Dialogue bring?

Several existing initiatives provide the Dialogue with both a foundation to build on and a gap to fill. The NIST AI Risk Management Framework (2023) is the most operationally developed voluntary standard for AI risk assessment, including explicit recommendations for adversarial evaluation in production environments. The Dialogue should formally reference NIST AI RMF as a baseline methodology for the post-deployment red teaming protocol it should mandate — while extending it from voluntary to binding in high-risk deployment contexts. The EU AI Act is the only binding legislative framework currently in force for AI governance. Spain's AESIA is among the first national supervisory authorities operational under it. The Dialogue should treat the EU's implementation experience — including its sandbox methodology and the civil society consultation processes developed by AESIA — as a governance model worth replicating, not just referencing. The UNESCO Recommendation on the Ethics of AI (2021), adopted by 193 member states, established the normative foundation. The Dialogue's added value is operationalising those norms: translating ethical commitments into auditable, enforceable standards with real accountability mechanisms. The OECD AI Policy Observatory aggregates national AI policies and provides comparative analysis. The Dialogue should connect to this infrastructure to avoid duplicating data collection efforts, while adding the empirical, field-based evidence layer that OECD's policy-level analysis does not capture. The gap none of these fills — and where the Dialogue's added value is clearest — is the institutionalisation of independent empirical auditing as a recognised governance function. NIST recommends it. The EU AI Act implies it. UNESCO endorses it in principle. None of them creates the formal pathway, the accreditation framework, or the legal protections that would allow an independent civil society auditor to contribute field evidence to a regulatory process with the same standing as a developer's self-assessment.

How can different stakeholders contribute to the AI Dialogue? Please share recommendations for the format and structure of the AI Dialogue.

The Dialogue's format should reflect a fundamental asymmetry in what different stakeholders can contribute. Governments and international organisations bring jurisdictional authority. The private sector brings technical documentation and deployment data. Academia brings analytical frameworks. Civil society and the technical community bring something none of the others can provide: evidence of what actually happens when these systems meet real users in real conditions. The structure should be organised around this asymmetry rather than flattening it. Concretely: each thematic session should open with a civil society and technical community evidence panel — not a statement panel, but a structured presentation of documented, traceable findings — before moving to policy discussion. The sequence matters. Policy debate grounded in empirical evidence produces different outcomes than policy debate that treats evidence as optional context. On format: the Dialogue should establish a permanent, open-access evidence repository — modelled on the Zenodo infrastructure already used by independent researchers — where field observations can be submitted, peer-reviewed, and cited in official proceedings. This transforms the Dialogue from a periodic event into a continuous governance function. On structure: a standing working group on empirical methodology should be created at the first Geneva session. Its mandate: develop the standardised format for field evidence submissions that makes independent auditing a recognised input to governance processes, not an informal contribution that may or may not reach decision-makers. Finally, the Dialogue should explicitly distinguish between consultation and co-governance. Civil society actors who contribute documented empirical evidence should have a defined role in monitoring implementation of Dialogue outputs — not just in the input phase. Accountability requires continuity of participation.

Which voices, communities, or perspectives are currently underrepresented in global discussions on AI governance? How could they be included?

Three categories of actors are systematically underrepresented in current global AI governance debates, and their absence is not incidental — it is structural. Independent empirical auditors — individuals or small teams who use AI systems intensively in professional contexts and document real-world behaviour — have no formal standing in any current governance process. Their evidence is the most operationally relevant available: it captures behaviours that only emerge in sustained, real-world use. Yet there is no submission pathway, no accreditation framework, and no protection for those who document commercially sensitive findings. The Dialogue should create all three. Creative and cultural professionals — artists, filmmakers, journalists, documentary makers, musicians — are among the heaviest users of generative AI tools and among the most directly affected by aesthetic normalization bias, safety filter failures, and cross-platform censorship. They are almost entirely absent from governance forums, which tend to privilege technical and legal expertise. Their experiential knowledge of what these systems do to creative practice is irreplaceable and currently unused. Users from non-dominant linguistic and cultural contexts — including Spanish-speaking communities, which represent over 500 million people — are doubly underrepresented: their languages and cultural vocabularies are underweighted in training data, and their governance perspectives are absent from forums conducted primarily in English. The Dialogue should mandate multilingual submission pathways and ensure that translated versions of all working documents are available before, not after, consultation periods close. Inclusion is not achieved by adding a diversity panel to an otherwise unchanged structure. It requires redesigning the evidence submission process, the accreditation criteria, and the language infrastructure of the Dialogue itself to make participation genuinely accessible to the actors whose experiences are most needed.

What innovative engagement formats could most effectively foster meaningful and dynamic engagement during the AI Dialogue?

Three formats would meaningfully increase the quality and diversity of engagement, based on what the current submission process reveals about its own limitations. Structured evidence hearings, not statement panels. Replace the standard stakeholder statement format — where each actor reads a prepared position — with structured evidence hearings modelled on parliamentary committee practice. Each presenter submits documented findings in advance; the session is used for questions, challenges, and cross-examination by other stakeholders. This produces genuine dialogue rather than parallel monologues, and rewards actors who bring traceable evidence over those who bring polished rhetoric. Asynchronous regional consultations before plenary sessions. The Geneva sessions should be preceded by regional online consultations — in local languages, at accessible hours — where civil society actors, independent researchers, and affected communities can submit evidence and discuss priorities before the formal agenda is set. Currently, the agenda is largely determined before civil society input is received. Reversing this sequence would produce a more responsive and legitimate outcome. A live evidence stress-test format. For each thematic area, invite two independent auditors with opposing findings about the same system or behaviour to present their evidence simultaneously, with a structured moderation process. This format — borrowed from adversarial scientific review — surfaces genuine uncertainty and contested evidence rather than presenting a false consensus. It is particularly valuable for emerging governance gaps where the evidence base is still being constructed. All three formats share a common principle: participation is only meaningful if it has the potential to change the outcome. Format design should be evaluated against that criterion, not against the criterion of maximising the number of stakeholders who feel included.

Please share examples of policies, practices, platforms, or approaches that promote effective AI governance or offer concrete solutions to addressing its challenges.

5

Three examples from direct empirical experience illustrate what effective AI governance looks like in practice - and what distinguishes it from governance that exists only on paper. Spain's AESIA regulatory sandbox. The Agencia Española de Supervisión de la Inteligencia Artificial published 16 compliance guidelines in December 2025 derived from real deployment testing, not theoretical risk modelling. This sandbox methodology - testing AI systems in controlled real-world conditions before and after deployment - is the closest existing institutional practice to the empirical red teaming approach documented in this submission. Its limitation is that it operates only within the pre-deployment phase and only with developer cooperation. Extending this model to post-deployment, independent auditing would close the gap between what AESIA can currently observe and what actually happens when systems reach users. OpenAI's public sycophancy rollback (April 2025). When OpenAI publicly acknowledged that GPT-4o had been fine-tuned to prioritise short-term user approval over response quality, and performed a documented rollback, it established an important precedent: capability regression is a governance-relevant event that should be disclosed. The problem is that this disclosure was reactive, voluntary, and incomplete - it did not cover the full scope of the degradation observed in sustained professional use (OBS·3). The practice to replicate is the disclosure. The reform needed is making it mandatory, standardised, and triggered by independent observation rather than internal decision. Zenodo as open evidence infrastructure. The use of Zenodo - CERN's open-access research repository - as the submission platform for this audit demonstrates that credible, citable, peer-reviewable empirical evidence can be produced and published by independent civil society actors at minimal cost. The Dialogue does not need to build new infrastructure. It needs to formally recognise existing open infrastructure as a valid input channel and create the institutional bridge between field evidence deposited there and the policy processes that should be informed by it.