The analysis covers all 1,534 written submissions
published by the UN Global Dialogue on AI Governance.
The categories were fixed before the corpus was read: 31 risk
codes in 9 domains, most taken from the MIT AI Risk
Repository1, and 15 forms a governance measure can take,
adapted from OECD and academic taxonomies of policy instruments2.
Each submission was then read independently by three language models from
three different developers, so their mistakes are less likely to line up.
A model cannot simply assert an annotation: every annotation must come with a word-for-word
quote from the submission, and software checks that the quote is actually
there (99.8% of the 36,843 quotes
cited across the corpus were). An annotation is accepted directly when the two
primary models both support it with verified quotes. Contested codes (52%
of the 18,667 decisions) went to a separate arbiter model,
Claude Opus 4.8, which accepted 46% of them, with its
own quote checked the same way. The result is 13,425
annotations.
To measure how accurate that process is, a human labeled a sample of
60 submissions. Against that benchmark, the system
reaches 94% precision and 93% recall
on risks, and 92% precision and 96% recall on governance measures.
Every share on this page is corrected for that measured error rate using a
Rogan-Gladen prevalence correction3. Differences between stakeholder groups
or regions are only reported when they survive a false-discovery correction
(Benjamini-Hochberg4) across every comparison tested, and they should still be
read as exploratory.
Two limits apply to everything above. Participation was self-selected, so
these figures describe what the people and organizations who wrote in chose
to say, not world opinion, and each submission counts once whether it came
from an individual or a ministry. And absence is weak evidence: a submission
that never mentions a risk has not said the risk does not matter. Every
accepted code links back to its quote, and any submission can be audited in
the explorer.