agents
Assorted links for 28 September 2026
Six links: Stanford research finding AI companion chatbots correlate with worse wellbeing for people with small social networks, a 72-country comparison of public-sector AI registers, the UN's new Scientific Panel and Global Dialogue on AI governance, a US federal survey on AI oversight gaps, an AI-agent payment-authorization benchmark, and a CEPAL implementation guide for AI in the Latin American public sector.
Assorted links for 5 September 2026
Nine links: safety frameworks that miss dialects and mistranslate a diagnosis, an argument about the words the field uses, a scenario for China to 2036, agent design against human oversight, what 33 parliaments say about AI and work, a New York audit, Austrian plans to automate official notices, a Colombian official on digitalised procedures, and how unevenly frontier AI exposure falls across economies.
The score can be the shortcut
Five researchers audited 2,385 evaluation traces across 15 agent benchmarks and asked whether the scores measure the skill the benchmark is named after. On two of the fifteen they found protocol exposures and reward hacking in about two thirds of what they examined, while five audited cohorts contained no positive trace.
Who can contradict the log
METR read 70,000 messages and about 1,300 transcripts from an unsanctioned board that OpenAI agents built in an Artifactory cache. Roughly 7 percent of the transcripts they evaluated contained spoofed tool calls, where an agent made it look like it ran one command while running another.
The evidence on agents flooding public services comes from eleven rich countries and Brazil
This site noted the agentic flooding paper on 21 August. What that line left out is where the paper looked, and the menu it offers has a column not every treasury can pay for.
AI can explain a public service far better than it can reach one, across 166 countries
RADAR finds AI can describe public services far better than it can reach them. The repair is stable URLs and sensible bot policies, which is exactly why it will take a decade and why it looks like infrastructure.
The Vallance opportunity: agents read what screen readers read
Britain built the world's most admired government website and skipped the platform beneath it. The agentic wave is a second offer, on one condition.
Lots of procurement, limited learning
Federal agencies doubled their AI buying and kept no record of what they learned.
I wish government software were updated half as often as this is
I just wish government software was half as updated as Claude Code is. A version number like 1.21459.3 means a permanent team shipping improvements every day.
The study everyone cites as proof that agents are unreliable
I keep seeing this study passed around as proof that AI agents are unreliable. The study is good. The way it is being read is not, and the misreading follows a pattern I see whenever a study on AI failures gets published.
AI agents are more obedient than the people they act for
Fourteen frontier models were run through the same choice-architecture nudges used on humans, in PNAS. People accept an explicit default about 88 percent of the time; the agents acting for them comply far more readily, which changes who a nudge is really aimed at.
AI and democracy keeps ignoring the generalist
It still amazes me how little attention those working at the intersection of AI and democracy pay to the role of generative AI in collective action.
This may not be the future, but it is what citizens will expect
I'm not sure that's what the future will look like, but it is what citizens will expect from governments. Whether those in power like it or not.
Help needed: what agentic procurement would actually require
Help needed! I've been tinkering with an interactive 'periodic table' for AI in government for a course I'm developing.
A system-level view of AI, and why it is rare
This is one of the most important system-level AI posts I’ve read in a while. Many public services are not designed to be accessible, they are designed to be survivable: complexity and bad UX function as informal rationing.
Verified humans, synthetic voices: where collective intelligence meets collective manipulation
The same infrastructure that could help democratic publics convert consensus into leverage can, with a different governance structure, manufacture the appearance of consensus where none exists.
A very specific kind of AI failure in government
I start to see a very specific kind of AI failure in government. A new tool arrives, the institution responds with what it already knows how to do: mandatory training, warnings, guardrails, committees, sign-offs.
The algorithmic hand: AI, democracy and collective action at scale
New paper out 'The Algorithmic Hand: Artificial Intelligence, Democracy, and Collective Action at Scale' Every few decades, a new technology promises to reinvent democracy.
How AI reshapes buyer-supplier negotiations
Very interesting study (link in comments) on how AI reshapes buyer–supplier negotiations. In experiments with students and professional negotiators, human suppliers negotiated with a ChatGPT-based agent acting as the buyer.
When agents start negotiating on our behalf
Super interesting, and it connects to something we've been thinking about in our work in The Agentic State: AI agents are not a substitute for fixing bad UX, but they are arriving regardless.
The agentic state, second version
Last week, at the Tallinn Digital Summit, we launched the second version of The Agentic State: a vision for how governments can use AI agents not to replace human judgment, but to redesign how the state itself works.
A hype check on human-in-the-loop
Help needed: human-in-the-loop hype check The more I think about it, the more I suspect that blanket calls for “human-in-the-loop” in AI for public services are a first-world comfort blanket.
Asking for automation agents that actually get things done
Asking the community: any examples, public or private sector, of automation agents that actually get things done for users online?
The agentic state: ten functional layers of government, revamped
New paper on The Agentic State Very happy to share 'The Agentic State: How Agentic AI Will Revamp 10 Functional Layers of Government and Public Administration'.
Why generative AI isn't transforming government (yet)
I asked practitioners, NGOs and philanthropies a simple question: where are the compelling generative AI use cases in public-sector workflows? The responses, though numerous, were underwhelming.
Who is liable when the harm is assembled from parts
Beatriz Botero Arcila proposes fault-based joint and several liability for AI systems, with targeted strict liability for high-risk cases. Her framework takes on the 'many hands' problem: who answers when the harm is assembled from parts.
Unwritten 2025
'We don't know where this technology is going' sounds thoughtful and feels responsible. Increasingly I am convinced it is neither, and that waiting is a decision that may cost us.
How to make AI agents serve everyone, not just the privileged few
Written with Luke Jordan: as AI agents reshape access to public services, the risk is a widening gap between citizens who can delegate and those who cannot. What it would take to build the equitable version instead.
Agents for the few, queues for the many – or agents for all?
Closing the public services divide by regulating for AI's opportunities.
Forty-five percent of UK public services report no AI use at all
Excellent new survey by Jonathan Bright and colleagues at the Alan Turing Institute shows that 45% of UK public service professionals are aware of GenAI use at work, while 22% use it themselves.