evals
No AI safety index caught a chatbot mistranslating smallpox as syphilis
AI safety evaluation budgets concentrate in a handful of Western labs testing for model-level risks, while deployment failures in low-resource languages go untested until a user reports the harm, including a Tigrinya medical chatbot that rendered smallpox as syphilis and gonorrhea as diabetes.
Generative AI raised homework scores and lowered exam scores for 26,811 Chinese students
A CEPR working paper tracked 26,811 Chinese secondary students for 30 months and found generative AI raised homework scores while lowering exam and entrance-exam performance, concentrated among the students whose usage pattern looked like outsourcing the work.
Assorted links for 14 September 2026
Five links: at least 39 submissions to Australian parliamentary inquiries carrying apparently hallucinated references, a retrieval system whose recall falls from 84 to 16 per cent when a benefits question is asked in plain English, federal AI obligations reaching $7.2 billion, 300,000 hours saved by bots at the Defense Logistics Agency, and Brazil's open platform for reusable public-sector AI.
AI trials involve governments less often than trials in general
Five researchers tracked AI interventions across two trial registries from 2019 into a partial 2026. AI is on track to make up one in five registered trials, up from one in thirty in under four years, the average planned AI trial runs 11.3 months, and AI trials in the AEA registry are government-related 10.7 per cent of the time against 14.1 per cent for trials overall.
Access to an AI assistant built on World Bank reports saved no significant time
Researchers deployed an AI assistant restricted to a curated library of World Bank reports to more than 2,200 professionals across 116 countries. Its refusal rate fell from between 40 and 70 percent in its early weeks to under 10 percent once the library grew from about 50 reports to over 4,000, and the broader experiment found no significant time savings for the overall group given access.
The score can be the shortcut
Five researchers audited 2,385 evaluation traces across 15 agent benchmarks and asked whether the scores measure the skill the benchmark is named after. On two of the fifteen they found protocol exposures and reward hacking in about two thirds of what they examined, while five audited cohorts contained no positive trace.
Brazil and Portugal account for about 69% of Portuguese dataset records
Twenty researchers built the first country-level atlas of who is represented in the datasets behind language models. Brazil and Portugal hold about 69% of the Portuguese records, France, Switzerland and Canada about 63% of the French, and 121 of 197 countries have ten or fewer records attributed to them at all.
Assorted links for 31 August 2026
Nine links: what two frontier labs now book in a year, the three figures a Mexican court got when it asked three chatbots the same question, what a smaller model does when it is handed skills another model worked out, an 825 million euro fine with no AI in it.
Assorted links for 29 August 2026
Nine links: which of England's two planning AI tools is actually live, what MIT's committee found had changed on campus in under three years, an Ebola nowcast in the Democratic Republic of the Congo, and the first evaluation of a proprietary model by someone who never saw its weights.
Who can contradict the log
METR read 70,000 messages and about 1,300 transcripts from an unsanctioned board that OpenAI agents built in an Artifactory cache. Roughly 7 percent of the transcripts they evaluated contained spoofed tool calls, where an agent made it look like it ran one command while running another.
A new gov.uk benchmark finds chatbots usually answer well and almost never say I don't know
Researchers generated 22,066 questions from 2,781 gov.uk pages and tested 11 models on them, scoring each answer claim by claim against the page. Most answers were good, a small tail of bad misses drags the averages down, every model volunteered more than the page held, and almost none ever refused to answer.
The government share in Pew's AI-writing study is about five flagged pages out of 669
Pew ran an AI-detection model over nearly 490,000 web pages and found ten times more AI writing on commercial domains than on government ones. The government figure everyone will quote is drawn from 669 sampled pages, and it has been falling since 2024 while every other domain rises.
Counting where migrants work moves Tajikistan from the 21st to the 75th percentile of AI exposure
A new country-level measure follows AI exposure through migrant work and remittance income. Tajikistan shows why a national labour market does not stop at the border.
Claude's content mark shows that a model touched a text, not whether it supplied the argument
Claude's new mark can show that a model touched a text. It cannot tell whether the model supplied the argument, or only the English through which the argument had to travel.
An AI exposure estimate was checked against vacancy data, and it moved the same way
AI exposure scores are hypotheses, and their value lies in predicting correctly. Here is one that was checked against vacancy data, and moved the same way.
My two cents on AI detectors
My two cents on AI detectors. My worry is not whether they work, but what they measure and how people read it.
The study everyone cites as proof that agents are unreliable
I keep seeing this study passed around as proof that AI agents are unreliable. The study is good. The way it is being read is not, and the misreading follows a pattern I see whenever a study on AI failures gets published.
A place where every claim about AI and work points somewhere
Johanna Einsiedler just launched what the argument about AI and work has been missing: a place where every claim points back to a number on a page.
The interesting part is the volatility
The interesting part is the volatility. A mix that swings this much in twelve months is a warning to anyone signing a multi-year exclusive deal.
Theoretical AI capability against observed usage
This figure from Anthropic comparing “theoretical AI capability” with observed usage across occupations has been circulating widely in the AI policy bubble.
ChatGPT is splitting into work and everything else
OpenAI recently released data on how people use ChatGPT. One number stood out. In December 2025, about 75% of messages from paid Pro users were work-related.
An RCT on patient-facing medical AI, and what it measured
A randomised trial in China (n = 2,069), published in Nature Medicine, tested an LLM chatbot that conducts the patient intake interview before a specialist visit and hands the clinician a structured summary.
We have far more AI policy trackers than AI deployment trackers
I wish we had at least half as many AI deployment trackers as we have AI policy trackers. Especially in contexts where deployments are likely to have major consequences….
Looking for machine learning systems that are actually running in government
Help needed: looking for real-world Machine Learning systems in government A few weeks ago, I reached out to this network asking for compelling GenAI use cases in public-sector workflows.
The UK's Copilot experiment with 20,000 civil servants deserves more attention
The UK's Copilot experiment with 20,000 civil servants deserves way more attention than it's gotten. The results, 26 minutes saved per day, might seem modest, but they reveal something crucial about AI in government.
Where generative AI is actually hitting labour markets
Most studies of generative AI and jobs rest on exposure estimates rather than observed effects. New World Bank research asks where the impact on labour markets is actually landing.
The first clinical trial of a generative AI therapy chatbot
The first clinical trial of a generative AI therapy chatbot, from Dartmouth in NEJM AI: 51% average reduction in depression symptoms, 31% in generalised anxiety, 19% in eating-disorder concerns.