Natural language processing1 datasets usually label the language they use but often omit which countries or populations are represented. To map this gap, researchers built AtlasNLP, a database of over 13,000 dataset records extracted from academic papers. Using automated text extraction and human audits, they separated the language of a dataset from its geographic origin. They then compared the represented country, whose data is included, against the producer country, where the institutions are located, and analysed coverage across 30 task categories such as machine translation and sentiment analysis.
Language is an incomplete proxy for geography. Brazil and Portugal account for about 69% of Portuguese dataset records, while France, Switzerland, and Canada account for about 63% of the French records. Dataset coverage is highly concentrated, with 121 of 197 countries having ten or fewer records attributed to them. Even when the authors included inferred geographic origin, 74% of the records remain geographically unattributed. Consequently, the authors interpret low or zero country coverage as missing documentation within the academic papers they analysed rather than proof that no local data exists. Explicit country-aware metadata is necessary for observing population-level gaps that language metadata alone can obscure.
An AI model that supports a language is not necessarily built on data from every country that speaks it. Buyers consulting a model card2 see the language but cannot find the specific country. Without geographic metadata, developers cannot measure how a system performs for regional populations. From the outside, missing datasets and missing metadata are indistinguishable.
Today’s links: Assorted links for 2 September 2026.
Footnotes
-
Natural language processing is a branch of computer science that gives machines the ability to read and understand human words. It is used to build software that can translate text, answer questions, or summarise documents. ↩
-
A model card is a short reference document that provides key details about an artificial intelligence system. It is used to explain how the software was built, what it is designed to do, and what its limits are. ↩