OPEN DATA INSIGHTS:
MAKING OFFICIAL STATISTICS WORK IN AN AI-DRIVEN WORLD
Edition: September 2026
Photo: Shutterstock.com
This brief is part of Open Data Watch’s Open Data Insights Series, which explores key topics in open data in greater depth, drawing on findings from the Open Data Inventory to support more effective and inclusive data systems. The primary author is Jamison Henninger, with input from the Open Data Watch team.
Today, a person looking for a country’s official unemployment rate or population figures may get an answer without ever visiting a government website or data portal. According to the Organisation for Economic Co-operation and Development (OECD), a growing share of citizens, journalists, and policymakers are accessing official statistics through chatbots and AI assistants rather than official government or statistical office websites. The World Bank has raised a related concern. General-purpose AI systems frequently draw on general internet content and web search results rather than prioritizing trusted statistical sources such as national statistical offices. If official statistics are not easy to find and access, they risk being overlooked in favor of other information that is easier for AI tools to retrieve and use. How, then, can governments make sure the statistics they produce continue to reach users as the way people search for information changes?
While AI changes how users reach official statistics, it raises familiar questions about whether those statistics are easy to find, understand, and use. Open Data Watch’s Open Data Inventory (ODIN) has tracked how well countries make official statistics available and usable for more than a decade. These findings are increasingly relevant in the age of AI. Countries have made strong progress in making data easier to access, including through machine-readable formats and better download options. But progress has been slower in providing the metadata and documentation needed to understand and use those data effectively, and gaps remain in data coverage and open licensing. ODIN shows that years of investment in open data have created an important foundation for AI readiness in official statistics. At the same time, emerging AI-readiness guidance is raising the bar, requiring stronger metadata, documentation, licensing, and data governance to ensure that official statistics can be used reliably and responsibly through AI tools.
WHAT DOES AI READINESS MEAN FOR OFFICIAL STATISTICS?
AI readiness refers to whether official statistics are prepared for an environment in which users access information through AI tools. There are multiple dimensions to AI readiness. The United Nations Statistics Division (UNSD) frames AI readiness for official data and statistics around data and metadata quality, discoverability, accessibility, interoperability, dissemination, and governance.
Some of the requirements for making official statistics AI-ready build directly on established international open data principles. Data need to be publicly available in machine-readable formats that computers can process, with practical ways to access and download them. Metadata need to provide enough information to understand what is being measured and to interpret and use the statistics correctly. Under open data principles, data also need clear open licenses that establish whether and how the data can be used, reused, and redistributed, including any conditions that apply. Recommendations from the World Bank and the Committee for the Coordination of Statistical Activities (CCSA) similarly connect AI-ready dissemination with machine-readable data, application programming interfaces (APIs), comprehensive metadata, and clear licensing. These are areas where ODIN can provide useful evidence of how statistical systems are currently performing.
WHAT ODIN SHOWS ABOUT DATA READINESS FOR AI
Several of the AI-readiness requirements described above overlap directly with what ODIN already measures. ODIN’s Coverage Index assesses whether relevant statistics are publicly available and whether they include relevant disaggregation, recent and historical data, and geographic detail. ODIN’s Openness Index assesses whether published data are available in machine-readable and non-proprietary formats, can be accessed through useful download options, are accompanied by reference metadata, and have an open license.
Data availability.
AI tools can only retrieve official statistics that have been published. In the 2024/25 ODIN assessment, the average overall coverage score across all 197 countries was 53 out of 100, the first time the average exceeded 50 in ODIN’s history. Substantial gaps nevertheless remain in the availability and coverage of official statistics. Better search or AI tools can make existing statistics easier to reach, but they cannot provide official statistics that have not been made available. So, before data can be AI-ready, they have to be available.
Data openness.
For data that are available, ODIN results show substantial improvements in how easily they can be accessed and processed. Looking at a consistent set of 171 countries assessed in 2018, 2020, 2022, and 2024, scores for machine readability, download options, and licensing increased considerably, while metadata availability improved much more slowly.
Figure 1. Metadata availability has improved more slowly than other openness measures, 2018-2024
Source: Open Data Watch calculations using ODIN assessments. Average scores for the 171 countries included in all four assessments.
Between 2018 and 2024, average licensing scores increased by 18.6 points, machine readability by 16.9 points, and download options by 14.2 points. Metadata availability increased by only 4.0 points. In other words, countries have made considerably more progress on several other aspects of data openness than on consistently providing the information users need to understand and use those data correctly—a gap that becomes more important when AI systems are expected to retrieve and interpret statistics on users’ behalf.
Metadata.
The slower progress on metadata deserves particular attention. ODIN assesses whether three basic pieces of reference information are available: a definition of the indicator, the date the data was made publicly available, and the data source.
Table 1. Basic metadata remain uneven across available ODIN datasets, 2024/25
| Metadata information | Share of records (%) |
| Data Source | 87.5 |
| Publication Date | 51.6 |
| Indicator Definition | 44.0 |
| All three pieces of information | 25.0 |
Source: Open Data Watch calculations from 14,136 available dataset records in the 2024/25 ODIN assessment; these records are not the universe of official statistics published by countries.
Only one quarter of the datasets assessed in ODIN 2024/25 included all three basic reference items. Yet current AI-readiness guidance calls for substantially richer metadata. World Bank and CCSA recommendations emphasize comprehensive machine-readable metadata, recognized standards and code lists, stronger information on provenance and versions, and better connections between data and metadata. PARIS21 has similarly emphasized structured, machine-readable metadata to improve discoverability and interoperability.
These recommendations are built on standards that the international statistical community has promoted for years. The Statistical Data and Metadata Exchange (SDMX) initiative was established in 2001 to create common technical and statistical standards for data and metadata exchange, and the UN Statistical Commission recognized SDMX as the preferred standard for exchanging and sharing data and metadata in 2008. Its standardized representations of statistical concepts, dimensions, classifications, and relationships is particularly valuable in an AI environment because they give users, including AI agents, more consistent information about how statistics are structured and what they represent.
ODIN does not require this higher level of metadata maturity. A perfect ODIN metadata score means that basic reference information is available; it does not mean that the metadata are structured, standardized, or machine-readable. Statistical systems are therefore still working toward consistently providing basic metadata while AI readiness is raising expectations for metadata that machines can process and use directly.
APIs.
ODIN also considers several download options that make published data easier to access, including bulk downloads, customizable downloads, and APIs. APIs are particularly relevant to AI because they provide a structured way for software to request data directly from an official source. In 2024/25, 20 percent of available ODIN dataset records were accessible through an API. The International Monetary Fund’s StatGPT shows how an API can be paired with an AI assistant: users ask for statistics in natural language, the AI interprets the request and translates it into a structured query, and the requested statistics are retrieved through the API. This allows users who do not know how to work with an API directly to access official statistics through an AI interface. The existence of an API does not mean that general-purpose AI assistants will automatically use it, but it provides the machine-accessible infrastructure that AI applications can be built or configured to use.
Model Context Protocol (MCP).
MCP servers provide one emerging way to build on this infrastructure. An API gives software structured access to official data, while an MCP server gives AI applications a standardized way to understand what the API offers and how to query it. This can provide AI tools with a more direct connection to official statistics without requiring users to know how to interact with the underlying API themselves. The use of MCP for official statistics remains limited. An August 2026 ODW update to an initial tally by UNICEF Chief Statistician João Pedro Azevedo identified five officially maintained national statistical or data agency MCP servers in the United States, United Kingdom, France, India, and Singapore. At the subnational level, Idescat operates one for Catalonia, and community-built servers exist for additional countries. These early examples illustrate why existing API infrastructure is increasingly important for AI-ready dissemination: APIs provide the underlying machine access that AI tools such as MCP can build on.
Open licensing.
Open licensing is another established open-data principle that becomes more important when official statistics are accessed and redistributed through third-party services. Clear open licenses tell users and redistributors whether and how data may be used, reused, and redistributed, including conditions such as attribution. The World Bank and CCSA recommendations therefore include clear licensing as part of AI-ready dissemination and call on third-party redistributors and AI developers to respect the terms attached to official data. Yet substantial gaps remain. In the 2024/25 ODIN assessment, ODIN found no license for about 42 percent of datasets, and in 40 of the 197 countries assessed, ODIN found no license for any available dataset. Licensing also varied across official sources: among the 112 countries where ODIN found at least one openly licensed dataset, 67 countries also published data under non-open licenses. This means that the conditions for reusing official statistics can differ depending on which government source publishes the data. Current AI-readiness recommendations add another consideration: licensing and attribution information should be available in machine-readable form, so automated systems that retrieve or redistribute official statistics can identify the applicable license and preserve information such as attribution, source, and version. This does not require replacing existing open licenses: standard licenses such as Creative Commons can already be accompanied by machine-readable licensing information.
WHERE AI READINESS RAISES THE BAR
ODIN measures important conditions for making official statistics available and open. But making those statistics easier to access and use through AI tools requires more than meeting those conditions.
Reliable machine-to-machine access.
Having an API is an important starting point, but external systems also need to be able to rely on it as data change. Current World Bank and CCSA recommendations emphasize stable, versioned, and well-documented APIs, as well as ways to communicate updates, revisions, and new versions. This helps ensure that systems connected to official data can recognize when statistics have changed and retrieve the current version rather than continue using outdated data. AI readiness also raises expectations for metadata. Beyond being available, metadata increasingly need to be structured and machine-readable so AI systems can more reliably discover, interpret, and use official statistics.
Making sure AI tools use official data correctly.
Even when an AI tool is connected to official data, it may still select the wrong series, unit, geography, time period, or version, or present the result without accurately identifying its source. Current recommendations therefore go beyond making data accessible. They call for AI tools used to disseminate official statistics to rely on official databases and validated metadata, cite the source of numerical results, and be tested to ensure that they retrieve and present statistics correctly. Where possible, users should also be able to trace a result back to the original official source.
Making official statistics visible in external AI tools.
Statistical offices can improve their own data and build AI tools that connect directly to them such as MCP servers, but many users will rely on general-purpose AI assistants and search tools that governments do not control. Making official data technically accessible does not guarantee that these services will find or prioritize them. Current recommendations therefore call for closer collaboration between official data producers and AI platforms, search engines, and other data intermediaries on how official statistics are discovered, cited, updated, and corrected. This builds on earlier work by Open Data Watch and PARIS21 that argued that statistical and technology communities need to work together to ensure that quality official statistics remain visible as AI changes how people find information.
Meet users where they are
The user-centered premise behind ODIN therefore remains the same even as the way people find information changes. Official statistics should reach users where they are and give them what they need to use the data correctly and confidently. AI does not replace the principles of open data; it makes them more important and raises the standard for putting them into practice.







