# Which primary AI economy datasets can a founder actually reuse?

Canonical: https://mudpie.ai/takes/primary-ai-economy-datasets-for-founders/
Breadcrumb: [Home](https://mudpie.ai/) / [Takes](https://mudpie.ai/takes/) / [Which primary AI economy datasets can a founder actually reuse?](https://mudpie.ai/takes/primary-ai-economy-datasets-for-founders/)
Author: Ali Abouelatta (https://mudpie.ai/authors/ali-abouelatta/)
Published: 2026-09-19
Updated: 2026-09-19
Research type: curated directory
Method: Source-led curation by founder question; every entry records actual access, date, reuse-term status, unit, and limitation.

Most AI data lists are reading lists. A useful one tells you which source can answer the decision in front of you.

The source matters because “AI adoption” can mean model use, workflow deployment, task success, falling cost, or occupational exposure. These datasets do not measure the same thing.

Here is the shortlist I would actually keep.

## The founder's dataset shelf

| Source | Access, date, and terms | What it can answer | What it cannot answer |
| --- | --- | --- | --- |
| [Anthropic Economic Index](https://www.anthropic.com/economic-index) | Landing page last updated June 26, 2026; full datasets are freely available from Anthropic's site. The landing page does not state a redistribution licence, so check terms before republishing raw data. | Which tasks, occupations, countries, and API/consumer patterns appear in Claude use? | It is Claude usage, not the labour market as a whole. |
| [Stanford AI Index 2026 public data](https://hai.stanford.edu/ai-index/2026-ai-index-report) | Report and public-data link are current in the 2026 edition; the public data lives in a linked Google Drive folder. Reuse terms are not stated on the landing page; cite Stanford and verify before redistributing files. | What do the tracked series say about model cost, investment, adoption, agents, and productivity? | It is a compiled index with source-specific methods, not one raw experiment. |
| [METR Time Horizon](https://metr.org/horizon-chart-embed) | Current chart exposes Time Horizon 1.1 and links raw YAML plus the [public analysis repository](https://github.com/METR/eval-analysis-public). The repo says to see its `LICENSE` file; confirm the exact reuse terms before redistribution. | How does measured agent success change with the estimated human time required for a task? | It is a benchmark task suite, not proof of customer demand or business productivity. |
| [U.S. Census BTOS AI data](https://www.census.gov/hfp/btos/data) | Public Census survey and experimental AI supplement; the latest cited collection runs through May 3, 2026. Use the official Census files and attribution; the questionnaire and release history show when wording changed. | Which firm sizes, sectors, and business functions report current or expected AI use, and why are non-users holding back? | Survey answers are not product usage logs, and the question revision creates a time-series break. |
| [OECD.AI OpenRouter data](https://oecd.ai/en/openrouter-data) | Public OECD.AI methodological page; monthly observations are updated quarterly. OECD terms apply, and the page does not promise that raw data can be redistributed. | How do model requests vary by requester country, model developer, agentic/human/mixed classification, tokens, and cost? | It sees OpenRouter traffic, not direct provider APIs, private enterprise use, or all consumer activity. |
| [OECD AIKoD](https://oecd.ai/en/aikod) | Public OECD.AI methodology and visualisation; data covers 2023 onward and is updated regularly. The page calls it an experimental OECD resource; do not assume raw-data redistribution rights. | How do public model catalogues, providers, prices, modalities, and benchmark-linked quality move together? | It misses undisclosed and strictly on-premise models and can lag fast provider changes. |
| [OECD AI Exposure Measure](https://www.oecd.org/en/about/projects/artificial-intelligence-and-future-of-skills.html) | OECD page dated May 26, 2026 for the working paper and XLSX download. OECD terms apply; use the official download and attribution. | Which occupations look closer to current AI capability profiles across cognitive, social, and physical domains? | Exposure is not adoption, job loss, or a startup market by itself. |

## Pick the source by the question

For current task use, start with Anthropic. Its data is close to model interactions, and its connector lets readers query the Index directly. Keep the boundary visible: Claude use is not the whole labour market.

For the U.S. business environment, use Census. It answers “Are firms in this sector using AI, and what are they doing with it?” Its latest BTOS summary reports 17%–20% current use in the December 2025–May 2026 window, with higher use among larger firms. Use the OECD sources below for cross-country request, model, and exposure questions.

For model economics, use Stanford for the long trend and OECD AIKoD for a model/provider catalogue. AIKoD puts price, quality, modality, provider, and update date together, but misses undisclosed and strictly on-premise models.

For task horizon, use METR. Its raw runs and analysis code are closer to an evaluation fixture than a market survey. Do not turn a horizon curve into a willingness-to-pay claim.

For a labour or vertical wedge, use the OECD Exposure Measure as a first filter. It maps capabilities to occupations; you still need the buyer, workflow, data access, and budget.

## My rule for using them

Use one source to answer one question, then add a second source only when it measures a different unit.

For example: Census shows reported finance-and-insurance use, Anthropic shows tasks Claude users bring, and METR shows task difficulty by human time. None proves finance-agent product-market fit.

That is enough to beat a generic “AI is growing” chart. Save the source date, access terms, unit, and question beside every number.


## Author disclosure

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.
