Section 1
Key Findings / Plain Summary
Talvio scores 894 U.S. occupations, held by 96% of US workers, by how much of their work today's AI can help with, and adds those scores up by department for an uploaded roster. The product calls it the Training Priority score; formulas call it TAP (methodology version TPS v2.2). It helps a learning and development team decide which roles to start AI training with, which work activities explain that order, and which AI skills the training should cover.
The headline validation results: the ranking agrees with independent human expert ratings of AI exposure at Spearman ρ = 0.858 across all 894 scored occupations: the full population, not a favorable sample. It is corroborated by observed real-world AI usage (Anthropic Economic Index) at ρ = 0.458 on 479 occupations; notably, that usage data agrees with Talvio's ranking at least as strongly as it agrees with the expert benchmark itself (0.363 on the same set). Structural concordance with the widely cited Felten AIOE index is 0.972, reported as methodological consistency rather than independent proof, because AIOE is built from the same O*NET foundation.
Three properties matter as much as the correlations. The score is deterministic: published constants, no model inference at scoring time, every step arithmetic on public data. It is robust: in a thirteen-variant sensitivity sweep, eleven alternative choices barely move the ranking, and the two that do both change which abilities count as physical and both lower agreement with the expert ratings. And its limits are disclosed: two known blind spots are documented in Section 7 with named examples, not buried.
Scores display on a calibrated 0.5–10 scale. No occupation scores zero, because every job includes some AI-assistable duties. The score is not an individual performance measure, not a job-loss forecast, and small differences between occupations are not meaningful.
Weighted by how many people hold each occupation (BLS, May 2025), the average US worker scores 4.2. Industry averages built this way rank the 20 industry sectors in nearly the same order as the same averages built from the expert ratings (0.94) and in a similar order to a measure built from Copilot use (0.77). Section 5 gives the detail and its limits.
Section 2
What the Score Measures (and What It Does Not)
The construct is augmentation: how reliably today's deployed AI can help a trained person perform the kinds of work an occupation involves. That construct was pinned in writing before scoring. It is deliberately not automation risk (whether AI replaces the job), not adoption (whether the organization has rolled AI out), and not skill gap (whether particular employees need training more than others).
Nearby constructs are easy to confuse, and the confusion is where most AI-workforce numbers go wrong. Task exposure, observed AI usage, and model performance on benchmark deliverables are three different things; Talvio validates against the first two and treats the third (peak-deliverable-quality benchmarks) as a different construct rather than a validation target. Talvio sits in the task-based exposure lineage of Felten, Raj, and Seamans and of Eloundou, Manning, Mishkin, and Rock; it is scoped to training prioritization rather than staffing-impact prediction.
What follows from the construct: a high score means much of the role's work is the kind AI tools can help with, so learning to use them well matters most there. It does not mean the role is at risk, and it does not mean AI can do the role's job. A low score means the role's work is mostly physical or in-person; it does not mean zero opportunity, which is why the scale floors at 0.5 rather than 0.
Section 3
How the Score Is Built
The pipeline is six deterministic steps on public data. For each of the 41 O*NET Generalized Work Activities an occupation performs, Talvio asks which of twelve AI capabilities is best suited to assist it and how mature that capability is today; activities are weighted by O*NET importance × level. The weighted sum is the cognitive base. The base is then discounted by the squared share of the occupation's O*NET ability importance in the psychomotor and physical branches (AI assists thinking and communicating, not lifting and touching), and the result is mapped to the display scale by a fixed, rank-preserving calibration with frozen anchors.
A worked example, Registered Nurses (29-1141.00): the activity-weighted cognitive base is 35.3 of 100; nursing is information-rich (documenting, monitoring, coordinating). But 29% of the occupation's ability importance is psychomotor/physical, so the base is multiplied by (1 − 0.29)² = 0.504, giving a raw score of 1.78, which calibrates to a displayed 3.1.
Two formula choices deserve their rationale stated plainly. The best-fit rule (an activity is scored by its single best capability, not a combination of all twelve) prevents unrelated capabilities from inflating cognitive activities, a failure mode in which every cognitive activity saturates near the top of the scale. The squared gate encodes that heavily physical work disproportionately limits AI assistance: the bodily core of such jobs also makes the surrounding cognitive work embodied and context-bound.
Section 4
Where the Inputs Come From
The occupational data is O*NET, the U.S. Department of Labor's occupational database: Work Activities with importance and level ratings (30.2) and the Abilities taxonomy (30.3) for the physical gate. Both are public, versioned, and independently maintained. Talvio does not author any claim about what a job involves.
The two judgment layers (how well each capability fits each activity, and how mature each capability is today) carry panel provenance rather than single-author opinion. The activity-to-capability matrix (v2) is the median of five blinded raters with distinct professional framings who saw only official O*NET definitions and the pinned construct, never the existing values or the purpose; inter-rater agreement was r = 0.94, and the panel's matrix reproduces the prior human-reviewed ranking at 0.9986. The twelve capability maturity scores (v2) come from a second five-rater panel required to ground every score in named, dated 2025–2026 deployment evidence, with “insufficient evidence” as an allowed answer. The largest correction from that evidence pass: coding assistance, which the blind panel had rated 8.5, its highest rating, fell to 6.5 after all five raters independently surfaced a randomized controlled trial (METR, July 2025) finding experienced developers were ~19% slower with AI assistance on mature codebases while believing they were faster.
Capability contexts and primary benchmark anchors
Each capability context carries a primary public benchmark anchor. The anchors are reference points for tracking capability progress between refreshes; the maturity scores themselves come from the deployed-evidence panel, because benchmark performance and real-world assistance value can diverge (coding is the documented example).
| ID | Capability context | Maturity (v2, 0–10) | Primary benchmark anchor |
|---|---|---|---|
| cap_01 | Written content generation and editing | 7.5 | HELM Capabilities |
| cap_02 | Information synthesis and research | 6.0 | Artificial Analysis Intelligence Index |
| cap_03 | Structured data analysis and quantitative reasoning | 5.5 | Artificial Analysis Intelligence Benchmarking |
| cap_04 | Coding and software engineering | 6.5 | SWE-bench |
| cap_05 | Conversational support and customer interaction | 6.5 | Arena Text Leaderboard |
| cap_06 | Translation and cross-language work | 7.0 | Artificial Analysis Multilingual Index |
| cap_07 | Speech and audio processing | 6.5 | Artificial Analysis Speech to Text Leaderboard |
| cap_08 | Image and document understanding | 7.5 | MMMU-Pro |
| cap_09 | Image, video, and design generation | 7.0 | Artificial Analysis Image Model Leaderboard |
| cap_10 | Planning, scheduling, and structured decision support | 5.0 | GDPval |
| cap_11 | Tool use and autonomous agents | 4.0 | METR Time Horizons |
| cap_12 | Domain-specialist reasoning | 6.5 | GPQA |
Because AI capability is a moving target, maturity scores are re-assessed on a fixed cadence (each major frontier-model release, or quarterly, whichever comes first) using the same evidence-verified panel process. Each refresh re-runs the ablation and regression suite; material changes bump the methodology version. A score without a version stamp should not be trusted, including Talvio's.
Section 5
External Validation
The validation hierarchy is ordered by independence, not by flattery. The primary referent is human-rated exposure: Eloundou et al.'s beta measure, the share of an occupation's tasks where AI assistance cuts completion time by half or more, the closest published construct to Talvio's. Agreement is Spearman ρ = 0.858 on all 894 scored occupations (bootstrap 95% CI [0.839, 0.873]).
The second referent is behavioral: the Anthropic Economic Index's observed AI task-use (release 2026-03-24, US). AEI task records were matched to O*NET task statements (1,303 tasks matched) and rolled up to 479 occupations as a log-scaled usage intensity. Agreement is ρ = 0.458; on the same occupations, AEI agrees with the human-rated benchmark at only 0.363. Real-world usage corroborates Talvio's ranking at least as strongly as it corroborates expert judgment. The claim is deliberately modest: correspondence with observed use, not high prediction of it. Other reasonable ways of joining the usage data (global rather than US records, SOC codes rather than O*NET codes, average rather than summed task shares) give 0.34 to 0.58. In every one, usage agrees more closely with Talvio's ranking than with the expert ratings.
The third referent is structural: Felten AIOE, ρ = 0.972 on 781 occupations, each taking the index value of its 6-digit SOC code (682 distinct codes; one value per code gives 0.982). This is reported last and always with its caveat: AIOE is itself O*NET-structure-derived, so high agreement demonstrates that Talvio's use of the O*NET spine is methodologically consistent with the academic standard, not that an independent measurement agrees.
Industry averages
Occupation scores were weighted by BLS worker counts (Occupational Employment and Wage Statistics, May 2025). For every industry, BLS surveys employers and publishes how many people work in each occupation; Talvio uses those published counts and does not estimate which jobs an industry has. An industry's average is each occupation's score weighted by how many people in that industry hold it. 96.3% of US workers hold a scored occupation; the rest are left out, not imputed. Where O*NET splits a BLS occupation into specialties, the general occupation's score is used, so registered nurses count at 3.1, not at the average of the nurse specialties. Where BLS merges several O*NET occupations, the BLS occupation takes their simple average.
The average US worker scores 4.25. The 20 industry sectors run from 1.5 (Agriculture, Forestry, Fishing and Hunting) to 7.4 (Finance and Insurance).
The same worker counts were applied to three published measures. Rank agreement between the industry averages:
| Industries compared | Expert ratings (Eloundou et al.) | Felten index, language-model version | Copilot use (Microsoft) |
|---|---|---|---|
| Sectors (20) | 0.94 | 0.95 | 0.77 |
| Industries, 3-digit NAICS (85) | 0.92 | 0.94 | 0.78 |
| Industries, 4-digit NAICS (246) | 0.93 | 0.95 | 0.78 |
For single occupations, agreement with Microsoft's AI applicability measure is 0.71 (751 occupations). Microsoft's measure comes from how people use Copilot, which trails what AI could do in some industries.
Four limits apply. Averages agree more closely than single occupations do, because each covers thousands of workers, and most of the agreement reflects how much of each industry's work is office and desk work. The Felten figure is partly circular, as above. Industry averages support ordering claims only, never claims that one industry is some multiple of another. And the BLS survey does not cover the self-employed or farms, so the agriculture sector reflects forestry, logging and farm support services, not farmworkers.
Section 6
Robustness: The Ranking Does Not Depend on Our Choices
Any scoring methodology accumulates discretionary choices, and a fair critic asks whether the answers depend on them. Talvio tests this directly. Each choice was flipped one at a time, and then all of them at once, and the resulting ranking compared to production:
| Variant (one choice flipped) | Median rank shift (of 894) | Rank corr vs v2.2 | ρ vs human benchmark | Verdict |
|---|---|---|---|---|
| weights: IM only | 7.0 | 0.9987 | 0.857 | PASS |
| weights: IM x LV^2 | 7.0 | 0.9989 | 0.857 | PASS |
| gate form: linear k=1 | 9.0 | 0.9981 | 0.858 | PASS |
| gate form: k=3 | 4.0 | 0.9996 | 0.857 | PASS |
| gate set: + sensory | 21.0 | 0.9903 | 0.856 | FAIL |
| gate set: physical only | 48.0 | 0.9598 | 0.815 | FAIL |
| gate weight: raw IM | 10.0 | 0.9973 | 0.855 | PASS |
| addressability: top-2 softOR | 3.0 | 0.9998 | 0.857 | PASS |
| addressability: top-3 mean | 3.0 | 0.9998 | 0.858 | PASS |
| matrix: alternative reviewed values | 7.0 | 0.9989 | 0.861 | PASS |
| capabilities: alternative reviewed set | 2.0 | 0.9999 | 0.858 | PASS |
| capabilities: blind panel | 4.0 | 0.9997 | 0.86 | PASS |
| WORST CASE combo | 14.0 | 0.9955 | 0.862 | PASS |
The bar was set before the sweep ran: a choice passes if flipping it moves the median occupation fewer than 20 of 894 rank positions. Eleven of thirteen variants pass, and nine move the median occupation fewer than 10 positions; even the worst-case combination of every alternative choice shifts it by 14. The one genuinely load-bearing choice is which abilities count as “needs a body.” That choice is justified rather than tuned: the psychomotor + physical definition exactly reproduces the benchmark-validated gates, and both alternatives measurably degrade agreement with the human benchmark (dropping psychomotor: ρ falls from 0.858 to 0.815).
A second robustness result concerns the capability maturity scores. Giving every AI skill the same rating barely changes the ranking. Before the physical gate, rank agreement with the real ratings is 0.9948 (0.9988 and 0.9952 with the two earlier rating sets); after the gate it is 0.9997. The ranking is driven by work-activity structure, importance/level weighting, and the physical gate; capability maturity is training-design context: it describes what to teach, not who ranks first. The product therefore does not use capability maturity to explain the score, and that decision is re-tested at every capability refresh.
Section 7
Known Blind Spots and Limitations
Two systematic errors survive every version of an activity-structure method, and Talvio discloses them with examples rather than patching them cosmetically.
Communication-heavy roles are likely under-rated. Jobs done largely by phone and email (sales representatives, scored 4.7 against a human-rated 7.1; executive assistants; customer service) score below what expert references suggest, because activity labels cannot see that this work flows through channels where AI assistance is strongest. Phone and email intensity measurably predicts this under-rating, and the product's limits name it.
Judgment-heavy roles may be over-rated. Administrative law judges (8.6 vs 3.5) and counselors score above 2023 expert ratings: the structure sees information work and cannot see that high-stakes human judgment resists delegation. The 2026 usage data complicates this in an interesting way: counseling-type work shows among the heaviest real AI use in AEI, so the truth likely sits between Talvio's score and the older expert view.
Beyond the blind spots: heavily physical jobs floor near 0.5–1 rather than zero by design. The primary benchmark is expert judgment from 2023, not a measured outcome. No one in this field, Talvio included, yet has longitudinal evidence that training the highest-scored roles first produces the best results; match-review accuracy tracking is the first step toward outcome data. National occupation profiles are not any single organization: titles and real workflows differ, which is why every analysis passes through human match review with one-click overrides before results are relied on. And ordinal survey inputs cannot support decimal precision: one-decimal display is a readability choice, and small differences between occupations do not mean much.
Section 8
Coverage, Governance, and Versioning
Talvio directly scores 894 occupations and explicitly excludes 122 rows: 77 residual SOC codes without descriptor data, 26 split codes without O*NET Work Activities, and 19 military occupations out of civilian scope by design. Excluded rows remain visible with provenance rather than being silently scored or imputed.
Every input carries a version and a date: O*NET releases, matrix v2 and capabilities v2 (both panel-adopted 2026-07-07), the calibration anchors (frozen; re-fit only at version bumps), and the score itself (TPS v2.2). The methodology is versioned and updated when the evidence warrants, and every revision is recorded. The full record is kept: the pre-registration with its amendment log, both panel studies with raw ratings, the sensitivity sweep, and a calculator workbook with the formulas. The US comparison is rebuilt whenever scores change and each spring when BLS publishes new worker counts.
Product copy is governed by a scoped-claims contract: every validation statement must cite the human-rated benchmark and the observed-use figure together; the structural concordance figure may not appear without its circularity note; and copy implying job-loss prediction or capability-driven ranking is nonconforming by rule.
Section 9
References
External Literature and Data Sources
- Handa, K., Tamkin, A., McCain, M., Huang, S., Durmus, E., Heck, S., Mueller, J., Hong, J., Ritchie, S., Belonax, T., Troy, K. K., Amodei, D., Kaplan, J., Clark, J., & Ganguli, D. (2025). Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations. arXiv:2503.04761. https://arxiv.org/abs/2503.04761
- Anthropic. (2026). Anthropic Economic Index, March 24, 2026 dataset release. Dataset and source registry: https://huggingface.co/datasets/Anthropic/EconomicIndex
- Blinder, A. S., & Krueger, A. B. (2013). Alternative measures of offshorability: A survey approach. Journal of Labor Economics, 31(2), S97-S128. https://doi.org/10.1086/669061
- Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science, 384(6702), 1306-1308. https://doi.org/10.1126/science.adj0998. Data: https://github.com/openai/GPTs-are-GPTs
- Felten, E., Raj, M., & Seamans, R. (2023). How will language modelers like ChatGPT affect occupations and industries? arXiv:2303.01157. https://arxiv.org/abs/2303.01157
- Tomlinson, K., Jaffe, S., Wang, W., Counts, S., & Suri, S. (2025). Working with AI: Measuring the occupational implications of generative AI. arXiv:2507.07935. https://arxiv.org/abs/2507.07935. Data: https://github.com/microsoft/working-with-ai
- U.S. Bureau of Labor Statistics. Occupational Employment and Wage Statistics, May 2025. https://www.bls.gov/oes/
- Felten, E., Raj, M., & Seamans, R. (2021). Occupational, industry, and geographic exposure to artificial intelligence: A novel dataset and its potential uses. Strategic Management Journal, 42(12), 2195-2217. https://doi.org/10.1002/smj.3286
- Felten, Raj, and Seamans AIOE data repository. https://github.com/AIOE-Data/AIOE
- Frey, C. B., & Osborne, M. A. (2017). The future of employment: How susceptible are jobs to computerisation? Technological Forecasting and Social Change, 114, 254-280. https://doi.org/10.1016/j.techfore.2016.08.019
- Kochhar, R. (2023). Which U.S. workers are more exposed to AI on their jobs? Pew Research Center. https://www.pewresearch.org/social-trends/2023/07/26/which-u-s-workers-are-more-exposed-to-ai-on-their-jobs/
- Massenkoff, M., & McCrory, P. (2026). Labor market impacts of AI: A new measure and early evidence. Anthropic. https://www.anthropic.com/research/labor-market-impacts
- National Center for ONET Development. (2014). A Multi-Phase Rational Method for Developing Area Work Activities*. https://www.onetcenter.org/dl_files/DWA_2014.pdf
- National Center for ONET Development. ONET 30.2 Database. https://www.onetcenter.org/database.html
- National Center for ONET Development. ONET License Agreements. https://www.onetcenter.org/license_agreements.html
- OpenAI. (2025). GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. https://openai.com/index/gdpval/ and https://cdn.openai.com/pdf/d5eb7428-c4e9-4a33-bd86-86dd4bcf12ce/GDPval.pdf
- OpenAI GDPval public dataset. https://huggingface.co/datasets/openai/gdpval
- U.S. Council of Economic Advisers. (2024). Potential Labor Market Impacts of Artificial Intelligence: An Empirical Analysis. https://bidenwhitehouse.archives.gov/wp-content/uploads/2024/07/Potential-Labor-Market-Impacts-of-Artificial-Intelligence-An-Empirical-Analysis-July-2024.pdf
- Webb, M. (2019). The Impact of Artificial Intelligence on the Labor Market. SSRN. https://doi.org/10.2139/ssrn.3482150
- METR. (2025, updated 2026). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
- National Center for O*NET Development. O*NET Database, releases 30.2 and 30.3. U.S. Department of Labor. https://www.onetcenter.org/database.html
Capability Benchmark Sources
- Artificial Analysis. Model leaderboards. https://artificialanalysis.ai/
- Artificial Analysis. Intelligence Benchmarking Methodology. https://artificialanalysis.ai/methodology/intelligence-benchmarking
- Artificial Analysis. Speech to Text Leaderboard. https://artificialanalysis.ai/speech-to-text
- Artificial Analysis. Image Model Leaderboard. https://artificialanalysis.ai/text-to-image
- Arena. Text Leaderboard. https://arena.ai/leaderboard
- Center for Research on Foundation Models. HELM Capabilities. https://crfm.stanford.edu/helm/capabilities/latest/
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770. https://arxiv.org/abs/2310.06770
- METR. (2026). Time Horizon 1.1. https://metr.org/blog/2026-1-29-time-horizon-1-1/
- Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., & Bowman, S. R. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv:2311.12022. https://arxiv.org/abs/2311.12022
- SWE-bench leaderboards. https://www.swebench.com/
- Yue, X., Zheng, T., Ni, Y., Wang, Y., Zhang, K., Tong, S., Sun, Y., Yu, B., Zhang, G., Sun, H., Su, Y., Chen, W., & Neubig, G. (2024). MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark. arXiv:2409.02813. https://arxiv.org/abs/2409.02813