Section 1
Key Findings / Plain Summary
Talvio scores every U.S. occupation by how much of its duties would benefit from AI assistance today, and rolls those scores up to the departments in an uploaded roster. The score, Talvio Augmentation Potential (TAP, methodology version TPS v2.2), is a training-prioritization signal: it tells a learning and development team where structured AI enablement is most worth examining first, which work activities explain that priority, and which AI capabilities should shape the curriculum.
The headline validation results: the ranking agrees with independent human expert ratings of AI exposure at Spearman ρ = 0.858 across all 894 scored occupations: the full population, not a favorable sample. It is corroborated by observed real-world AI usage (Anthropic Economic Index) at ρ = 0.458 on 479 occupations; notably, that usage data agrees with Talvio's ranking at least as strongly as it agrees with the expert benchmark itself (0.363 on the same set). Structural concordance with the widely cited Felten AIOE index is 0.972, reported as methodological consistency rather than independent proof, because AIOE is built from the same O*NET foundation.
Three properties matter as much as the correlations. The score is deterministic: published constants, no model inference at scoring time, every step reproducible from public data. It is robust: a thirteen-variant sensitivity sweep shows no discretionary methodology choice materially moves the ranking. And its limits are disclosed: two known blind spots are documented in Section 7 with named examples, not buried.
Scores display on a calibrated 0.5–10 scale. No occupation scores zero, because every job includes some AI-assistable duties. The score is not an individual performance measure, not a job-loss forecast, and small differences between occupations are not meaningful.
Section 2
What the Score Measures (and What It Does Not)
The construct is augmentation: how reliably today's deployed AI can help a trained person perform the kinds of work an occupation involves. That construct was pinned in writing before scoring. It is deliberately not automation risk (whether AI replaces the job), not adoption (whether the organization has rolled AI out), and not skill gap (whether particular employees need training more than others).
Nearby constructs are easy to confuse, and the confusion is where most AI-workforce numbers go wrong. Task exposure, observed AI usage, and model performance on benchmark deliverables are three different things; Talvio validates against the first two and treats the third (peak-deliverable-quality benchmarks) as a different construct rather than a validation target. Talvio sits in the task-based exposure lineage of Felten, Raj, and Seamans and of Eloundou, Manning, Mishkin, and Rock; it is scoped to training prioritization rather than staffing-impact prediction.
What follows from the construct: a high score means employees in that role would likely get substantial leverage from learning to use AI tools well. It does not mean the role is at risk, and it does not mean AI can do the role's job. A low score means the role's work is mostly physical or in-person; it does not mean zero opportunity, which is why the scale floors at 0.5 rather than 0.
Section 3
How the Score Is Built
The pipeline is six deterministic steps on public data. For each of the 41 O*NET Generalized Work Activities an occupation performs, Talvio asks which of twelve AI capabilities is best suited to assist it and how mature that capability is today; activities are weighted by O*NET importance × level. The weighted sum is the cognitive base. The base is then discounted by the squared share of the occupation's O*NET ability importance in the psychomotor and physical branches (AI assists thinking and communicating, not lifting and touching), and the result is mapped to the display scale by a fixed, rank-preserving calibration with frozen anchors. All constants are published in methodology/tps_v2_2_config.json.
A worked example, Registered Nurses (29-1141.00): the activity-weighted cognitive base is 35.3 of 100; nursing is information-rich (documenting, monitoring, coordinating). But 29% of the occupation's ability importance is psychomotor/physical, so the base is multiplied by (1 − 0.29)² = 0.504, giving a raw score of 1.78, which calibrates to a displayed 3.1. The same walk-through, with live formulas, ships in the score-calculator workbook.
Two formula choices deserve their rationale stated plainly. The best-fit rule (an activity is scored by its single best capability, not a combination of all twelve) prevents unrelated capabilities from inflating cognitive activities, a failure mode in which every cognitive activity saturates near the top of the scale. The squared gate encodes that heavily physical work disproportionately limits AI assistance: the bodily core of such jobs also makes the surrounding cognitive work embodied and context-bound.
Section 4
Where the Inputs Come From
The occupational data is O*NET, the U.S. Department of Labor's occupational database: Work Activities with importance and level ratings (30.2) and the Abilities taxonomy (30.3) for the physical gate. Both are public, versioned, and independently maintained. Talvio does not author any claim about what a job involves.
The two judgment layers (how well each capability fits each activity, and how mature each capability is today) carry panel provenance rather than single-author opinion. The activity-to-capability matrix (v2) is the median of five blinded raters with distinct professional framings who saw only official O*NET definitions and the pinned construct, never the existing values or the purpose; inter-rater agreement was r = 0.94, and the panel's matrix reproduces the prior human-reviewed ranking at 0.9987. The twelve capability maturity scores (v2) come from a second five-rater panel required to ground every score in named, dated 2025–2026 deployment evidence, with “insufficient evidence” as an allowed answer. The flagship correction from that evidence pass: coding assistance, rated highest by every prior assessment, was cut from 8.4 to 6.5 after all five raters independently surfaced a randomized controlled trial (METR, July 2025) finding experienced developers were ~19% slower with AI assistance on mature codebases while believing they were faster.
Capability contexts and primary benchmark anchors
Each capability context carries a primary public benchmark anchor. The anchors are reference points for tracking capability progress between refreshes; the maturity scores themselves come from the deployed-evidence panel, because benchmark performance and real-world assistance value can diverge (coding is the documented example).
| ID | Capability context | Maturity (v2, 0–10) | Primary benchmark anchor |
|---|---|---|---|
| cap_01 | Written content generation and editing | 7.5 | HELM Capabilities |
| cap_02 | Information synthesis and research | 6.0 | Artificial Analysis Intelligence Index |
| cap_03 | Structured data analysis and quantitative reasoning | 5.5 | Artificial Analysis Intelligence Benchmarking |
| cap_04 | Coding and software engineering | 6.5 | SWE-bench |
| cap_05 | Conversational support and customer interaction | 6.5 | Arena Text Leaderboard |
| cap_06 | Translation and cross-language work | 7.0 | Artificial Analysis Multilingual Index |
| cap_07 | Speech and audio processing | 6.5 | Artificial Analysis Speech to Text Leaderboard |
| cap_08 | Image and document understanding | 7.5 | MMMU-Pro |
| cap_09 | Image, video, and design generation | 7.0 | Artificial Analysis Image Model Leaderboard |
| cap_10 | Planning, scheduling, and structured decision support | 5.0 | GDPval |
| cap_11 | Tool use and autonomous agents | 4.0 | METR Time Horizons |
| cap_12 | Domain-specialist reasoning | 6.5 | GPQA |
Because AI capability is a moving target, maturity scores are re-assessed on a fixed cadence (each major frontier-model release, or quarterly, whichever comes first) using the same evidence-verified panel process. Each refresh re-runs the ablation and regression suite; material changes bump the methodology version. A score without a version stamp should not be trusted, including Talvio's.
Section 5
External Validation
The validation hierarchy is ordered by independence, not by flattery. The primary referent is human-rated exposure: Eloundou et al.'s beta measure, the share of an occupation's tasks where AI assistance cuts completion time by half or more, the closest published construct to Talvio's. Agreement is Spearman ρ = 0.858 on all 894 scored occupations (bootstrap 95% CI [0.844, 0.875] on the v2.0 baseline).
The second referent is behavioral: the Anthropic Economic Index's observed AI task-use (release 2026-03-24, US). AEI task records were matched to O*NET task statements (1,303 tasks matched) and rolled up to 479 occupations as a log-scaled usage intensity. Agreement is ρ = 0.458; on the same occupations, AEI agrees with the human-rated benchmark at only 0.363. Real-world usage corroborates Talvio's ranking at least as strongly as it corroborates expert judgment. The claim is deliberately modest: correspondence with observed use, not high prediction of it.
The third referent is structural: Felten AIOE, ρ = 0.972 on 781 SOC6 occupations. This is reported last and always with its caveat: AIOE is itself O*NET-structure-derived, so high agreement demonstrates that Talvio's use of the O*NET spine is methodologically consistent with the academic standard, not that an independent measurement agrees.
Section 6
Robustness: The Ranking Does Not Depend on Our Choices
Any scoring methodology accumulates discretionary choices, and a fair critic asks whether the answers depend on them. Talvio tests this directly. Each choice was flipped one at a time, and then all of them at once, and the resulting ranking compared to production:
| Variant (one choice flipped) | Median rank shift (of 894) | Rank corr vs v2.2 | ρ vs human benchmark | Verdict |
|---|---|---|---|---|
| weights: IM only | 7.0 | 0.9987 | 0.857 | PASS |
| weights: IM x LV^2 | 7.0 | 0.9989 | 0.857 | PASS |
| gate form: linear k=1 | 9.0 | 0.9981 | 0.858 | PASS |
| gate form: k=3 | 4.0 | 0.9996 | 0.857 | PASS |
| gate set: + sensory | 21.0 | 0.9903 | 0.856 | FAIL |
| gate set: physical only | 48.0 | 0.9598 | 0.815 | FAIL |
| gate weight: raw IM | 10.0 | 0.9973 | 0.855 | PASS |
| addressability: top-2 softOR | 3.0 | 0.9998 | 0.857 | PASS |
| addressability: top-3 mean | 3.0 | 0.9998 | 0.858 | PASS |
| matrix: alternative reviewed values | 7.0 | 0.9989 | 0.861 | PASS |
| capabilities: alternative reviewed set | 2.0 | 0.9999 | 0.858 | PASS |
| capabilities: blind panel | 4.0 | 0.9997 | 0.86 | PASS |
| WORST CASE combo | 14.0 | 0.9955 | 0.862 | PASS |
Eleven of thirteen variants move the median occupation by single-digit rank positions out of 894; even the worst-case combination of every alternative choice shifts the median occupation by 14 positions. The one genuinely load-bearing choice is which abilities count as “needs a body.” That choice is justified rather than tuned: the psychomotor + physical definition exactly reproduces the benchmark-validated gates, and both alternatives measurably degrade agreement with the human benchmark (dropping psychomotor: ρ falls from 0.858 to 0.815).
A second robustness result concerns the capability maturity scores. Replacing them with a uniform constant changes the ranking almost not at all (ρ = 0.9948), a result stable across three independently constructed maturity sets. The ranking is driven by work-activity structure, importance/level weighting, and the physical gate; capability maturity is training-design context: it describes what to teach, not who ranks first. The product's capability_layer_explains_score flag is false and is re-tested at every capability refresh.
Section 7
Known Blind Spots and Limitations
Two systematic errors survive every version of an activity-structure method, and Talvio discloses them with examples rather than patching them cosmetically.
Communication-heavy roles are likely under-rated. Jobs done largely by phone and email (sales representatives, scored 4.7 against a human-rated 7.1; executive assistants; customer service) score below what expert references suggest, because activity labels cannot see that this work flows through channels where AI assistance is strongest. Phone/email intensity measurably predicts this under-rating, so affected occupations are annotated in the product rather than silently mis-ranked.
Judgment-heavy roles may be over-rated. Administrative law judges (8.6 vs 3.5) and counselors score above 2023 expert ratings: the structure sees information work and cannot see that high-stakes human judgment resists delegation. The 2026 usage data complicates this in an interesting way: counseling-type work shows among the heaviest real AI use in AEI, so the truth likely sits between Talvio's score and the older expert view.
Beyond the blind spots: heavily physical jobs floor near 0.5–1 rather than zero by design. The primary benchmark is expert judgment from 2023, not a measured outcome. No one in this field, Talvio included, yet has longitudinal evidence that training the highest-scored roles first produces the best results; match-review accuracy tracking is the first step toward outcome data. National occupation profiles are not any single organization: titles and real workflows differ, which is why every analysis passes through human match review with one-click overrides before results are relied on. And ordinal survey inputs cannot support decimal precision: scores carry an uncertainty band, and one-decimal display is a readability choice, not a precision claim.
Section 8
Coverage, Governance, and Versioning
Talvio directly scores 894 occupations and explicitly excludes 122 rows: 77 residual SOC codes without descriptor data, 26 split codes without O*NET Work Activities, and 19 military occupations out of civilian scope by design. Excluded rows remain visible with provenance rather than being silently scored or imputed.
Every input carries a version and a date: O*NET releases, matrix v2 and capabilities v2 (both panel-adopted 2026-07-07), the calibration anchors (frozen; re-fit only at version bumps), and the score itself (TPS v2.2). The methodology is versioned and updated when the evidence warrants; every revision is documented in the methodology changelog. The full revision record (pre-registration with amendment log, both panel studies with raw ratings, the sensitivity sweep, and a formula-transparent calculator workbook) is retained in the methodology archive.
Product copy is governed by a scoped-claims contract: every validation statement must cite the human-rated benchmark and the observed-use figure together; the structural concordance figure may not appear without its circularity note; and copy implying job-loss prediction or capability-driven ranking is nonconforming by rule.
Section 9
References
External Literature and Data Sources
- Handa, K., Tamkin, A., McCain, M., Huang, S., Durmus, E., Heck, S., Mueller, J., Hong, J., Ritchie, S., Belonax, T., Troy, K. K., Amodei, D., Kaplan, J., Clark, J., & Ganguli, D. (2025). Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations. arXiv:2503.04761. https://arxiv.org/abs/2503.04761
- Anthropic. (2026). Anthropic Economic Index, March 24, 2026 dataset release. Dataset and source registry: https://huggingface.co/datasets/Anthropic/EconomicIndex
- Blinder, A. S., & Krueger, A. B. (2013). Alternative measures of offshorability: A survey approach. Journal of Labor Economics, 31(2), S97-S128. https://doi.org/10.1086/669061
- Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2023). GPTs are GPTs: An early look at the labor market impact potential of large language models. arXiv:2303.10130. https://arxiv.org/abs/2303.10130
- Felten, E., Raj, M., & Seamans, R. (2021). Occupational, industry, and geographic exposure to artificial intelligence: A novel dataset and its potential uses. Strategic Management Journal, 42(12), 2195-2217. https://doi.org/10.1002/smj.3286
- Felten, Raj, and Seamans AIOE data repository. https://github.com/AIOE-Data/AIOE
- Frey, C. B., & Osborne, M. A. (2017). The future of employment: How susceptible are jobs to computerisation? Technological Forecasting and Social Change, 114, 254-280. https://doi.org/10.1016/j.techfore.2016.08.019
- Kochhar, R. (2023). Which U.S. workers are more exposed to AI on their jobs? Pew Research Center. https://www.pewresearch.org/social-trends/2023/07/26/which-u-s-workers-are-more-exposed-to-ai-on-their-jobs/
- Massenkoff, M., & McCrory, P. (2026). Labor market impacts of AI: A new measure and early evidence. Anthropic. https://www.anthropic.com/research/labor-market-impacts
- National Center for ONET Development. (2014). A Multi-Phase Rational Method for Developing Area Work Activities*. https://www.onetcenter.org/dl_files/DWA_2014.pdf
- National Center for ONET Development. ONET 30.2 Database. https://www.onetcenter.org/database.html
- National Center for ONET Development. ONET License Agreements. https://www.onetcenter.org/license_agreements.html
- OpenAI. (2025). GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. https://openai.com/index/gdpval/ and https://cdn.openai.com/pdf/d5eb7428-c4e9-4a33-bd86-86dd4bcf12ce/GDPval.pdf
- OpenAI GDPval public dataset. https://huggingface.co/datasets/openai/gdpval
- U.S. Council of Economic Advisers. (2024). Potential Labor Market Impacts of Artificial Intelligence: An Empirical Analysis. https://bidenwhitehouse.archives.gov/wp-content/uploads/2024/07/Potential-Labor-Market-Impacts-of-Artificial-Intelligence-An-Empirical-Analysis-July-2024.pdf
- Webb, M. (2019). The Impact of Artificial Intelligence on the Labor Market. SSRN. https://doi.org/10.2139/ssrn.3482150
- METR. (2025, updated 2026). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
- National Center for O*NET Development. O*NET Database, releases 30.2 and 30.3. U.S. Department of Labor. https://www.onetcenter.org/database.html
Capability Benchmark Sources
- Artificial Analysis. Model leaderboards. https://artificialanalysis.ai/
- Artificial Analysis. Intelligence Benchmarking Methodology. https://artificialanalysis.ai/methodology/intelligence-benchmarking
- Artificial Analysis. Speech to Text Leaderboard. https://artificialanalysis.ai/speech-to-text
- Artificial Analysis. Image Model Leaderboard. https://artificialanalysis.ai/text-to-image
- Arena. Text Leaderboard. https://arena.ai/leaderboard
- Center for Research on Foundation Models. HELM Capabilities. https://crfm.stanford.edu/helm/capabilities/latest/
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770. https://arxiv.org/abs/2310.06770
- METR. (2026). Time Horizon 1.1. https://metr.org/blog/2026-1-29-time-horizon-1-1/
- Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., & Bowman, S. R. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv:2311.12022. https://arxiv.org/abs/2311.12022
- SWE-bench leaderboards. https://www.swebench.com/
- Yue, X., Zheng, T., Ni, Y., Wang, Y., Zhang, K., Tong, S., Sun, Y., Yu, B., Zhang, G., Sun, H., Su, Y., Chen, W., & Neubig, G. (2024). MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark. arXiv:2409.02813. https://arxiv.org/abs/2409.02813