Methodology
Talvio's score is a work-activity-exposure signal. It helps organizations decide where AI training and workflow discovery should start by reading the O*NET activities that make up each occupation, then weighting those activities by their importance and level in the role.
Talvio reads what each occupation actually does from the U.S. Department of Labor's O*NET database, asks which parts of that work today's AI can genuinely help a trained person with, discounts jobs whose work mostly requires hands and physical presence, and expresses the result as a Training Priority score from 0.5 to 10 — no job scores zero, because every job includes some AI-assistable duties. The ranking is checked against independent references: human expert ratings of AI exposure, real-world AI usage data, and a structural benchmark; it also survives a thirteen-variant stress test of every discretionary methodology choice. Use the score to decide where AI training and workflow discovery should start. It is not an individual performance measure, it is not a job-loss or staffing forecast, and small differences between occupations are not meaningful — the scale, the known blind spots, and every constant in the calculation are published on this page.
TAP is deterministic: the same reviewed inputs always produce the same score, and the scoring run is arithmetic. There is no model inference at scoring time. The score combines what an occupation does from O*NET with reviewed AI capability maturity values from public benchmarks.
TAP is a calibrated index of AI-assistable weighted work, not a rank and not an absolute capability claim.
Occupations with the same TAP have similar proportions of assistable weighted work, not necessarily similar job content.
Product score explanations are work-activity-based, so top O*NET work-activity drivers
are the honest explanation for a role's score. Methodology version: TPS v2.2 (2026-07-07); constants in
methodology/tps_v2_2_config.json.
The reviewed capability scores help shape training design and capability views. They answer "what AI skills are relevant here?" while the priority ranking remains work-activity-exposure-based in the current validated release.
| Capability | Reviewed maturity | Primary benchmark |
|---|---|---|
| Written content generation and editing | 7.5/10 | HELM Capabilities |
| Information synthesis and research | 6.0/10 | Artificial Analysis Intelligence Index |
| Structured data analysis and quantitative reasoning | 5.5/10 | Artificial Analysis Intelligence Benchmarking |
| Coding and software engineering | 6.5/10 | SWE-bench |
| Conversational support and customer interaction | 6.5/10 | Arena Text Leaderboard |
| Translation and cross-language work | 7.0/10 | Artificial Analysis Multilingual Index |
| Speech and audio processing | 6.5/10 | Artificial Analysis Speech to Text Leaderboard |
| Image and document understanding | 7.5/10 | MMMU-Pro |
| Image, video, and design generation | 7.0/10 | Artificial Analysis Image Model Leaderboard |
| Planning, scheduling, and structured decision support | 5.0/10 | GDPval |
| Tool use and autonomous agents | 4.0/10 | METR Time Horizons |
| Domain-specialist reasoning | 6.5/10 | GPQA |
The score measures how much of an occupation's duties would benefit from AI assistance today. The ranking is validated against human-rated exposure (Eloundou et al. beta, Spearman rho=0.858, n=894 occupations) and corroborated by Anthropic Economic Index observed AI task-use (Spearman rho=0.458, n=479 occupations, release_2026_03_24). The AEI correlation is moderate and computed on the subset with observed-use coverage; notably, AEI corroborates the score at least as strongly as it corroborates the human benchmark itself (rho=0.458 vs 0.363 on the same set). The claim is correspondence to observed use, not high prediction of AI use.
AEI is independent of O*NET structure: it is observed AI task-use rather than a structure-derived exposure score, and it corroborates the score at least as strongly as it corroborates the human-rated benchmark itself (rho=0.458 vs 0.363 on the same occupations), addressing the structural-circularity limitation.
GDPval near-zero correlation is explained by construct differences: the score shows strong agreement with human-rated exposure (rho=0.858, n=894), moderate agreement with observed use (rho=0.458, n=479), and near-orthogonality to peak-deliverable-quality benchmarks - the expected pattern for a work-activity augmentation measure. For Talvio's purpose, that orthogonality is correct behavior: GDPval measures peak deliverable quality on selected tasks, while TAP measures exposure of an occupation's work-activity mix.
The validated claim is deliberately scoped: Talvio corresponds to observed AI use at least as well as the human-rated benchmark does on the matched AEI set (Talvio-vs-AEI rho=0.458; benchmark-vs-AEI rho=0.363). It does not claim high prediction of observed use for every occupation or every organization. Robustness is documented rather than asserted: a 13-variant sensitivity sweep shows no discretionary methodology choice materially moves the ranking. Two disclosed blind spots: communication-heavy roles (phone/email-centric) are likely under-rated, and judgment-heavy roles may be over-rated relative to 2023 expert references.
For readers who want benchmark comparisons, validation narrative, limitations, references, and audit detail.
The production model carries 122 explicit exclusions across residual, split-code, and out-of-scope military rows. They remain visible as excluded rows with provenance rather than being silently scored.
Reason categories are A residual (residual_soc_no_descriptor_data), B split-code
(split_code_no_work_activities), and C military (out_of_scope_military).
Donor imputation was analyzed and is not included in the production method.