Methodology

How Talvio scores AI augmentation potential for training prioritization

Talvio's score is a work-activity-exposure signal. It helps organizations decide where AI training and workflow discovery should start by reading the O*NET activities that make up each occupation, then weighting those activities by their importance and level in the role.

What The Score Claims

Talvio reads what each occupation actually does from the U.S. Department of Labor's O*NET database, asks which parts of that work today's AI can genuinely help a trained person with, discounts jobs whose work mostly requires hands and physical presence, and expresses the result as a Training Priority score from 0.5 to 10 — no job scores zero, because every job includes some AI-assistable duties. The ranking is checked against independent references: human expert ratings of AI exposure, real-world AI usage data, and a structural benchmark; it also survives a thirteen-variant stress test of every discretionary methodology choice. Use the score to decide where AI training and workflow discovery should start. It is not an individual performance measure, it is not a job-loss or staffing forecast, and small differences between occupations are not meaningful — the scale, the known blind spots, and every constant in the calculation are published on this page.

How TAP Is Computed

TAP is deterministic: the same reviewed inputs always produce the same score, and the scoring run is arithmetic. There is no model inference at scoring time. The score combines what an occupation does from O*NET with reviewed AI capability maturity values from public benchmarks.

1. Occupation activity weights For each occupation and GWA, Talvio computes raw weight as O*NET Importance times Level, then normalizes within the occupation so all activity weights sum to 1.
wo,j = IMo,j × LVo,j j(IMo,j × LVo,j)
2. Activity-to-capability matrix The reviewed 41 x 12 matrix uses 0-1 values as assistable-work fractions for each GWA-capability pair. Matrix v2 values are medians of a five-rater blinded review panel (adopted 2026-07-07).
0 ≤ Mj,k ≤ 1
3. Capability maturity Each reviewed capability score is stored on a 0-10 scale and used in scoring as a fraction.
mk = Ck10
4. Activity addressability (best-fit capability) Each activity is scored by the single AI capability best suited to assist it. Using the best fit (rather than combining all twelve) prevents unrelated capabilities from inflating cognitive activities.
Aj = maxk(Mj,k · mk)
5. Physical gate The occupation’s cognitive base is discounted by the squared share of its O*NET ability importance (above baseline) in the psychomotor and physical branches. AI assists thinking and communicating, not lifting and touching; the square encodes that heavily physical work disproportionately limits AI assistance.
baseo = ∑jwo,jAj rawo = baseo · (1 − Go)2
6. Calibrated display score Raw scores are mapped to the display scale by a fixed, rank-preserving calibration (frozen anchors; re-fit only at version bumps). The floor is 0.5: no occupation scores zero, because every job includes some AI-assistable duties.
so = 9 · (rawop1) / (p99p1) + 0.5 TAPo = clip[0.5, 10](so)

TAP is a calibrated index of AI-assistable weighted work, not a rank and not an absolute capability claim. Occupations with the same TAP have similar proportions of assistable weighted work, not necessarily similar job content. Product score explanations are work-activity-based, so top O*NET work-activity drivers are the honest explanation for a role's score. Methodology version: TPS v2.2 (2026-07-07); constants in methodology/tps_v2_2_config.json.

TAP calculation example for Registered Nurses showing work-activity weights, best-fit capability addressability, the physical gate, and the calibrated score.
Real calculation example from O*NET 30.2: Registered Nurses. Source data is generated with matrix v2, capability maturities v2, the physical gate, and importance-times-level weighting (TPS v2.2).
Sorted distribution curve of TAP scores across all direct-scored O*NET occupations with six labeled anchor occupations.
Sorted TAP distribution for the 894 direct-scored occupations. Anchor labels are real O*NET occupations from the production master.

Capability Maturity As Training Context

The reviewed capability scores help shape training design and capability views. They answer "what AI skills are relevant here?" while the priority ranking remains work-activity-exposure-based in the current validated release.

Capability Reviewed maturity Primary benchmark
Written content generation and editing 7.5/10 HELM Capabilities
Information synthesis and research 6.0/10 Artificial Analysis Intelligence Index
Structured data analysis and quantitative reasoning 5.5/10 Artificial Analysis Intelligence Benchmarking
Coding and software engineering 6.5/10 SWE-bench
Conversational support and customer interaction 6.5/10 Arena Text Leaderboard
Translation and cross-language work 7.0/10 Artificial Analysis Multilingual Index
Speech and audio processing 6.5/10 Artificial Analysis Speech to Text Leaderboard
Image and document understanding 7.5/10 MMMU-Pro
Image, video, and design generation 7.0/10 Artificial Analysis Image Model Leaderboard
Planning, scheduling, and structured decision support 5.0/10 GDPval
Tool use and autonomous agents 4.0/10 METR Time Horizons
Domain-specialist reasoning 6.5/10 GPQA

External Validation

The score measures how much of an occupation's duties would benefit from AI assistance today. The ranking is validated against human-rated exposure (Eloundou et al. beta, Spearman rho=0.858, n=894 occupations) and corroborated by Anthropic Economic Index observed AI task-use (Spearman rho=0.458, n=479 occupations, release_2026_03_24). The AEI correlation is moderate and computed on the subset with observed-use coverage; notably, AEI corroborates the score at least as strongly as it corroborates the human benchmark itself (rho=0.458 vs 0.363 on the same set). The claim is correspondence to observed use, not high prediction of AI use.

rho=0.858
Human-rated exposure alignment Primary validation referent: Eloundou et al. beta (share of tasks where AI assistance halves time), n=894 occupations — the full scored population, bootstrap 95% CI [0.844, 0.875] on the v2.0 baseline.
rho=0.458
AEI observed-use corroboration Anthropic Economic Index observed AI task-use, n=479 occupations, release_2026_03_24. Notably, AEI agrees with the score at least as strongly as it agrees with the human benchmark itself (0.458 vs 0.363 on the same set).
rho=0.972
Structural concordance (Felten AIOE) n=781 SOC6. Reported as methodological consistency, not independent validation: AIOE is itself O*NET-structure-derived, so high agreement is expected and does not add independent evidence.

AEI is independent of O*NET structure: it is observed AI task-use rather than a structure-derived exposure score, and it corroborates the score at least as strongly as it corroborates the human-rated benchmark itself (rho=0.458 vs 0.363 on the same occupations), addressing the structural-circularity limitation.

GDPval near-zero correlation is explained by construct differences: the score shows strong agreement with human-rated exposure (rho=0.858, n=894), moderate agreement with observed use (rho=0.458, n=479), and near-orthogonality to peak-deliverable-quality benchmarks - the expected pattern for a work-activity augmentation measure. For Talvio's purpose, that orthogonality is correct behavior: GDPval measures peak deliverable quality on selected tasks, while TAP measures exposure of an occupation's work-activity mix.

The validated claim is deliberately scoped: Talvio corresponds to observed AI use at least as well as the human-rated benchmark does on the matched AEI set (Talvio-vs-AEI rho=0.458; benchmark-vs-AEI rho=0.363). It does not claim high prediction of observed use for every occupation or every organization. Robustness is documented rather than asserted: a 13-variant sensitivity sweep shows no discretionary methodology choice materially moves the ranking. Two disclosed blind spots: communication-heavy roles (phone/email-centric) are likely under-rated, and judgment-heavy roles may be over-rated relative to 2023 expert references.

Read the TAP external validation paper

For readers who want benchmark comparisons, validation narrative, limitations, references, and audit detail.

Coverage Limitation

The production model carries 122 explicit exclusions across residual, split-code, and out-of-scope military rows. They remain visible as excluded rows with provenance rather than being silently scored.

Reason categories are A residual (residual_soc_no_descriptor_data), B split-code (split_code_no_work_activities), and C military (out_of_scope_military). Donor imputation was analyzed and is not included in the production method.

894
Direct-scored occupations Rows with O*NET 30.2 Work Activities coverage.
77
A residual Excluded because residual SOC rows do not have occupation-specific O*NET Work Activities profiles.
26
B split-code Excluded because no O*NET 30.2 Work Activities row is available for this detailed split code.
19
C military Out of scope by design for Talvio's civilian occupation scope; not a coverage failure.