Talvio External Validation Paper

Methodology version TPS v2.2 (July 2026) · this edition September 2026. How Talvio scores how much of each job's work AI can help with, where the inputs come from, how the ranking compares with independent research, how US industries compare, and where its limits are. Every figure comes from public data. The three outside comparisons are rebuilt from saved copies of that data whenever the scores change.

Section 1

Key Findings / Plain Summary

Talvio scores 894 U.S. occupations, held by 96% of US workers, by how much of their work today's AI can help with, and adds those scores up by department for an uploaded roster. The product calls it the Training Priority score; formulas call it TAP (methodology version TPS v2.2). It helps a learning and development team decide which roles to start AI training with, which work activities explain that order, and which AI skills the training should cover.

The headline validation results: the ranking agrees with independent human expert ratings of AI exposure at Spearman ρ = 0.858 across all 894 scored occupations: the full population, not a favorable sample. It is corroborated by observed real-world AI usage (Anthropic Economic Index) at ρ = 0.458 on 479 occupations; notably, that usage data agrees with Talvio's ranking at least as strongly as it agrees with the expert benchmark itself (0.363 on the same set). Structural concordance with the widely cited Felten AIOE index is 0.972, reported as methodological consistency rather than independent proof, because AIOE is built from the same O*NET foundation.

Three properties matter as much as the correlations. The score is deterministic: published constants, no model inference at scoring time, every step arithmetic on public data. It is robust: in a thirteen-variant sensitivity sweep, eleven alternative choices barely move the ranking, and the two that do both change which abilities count as physical and both lower agreement with the expert ratings. And its limits are disclosed: two known blind spots are documented in Section 7 with named examples, not buried.

Scores display on a calibrated 0.5–10 scale. No occupation scores zero, because every job includes some AI-assistable duties. The score is not an individual performance measure, not a job-loss forecast, and small differences between occupations are not meaningful.

Weighted by how many people hold each occupation (BLS, May 2025), the average US worker scores 4.2. Industry averages built this way rank the 20 industry sectors in nearly the same order as the same averages built from the expert ratings (0.94) and in a similar order to a measure built from Copilot use (0.77). Section 5 gives the detail and its limits.

Section 2

What the Score Measures (and What It Does Not)

The construct is augmentation: how reliably today's deployed AI can help a trained person perform the kinds of work an occupation involves. That construct was pinned in writing before scoring. It is deliberately not automation risk (whether AI replaces the job), not adoption (whether the organization has rolled AI out), and not skill gap (whether particular employees need training more than others).

Nearby constructs are easy to confuse, and the confusion is where most AI-workforce numbers go wrong. Task exposure, observed AI usage, and model performance on benchmark deliverables are three different things; Talvio validates against the first two and treats the third (peak-deliverable-quality benchmarks) as a different construct rather than a validation target. Talvio sits in the task-based exposure lineage of Felten, Raj, and Seamans and of Eloundou, Manning, Mishkin, and Rock; it is scoped to training prioritization rather than staffing-impact prediction.

What follows from the construct: a high score means much of the role's work is the kind AI tools can help with, so learning to use them well matters most there. It does not mean the role is at risk, and it does not mean AI can do the role's job. A low score means the role's work is mostly physical or in-person; it does not mean zero opportunity, which is why the scale floors at 0.5 rather than 0.

Section 3

How the Score Is Built

The pipeline is six deterministic steps on public data. For each of the 41 O*NET Generalized Work Activities an occupation performs, Talvio asks which of twelve AI capabilities is best suited to assist it and how mature that capability is today; activities are weighted by O*NET importance × level. The weighted sum is the cognitive base. The base is then discounted by the squared share of the occupation's O*NET ability importance in the psychomotor and physical branches (AI assists thinking and communicating, not lifting and touching), and the result is mapped to the display scale by a fixed, rank-preserving calibration with frozen anchors.

Figure 1. The TPS v2.2 pipeline Six deterministic steps from O*NET data to the calibrated score, with Registered Nurses as a running example. Pipeline Running example: Registered Nurses 1 Read the occupation from O*NET 41 Generalized Work Activities per occupation, each rated for importance and level (O*NET 30.2). Weights are importance × level, normalized to sum to 1. 41 activities 2 Score each activity by its best-fitting AI capability Addressability = the highest of the twelve capability fits (matrix v2 × maturity v2, both panel-derived). Unrelated capabilities cannot stack. addressability 0–1 per activity 3 Sum to the cognitive base Weighted sum of activity addressability, scaled 0–100. This is how much of the job's important, high-level work AI can assist. base 35.3 / 100 4 Apply the physical gate Multiply by (1 − G)², where G is the occupation's psychomotor + physical share of O*NET ability importance (30.3). gate 29% → × 0.504 5 Calibrate to the display scale Fixed, rank-preserving map with frozen anchors; clipped to [0.5, 10]. No occupation scores zero. raw 1.78 → score 3.1 6 Roll up to departments Occupation scores attach to roster titles through human match review, then aggregate headcount-weighted. department averages
Figure 1. The TPS v2.2 pipeline with Registered Nurses as a running example. Every step is arithmetic on published O*NET data; no model inference occurs at scoring time.

A worked example, Registered Nurses (29-1141.00): the activity-weighted cognitive base is 35.3 of 100; nursing is information-rich (documenting, monitoring, coordinating). But 29% of the occupation's ability importance is psychomotor/physical, so the base is multiplied by (1 − 0.29)² = 0.504, giving a raw score of 1.78, which calibrates to a displayed 3.1.

Two formula choices deserve their rationale stated plainly. The best-fit rule (an activity is scored by its single best capability, not a combination of all twelve) prevents unrelated capabilities from inflating cognitive activities, a failure mode in which every cognitive activity saturates near the top of the scale. The squared gate encodes that heavily physical work disproportionately limits AI assistance: the bodily core of such jobs also makes the surrounding cognitive work embodied and context-bound.

Section 4

Where the Inputs Come From

The occupational data is O*NET, the U.S. Department of Labor's occupational database: Work Activities with importance and level ratings (30.2) and the Abilities taxonomy (30.3) for the physical gate. Both are public, versioned, and independently maintained. Talvio does not author any claim about what a job involves.

The two judgment layers (how well each capability fits each activity, and how mature each capability is today) carry panel provenance rather than single-author opinion. The activity-to-capability matrix (v2) is the median of five blinded raters with distinct professional framings who saw only official O*NET definitions and the pinned construct, never the existing values or the purpose; inter-rater agreement was r = 0.94, and the panel's matrix reproduces the prior human-reviewed ranking at 0.9986. The twelve capability maturity scores (v2) come from a second five-rater panel required to ground every score in named, dated 2025–2026 deployment evidence, with “insufficient evidence” as an allowed answer. The largest correction from that evidence pass: coding assistance, which the blind panel had rated 8.5, its highest rating, fell to 6.5 after all five raters independently surfaced a randomized controlled trial (METR, July 2025) finding experienced developers were ~19% slower with AI assistance on mature codebases while believing they were faster.

Capability contexts and primary benchmark anchors

Each capability context carries a primary public benchmark anchor. The anchors are reference points for tracking capability progress between refreshes; the maturity scores themselves come from the deployed-evidence panel, because benchmark performance and real-world assistance value can diverge (coding is the documented example).

IDCapability contextMaturity (v2, 0–10)Primary benchmark anchor
cap_01Written content generation and editing7.5HELM Capabilities
cap_02Information synthesis and research6.0Artificial Analysis Intelligence Index
cap_03Structured data analysis and quantitative reasoning5.5Artificial Analysis Intelligence Benchmarking
cap_04Coding and software engineering6.5SWE-bench
cap_05Conversational support and customer interaction6.5Arena Text Leaderboard
cap_06Translation and cross-language work7.0Artificial Analysis Multilingual Index
cap_07Speech and audio processing6.5Artificial Analysis Speech to Text Leaderboard
cap_08Image and document understanding7.5MMMU-Pro
cap_09Image, video, and design generation7.0Artificial Analysis Image Model Leaderboard
cap_10Planning, scheduling, and structured decision support5.0GDPval
cap_11Tool use and autonomous agents4.0METR Time Horizons
cap_12Domain-specialist reasoning6.5GPQA

Because AI capability is a moving target, maturity scores are re-assessed on a fixed cadence (each major frontier-model release, or quarterly, whichever comes first) using the same evidence-verified panel process. Each refresh re-runs the ablation and regression suite; material changes bump the methodology version. A score without a version stamp should not be trusted, including Talvio's.

Section 5

External Validation

The validation hierarchy is ordered by independence, not by flattery. The primary referent is human-rated exposure: Eloundou et al.'s beta measure, the share of an occupation's tasks where AI assistance cuts completion time by half or more, the closest published construct to Talvio's. Agreement is Spearman ρ = 0.858 on all 894 scored occupations (bootstrap 95% CI [0.839, 0.873]).

Figure 2. TPS v2.2 vs human-rated exposure, all 894 occupations Scatter of the calibrated score against human expert exposure ratings; Spearman rho 0.858; labeled outliers show the two disclosed blind spots. 0 0 2 2 4 4 6 6 8 8 10 10 Sales Reps Admin Law Judges Registered Nurses Software Developers Dishwashers Human-rated AI exposure (Eloundou et al. beta, scaled 0–10) TPS v2.2 (0.5–10) Spearman ρ = 0.858, n = 894
Figure 2. TPS v2.2 against human-rated exposure for all 894 scored occupations. Labeled points illustrate the two disclosed blind spots (Section 7): Sales Representatives sit far below the diagonal (under-rated), Administrative Law Judges far above it (over-rated).

The second referent is behavioral: the Anthropic Economic Index's observed AI task-use (release 2026-03-24, US). AEI task records were matched to O*NET task statements (1,303 tasks matched) and rolled up to 479 occupations as a log-scaled usage intensity. Agreement is ρ = 0.458; on the same occupations, AEI agrees with the human-rated benchmark at only 0.363. Real-world usage corroborates Talvio's ranking at least as strongly as it corroborates expert judgment. The claim is deliberately modest: correspondence with observed use, not high prediction of it. Other reasonable ways of joining the usage data (global rather than US records, SOC codes rather than O*NET codes, average rather than summed task shares) give 0.34 to 0.58. In every one, usage agrees more closely with Talvio's ranking than with the expert ratings.

The third referent is structural: Felten AIOE, ρ = 0.972 on 781 occupations, each taking the index value of its 6-digit SOC code (682 distinct codes; one value per code gives 0.982). This is reported last and always with its caveat: AIOE is itself O*NET-structure-derived, so high agreement demonstrates that Talvio's use of the O*NET spine is methodologically consistent with the academic standard, not that an independent measurement agrees.

Figure 3. US workers by score Bar chart of US workers by Talvio score in half-point bands, weighted by BLS worker counts (May 2025). Average 4.2. 0M 5M 10M 15M 20M Scores 0.5 to 0.9: 7.4 million workers Scores 1.0 to 1.4: 9.8 million workers Scores 1.5 to 1.9: 18.7 million workers Scores 2.0 to 2.4: 14.6 million workers Scores 2.5 to 2.9: 14.3 million workers Scores 3.0 to 3.4: 10.8 million workers Scores 3.5 to 3.9: 8.8 million workers Scores 4.0 to 4.4: 4.3 million workers Scores 4.5 to 4.9: 5.1 million workers Scores 5.0 to 5.4: 2.3 million workers Scores 5.5 to 5.9: 4.7 million workers Scores 6.0 to 6.4: 2.5 million workers Scores 6.5 to 6.9: 18.2 million workers Scores 7.0 to 7.4: 4.0 million workers Scores 7.5 to 7.9: 9.3 million workers Scores 8.0 to 8.4: 6.4 million workers Scores 8.5 to 8.9: 6.0 million workers Scores 9.0 to 9.4: 2.5 million workers Scores 9.5 to 9.9: 0.2 million workers Scores 10.0 to 10.4: 0.0 million workers 1 2 3 4 5 6 7 8 9 10 Average 4.2 Training Priority score (half-point bands); bar height is US workers
Figure 3. US workers by score, weighted by BLS worker counts (May 2025). The average US worker scores 4.2. 50% of workers are in jobs scoring under 3.5, mostly hands-on work, and 31% score 6.5 or higher, mostly office and professional work. Counted once per occupation instead, the median of the 894 scored occupations is 4.0, with deciles 1.3 and 8.6 and a floor of 0.5.

Industry averages

Occupation scores were weighted by BLS worker counts (Occupational Employment and Wage Statistics, May 2025). For every industry, BLS surveys employers and publishes how many people work in each occupation; Talvio uses those published counts and does not estimate which jobs an industry has. An industry's average is each occupation's score weighted by how many people in that industry hold it. 96.3% of US workers hold a scored occupation; the rest are left out, not imputed. Where O*NET splits a BLS occupation into specialties, the general occupation's score is used, so registered nurses count at 3.1, not at the average of the nurse specialties. Where BLS merges several O*NET occupations, the BLS occupation takes their simple average.

The average US worker scores 4.25. The 20 industry sectors run from 1.5 (Agriculture, Forestry, Fishing and Hunting) to 7.4 (Finance and Insurance).

The same worker counts were applied to three published measures. Rank agreement between the industry averages:

Industries comparedExpert ratings (Eloundou et al.)Felten index, language-model versionCopilot use (Microsoft)
Sectors (20)0.940.950.77
Industries, 3-digit NAICS (85)0.920.940.78
Industries, 4-digit NAICS (246)0.930.950.78

For single occupations, agreement with Microsoft's AI applicability measure is 0.71 (751 occupations). Microsoft's measure comes from how people use Copilot, which trails what AI could do in some industries.

Four limits apply. Averages agree more closely than single occupations do, because each covers thousands of workers, and most of the agreement reflects how much of each industry's work is office and desk work. The Felten figure is partly circular, as above. Industry averages support ordering claims only, never claims that one industry is some multiple of another. And the BLS survey does not cover the self-employed or farms, so the agriculture sector reflects forestry, logging and farm support services, not farmworkers.

Section 6

Robustness: The Ranking Does Not Depend on Our Choices

Any scoring methodology accumulates discretionary choices, and a fair critic asks whether the answers depend on them. Talvio tests this directly. Each choice was flipped one at a time, and then all of them at once, and the resulting ranking compared to production:

Variant (one choice flipped)Median rank shift (of 894)Rank corr vs v2.2ρ vs human benchmarkVerdict
weights: IM only7.00.99870.857PASS
weights: IM x LV^27.00.99890.857PASS
gate form: linear k=19.00.99810.858PASS
gate form: k=34.00.99960.857PASS
gate set: + sensory21.00.99030.856FAIL
gate set: physical only48.00.95980.815FAIL
gate weight: raw IM10.00.99730.855PASS
addressability: top-2 softOR3.00.99980.857PASS
addressability: top-3 mean3.00.99980.858PASS
matrix: alternative reviewed values7.00.99890.861PASS
capabilities: alternative reviewed set2.00.99990.858PASS
capabilities: blind panel4.00.99970.86PASS
WORST CASE combo14.00.99550.862PASS

The bar was set before the sweep ran: a choice passes if flipping it moves the median occupation fewer than 20 of 894 rank positions. Eleven of thirteen variants pass, and nine move the median occupation fewer than 10 positions; even the worst-case combination of every alternative choice shifts it by 14. The one genuinely load-bearing choice is which abilities count as “needs a body.” That choice is justified rather than tuned: the psychomotor + physical definition exactly reproduces the benchmark-validated gates, and both alternatives measurably degrade agreement with the human benchmark (dropping psychomotor: ρ falls from 0.858 to 0.815).

A second robustness result concerns the capability maturity scores. Giving every AI skill the same rating barely changes the ranking. Before the physical gate, rank agreement with the real ratings is 0.9948 (0.9988 and 0.9952 with the two earlier rating sets); after the gate it is 0.9997. The ranking is driven by work-activity structure, importance/level weighting, and the physical gate; capability maturity is training-design context: it describes what to teach, not who ranks first. The product therefore does not use capability maturity to explain the score, and that decision is re-tested at every capability refresh.

Section 7

Known Blind Spots and Limitations

Two systematic errors survive every version of an activity-structure method, and Talvio discloses them with examples rather than patching them cosmetically.

Communication-heavy roles are likely under-rated. Jobs done largely by phone and email (sales representatives, scored 4.7 against a human-rated 7.1; executive assistants; customer service) score below what expert references suggest, because activity labels cannot see that this work flows through channels where AI assistance is strongest. Phone and email intensity measurably predicts this under-rating, and the product's limits name it.

Judgment-heavy roles may be over-rated. Administrative law judges (8.6 vs 3.5) and counselors score above 2023 expert ratings: the structure sees information work and cannot see that high-stakes human judgment resists delegation. The 2026 usage data complicates this in an interesting way: counseling-type work shows among the heaviest real AI use in AEI, so the truth likely sits between Talvio's score and the older expert view.

Beyond the blind spots: heavily physical jobs floor near 0.5–1 rather than zero by design. The primary benchmark is expert judgment from 2023, not a measured outcome. No one in this field, Talvio included, yet has longitudinal evidence that training the highest-scored roles first produces the best results; match-review accuracy tracking is the first step toward outcome data. National occupation profiles are not any single organization: titles and real workflows differ, which is why every analysis passes through human match review with one-click overrides before results are relied on. And ordinal survey inputs cannot support decimal precision: one-decimal display is a readability choice, and small differences between occupations do not mean much.

Section 8

Coverage, Governance, and Versioning

Talvio directly scores 894 occupations and explicitly excludes 122 rows: 77 residual SOC codes without descriptor data, 26 split codes without O*NET Work Activities, and 19 military occupations out of civilian scope by design. Excluded rows remain visible with provenance rather than being silently scored or imputed.

Every input carries a version and a date: O*NET releases, matrix v2 and capabilities v2 (both panel-adopted 2026-07-07), the calibration anchors (frozen; re-fit only at version bumps), and the score itself (TPS v2.2). The methodology is versioned and updated when the evidence warrants, and every revision is recorded. The full record is kept: the pre-registration with its amendment log, both panel studies with raw ratings, the sensitivity sweep, and a calculator workbook with the formulas. The US comparison is rebuilt whenever scores change and each spring when BLS publishes new worker counts.

Product copy is governed by a scoped-claims contract: every validation statement must cite the human-rated benchmark and the observed-use figure together; the structural concordance figure may not appear without its circularity note; and copy implying job-loss prediction or capability-driven ranking is nonconforming by rule.

Section 9

References

External Literature and Data Sources

Capability Benchmark Sources