Talvio TAP External Validation Paper

Methodology version TPS v2.2 · 2026-07-08. How Talvio scores AI augmentation potential for training prioritization, where the inputs come from, how the ranking is validated against independent references, and exactly where its limits are. Every number in this paper is reproducible from published data and published constants.

Section 1

Key Findings / Plain Summary

Talvio scores every U.S. occupation by how much of its duties would benefit from AI assistance today, and rolls those scores up to the departments in an uploaded roster. The score, Talvio Augmentation Potential (TAP, methodology version TPS v2.2), is a training-prioritization signal: it tells a learning and development team where structured AI enablement is most worth examining first, which work activities explain that priority, and which AI capabilities should shape the curriculum.

The headline validation results: the ranking agrees with independent human expert ratings of AI exposure at Spearman ρ = 0.858 across all 894 scored occupations: the full population, not a favorable sample. It is corroborated by observed real-world AI usage (Anthropic Economic Index) at ρ = 0.458 on 479 occupations; notably, that usage data agrees with Talvio's ranking at least as strongly as it agrees with the expert benchmark itself (0.363 on the same set). Structural concordance with the widely cited Felten AIOE index is 0.972, reported as methodological consistency rather than independent proof, because AIOE is built from the same O*NET foundation.

Three properties matter as much as the correlations. The score is deterministic: published constants, no model inference at scoring time, every step reproducible from public data. It is robust: a thirteen-variant sensitivity sweep shows no discretionary methodology choice materially moves the ranking. And its limits are disclosed: two known blind spots are documented in Section 7 with named examples, not buried.

Scores display on a calibrated 0.5–10 scale. No occupation scores zero, because every job includes some AI-assistable duties. The score is not an individual performance measure, not a job-loss forecast, and small differences between occupations are not meaningful.

Section 2

What the Score Measures (and What It Does Not)

The construct is augmentation: how reliably today's deployed AI can help a trained person perform the kinds of work an occupation involves. That construct was pinned in writing before scoring. It is deliberately not automation risk (whether AI replaces the job), not adoption (whether the organization has rolled AI out), and not skill gap (whether particular employees need training more than others).

Nearby constructs are easy to confuse, and the confusion is where most AI-workforce numbers go wrong. Task exposure, observed AI usage, and model performance on benchmark deliverables are three different things; Talvio validates against the first two and treats the third (peak-deliverable-quality benchmarks) as a different construct rather than a validation target. Talvio sits in the task-based exposure lineage of Felten, Raj, and Seamans and of Eloundou, Manning, Mishkin, and Rock; it is scoped to training prioritization rather than staffing-impact prediction.

What follows from the construct: a high score means employees in that role would likely get substantial leverage from learning to use AI tools well. It does not mean the role is at risk, and it does not mean AI can do the role's job. A low score means the role's work is mostly physical or in-person; it does not mean zero opportunity, which is why the scale floors at 0.5 rather than 0.

Section 3

How the Score Is Built

The pipeline is six deterministic steps on public data. For each of the 41 O*NET Generalized Work Activities an occupation performs, Talvio asks which of twelve AI capabilities is best suited to assist it and how mature that capability is today; activities are weighted by O*NET importance × level. The weighted sum is the cognitive base. The base is then discounted by the squared share of the occupation's O*NET ability importance in the psychomotor and physical branches (AI assists thinking and communicating, not lifting and touching), and the result is mapped to the display scale by a fixed, rank-preserving calibration with frozen anchors. All constants are published in methodology/tps_v2_2_config.json.

Figure 1. The TPS v2.2 pipeline Six deterministic steps from O*NET data to the calibrated score, with Registered Nurses as a running example. Pipeline Running example: Registered Nurses 1 Read the occupation from O*NET 41 Generalized Work Activities per occupation, each rated for importance and level (O*NET 30.2). Weights are importance × level, normalized to sum to 1. 41 activities 2 Score each activity by its best-fitting AI capability Addressability = the highest of the twelve capability fits (matrix v2 × maturity v2, both panel-derived). Unrelated capabilities cannot stack. addressability 0–1 per activity 3 Sum to the cognitive base Weighted sum of activity addressability, scaled 0–100. This is how much of the job's important, high-level work AI can assist. base 35.3 / 100 4 Apply the physical gate Multiply by (1 − G)², where G is the occupation's psychomotor + physical share of O*NET ability importance (30.3). gate 29% → × 0.504 5 Calibrate to the display scale Fixed, rank-preserving map with frozen anchors; clipped to [0.5, 10]. No occupation scores zero. raw 1.78 → score 3.1 6 Roll up to departments Occupation scores attach to roster titles through human match review, then aggregate headcount-weighted. department averages
Figure 1. The TPS v2.2 pipeline with Registered Nurses as a running example. Every step is arithmetic on published O*NET data; no model inference occurs at scoring time.

A worked example, Registered Nurses (29-1141.00): the activity-weighted cognitive base is 35.3 of 100; nursing is information-rich (documenting, monitoring, coordinating). But 29% of the occupation's ability importance is psychomotor/physical, so the base is multiplied by (1 − 0.29)² = 0.504, giving a raw score of 1.78, which calibrates to a displayed 3.1. The same walk-through, with live formulas, ships in the score-calculator workbook.

Two formula choices deserve their rationale stated plainly. The best-fit rule (an activity is scored by its single best capability, not a combination of all twelve) prevents unrelated capabilities from inflating cognitive activities, a failure mode in which every cognitive activity saturates near the top of the scale. The squared gate encodes that heavily physical work disproportionately limits AI assistance: the bodily core of such jobs also makes the surrounding cognitive work embodied and context-bound.

Section 4

Where the Inputs Come From

The occupational data is O*NET, the U.S. Department of Labor's occupational database: Work Activities with importance and level ratings (30.2) and the Abilities taxonomy (30.3) for the physical gate. Both are public, versioned, and independently maintained. Talvio does not author any claim about what a job involves.

The two judgment layers (how well each capability fits each activity, and how mature each capability is today) carry panel provenance rather than single-author opinion. The activity-to-capability matrix (v2) is the median of five blinded raters with distinct professional framings who saw only official O*NET definitions and the pinned construct, never the existing values or the purpose; inter-rater agreement was r = 0.94, and the panel's matrix reproduces the prior human-reviewed ranking at 0.9987. The twelve capability maturity scores (v2) come from a second five-rater panel required to ground every score in named, dated 2025–2026 deployment evidence, with “insufficient evidence” as an allowed answer. The flagship correction from that evidence pass: coding assistance, rated highest by every prior assessment, was cut from 8.4 to 6.5 after all five raters independently surfaced a randomized controlled trial (METR, July 2025) finding experienced developers were ~19% slower with AI assistance on mature codebases while believing they were faster.

Capability contexts and primary benchmark anchors

Each capability context carries a primary public benchmark anchor. The anchors are reference points for tracking capability progress between refreshes; the maturity scores themselves come from the deployed-evidence panel, because benchmark performance and real-world assistance value can diverge (coding is the documented example).

IDCapability contextMaturity (v2, 0–10)Primary benchmark anchor
cap_01Written content generation and editing7.5HELM Capabilities
cap_02Information synthesis and research6.0Artificial Analysis Intelligence Index
cap_03Structured data analysis and quantitative reasoning5.5Artificial Analysis Intelligence Benchmarking
cap_04Coding and software engineering6.5SWE-bench
cap_05Conversational support and customer interaction6.5Arena Text Leaderboard
cap_06Translation and cross-language work7.0Artificial Analysis Multilingual Index
cap_07Speech and audio processing6.5Artificial Analysis Speech to Text Leaderboard
cap_08Image and document understanding7.5MMMU-Pro
cap_09Image, video, and design generation7.0Artificial Analysis Image Model Leaderboard
cap_10Planning, scheduling, and structured decision support5.0GDPval
cap_11Tool use and autonomous agents4.0METR Time Horizons
cap_12Domain-specialist reasoning6.5GPQA

Because AI capability is a moving target, maturity scores are re-assessed on a fixed cadence (each major frontier-model release, or quarterly, whichever comes first) using the same evidence-verified panel process. Each refresh re-runs the ablation and regression suite; material changes bump the methodology version. A score without a version stamp should not be trusted, including Talvio's.

Section 5

External Validation

The validation hierarchy is ordered by independence, not by flattery. The primary referent is human-rated exposure: Eloundou et al.'s beta measure, the share of an occupation's tasks where AI assistance cuts completion time by half or more, the closest published construct to Talvio's. Agreement is Spearman ρ = 0.858 on all 894 scored occupations (bootstrap 95% CI [0.844, 0.875] on the v2.0 baseline).

Figure 2. TPS v2.2 vs human-rated exposure, all 894 occupations Scatter of the calibrated score against human expert exposure ratings; Spearman rho 0.858; labeled outliers show the two disclosed blind spots. 0 0 2 2 4 4 6 6 8 8 10 10 Sales Reps Admin Law Judges Registered Nurses Software Developers Dishwashers Human-rated AI exposure (Eloundou et al. beta, scaled 0–10) TPS v2.2 (0.5–10) Spearman ρ = 0.858, n = 894
Figure 2. TPS v2.2 against human-rated exposure for all 894 scored occupations. Labeled points illustrate the two disclosed blind spots (Section 7): Sales Representatives sit far below the diagonal (under-rated), Administrative Law Judges far above it (over-rated).

The second referent is behavioral: the Anthropic Economic Index's observed AI task-use (release 2026-03-24, US). AEI task records were matched to O*NET task statements (1,303 tasks matched) and rolled up to 479 occupations as a log-scaled usage intensity. Agreement is ρ = 0.458; on the same occupations, AEI agrees with the human-rated benchmark at only 0.363. Real-world usage corroborates Talvio's ranking at least as strongly as it corroborates expert judgment. The claim is deliberately modest: correspondence with observed use, not high prediction of it.

The third referent is structural: Felten AIOE, ρ = 0.972 on 781 SOC6 occupations. This is reported last and always with its caveat: AIOE is itself O*NET-structure-derived, so high agreement demonstrates that Talvio's use of the O*NET spine is methodologically consistent with the academic standard, not that an independent measurement agrees.

Figure 3. Calibrated score distribution Sorted TPS v2.2 scores for the 894 direct-scored occupations, spanning 0.5 to 10. 0 2 4 6 8 10 894 occupations, sorted by score floor 0.5 — no job scores zero
Figure 3. The calibrated score distribution spans the full display scale: median 4.0, deciles 1.3 / 8.6, floor 0.5. Median 4.0, deciles 1.3 and 8.6, floor 0.5. The calibration is fixed and rank-preserving, so the spread reflects real differences in AI-assistable work, not display tuning.

Section 6

Robustness: The Ranking Does Not Depend on Our Choices

Any scoring methodology accumulates discretionary choices, and a fair critic asks whether the answers depend on them. Talvio tests this directly. Each choice was flipped one at a time, and then all of them at once, and the resulting ranking compared to production:

Variant (one choice flipped)Median rank shift (of 894)Rank corr vs v2.2ρ vs human benchmarkVerdict
weights: IM only7.00.99870.857PASS
weights: IM x LV^27.00.99890.857PASS
gate form: linear k=19.00.99810.858PASS
gate form: k=34.00.99960.857PASS
gate set: + sensory21.00.99030.856FAIL
gate set: physical only48.00.95980.815FAIL
gate weight: raw IM10.00.99730.855PASS
addressability: top-2 softOR3.00.99980.857PASS
addressability: top-3 mean3.00.99980.858PASS
matrix: alternative reviewed values7.00.99890.861PASS
capabilities: alternative reviewed set2.00.99990.858PASS
capabilities: blind panel4.00.99970.86PASS
WORST CASE combo14.00.99550.862PASS

Eleven of thirteen variants move the median occupation by single-digit rank positions out of 894; even the worst-case combination of every alternative choice shifts the median occupation by 14 positions. The one genuinely load-bearing choice is which abilities count as “needs a body.” That choice is justified rather than tuned: the psychomotor + physical definition exactly reproduces the benchmark-validated gates, and both alternatives measurably degrade agreement with the human benchmark (dropping psychomotor: ρ falls from 0.858 to 0.815).

A second robustness result concerns the capability maturity scores. Replacing them with a uniform constant changes the ranking almost not at all (ρ = 0.9948), a result stable across three independently constructed maturity sets. The ranking is driven by work-activity structure, importance/level weighting, and the physical gate; capability maturity is training-design context: it describes what to teach, not who ranks first. The product's capability_layer_explains_score flag is false and is re-tested at every capability refresh.

Section 7

Known Blind Spots and Limitations

Two systematic errors survive every version of an activity-structure method, and Talvio discloses them with examples rather than patching them cosmetically.

Communication-heavy roles are likely under-rated. Jobs done largely by phone and email (sales representatives, scored 4.7 against a human-rated 7.1; executive assistants; customer service) score below what expert references suggest, because activity labels cannot see that this work flows through channels where AI assistance is strongest. Phone/email intensity measurably predicts this under-rating, so affected occupations are annotated in the product rather than silently mis-ranked.

Judgment-heavy roles may be over-rated. Administrative law judges (8.6 vs 3.5) and counselors score above 2023 expert ratings: the structure sees information work and cannot see that high-stakes human judgment resists delegation. The 2026 usage data complicates this in an interesting way: counseling-type work shows among the heaviest real AI use in AEI, so the truth likely sits between Talvio's score and the older expert view.

Beyond the blind spots: heavily physical jobs floor near 0.5–1 rather than zero by design. The primary benchmark is expert judgment from 2023, not a measured outcome. No one in this field, Talvio included, yet has longitudinal evidence that training the highest-scored roles first produces the best results; match-review accuracy tracking is the first step toward outcome data. National occupation profiles are not any single organization: titles and real workflows differ, which is why every analysis passes through human match review with one-click overrides before results are relied on. And ordinal survey inputs cannot support decimal precision: scores carry an uncertainty band, and one-decimal display is a readability choice, not a precision claim.

Section 8

Coverage, Governance, and Versioning

Talvio directly scores 894 occupations and explicitly excludes 122 rows: 77 residual SOC codes without descriptor data, 26 split codes without O*NET Work Activities, and 19 military occupations out of civilian scope by design. Excluded rows remain visible with provenance rather than being silently scored or imputed.

Every input carries a version and a date: O*NET releases, matrix v2 and capabilities v2 (both panel-adopted 2026-07-07), the calibration anchors (frozen; re-fit only at version bumps), and the score itself (TPS v2.2). The methodology is versioned and updated when the evidence warrants; every revision is documented in the methodology changelog. The full revision record (pre-registration with amendment log, both panel studies with raw ratings, the sensitivity sweep, and a formula-transparent calculator workbook) is retained in the methodology archive.

Product copy is governed by a scoped-claims contract: every validation statement must cite the human-rated benchmark and the observed-use figure together; the structural concordance figure may not appear without its circularity note; and copy implying job-loss prediction or capability-driven ranking is nonconforming by rule.

Section 9

References

External Literature and Data Sources

Capability Benchmark Sources