How We Score a Job
This page explains, in full, how a Pathcast report arrives at its numbers. Nothing here is hidden behind "proprietary". If you want to argue with a score, this is where you find out what to argue with.
We score tasks, not titles
Every other tool we have seen takes a job title and returns a percentage. We think a title is too coarse to say anything useful.
A job is a bundle of roughly fifteen distinct tasks, and they do not share a fate. A financial analyst's variance commentary and their quarterly board-deck formatting sit under the same title and face completely different pressures. Even the standard occupational taxonomy does not help much: in our own data, two roles that share the same O*NET code shared zero of their fourteen tasks. Scoring the title would have given both the same number. Scoring the tasks gave them different maps.
So for each of our fifty roles, we build a task list by hand, then confirm it against the person's own résumé and their answers. Each task gets a share-of-week weight. Everything below is scored per task and rolled up through those weights.
The four scored dimensions
Each task is scored on four things. One measures what AI can do. Three measure what protects the task anyway.
Capability (0–4). What a frontier model plus ordinary, commercially available tooling can do on this task today. Not what a demo suggests, not what a roadmap promises. A 0 means it cannot meaningfully help; a 4 means it can do the task end to end at a quality a reasonable manager would accept. This is the only dimension that moves as the technology moves, and we say more about that below.
Verification cost (0–3). How expensive it is to find out the output was wrong. A misspelled meeting invite is caught in seconds. An error in a reconciliation may surface months later in an audit. High verification cost protects a task, because someone has to keep a human close enough to catch the mistakes, and that person tends to be the one who used to do the work.
Context dependency (0–3). How much of the task depends on organisational knowledge that is not written down anywhere. Which client hates being cc'd. Why the Q3 numbers are always restated. What "the usual" means on a purchase order. A model cannot use what it cannot read.
Accountability (0–3). Whether a named human must be answerable for the result. Some tasks legally or professionally require a person's signature, licence or judgement on the record. Others do not, even if a person happens to do them now.
The formula
We publish it because you should be able to check our work.
First, the three protective dimensions combine into a single protection value:
protection = 0.75 × (0.25 × ver/3 + 0.35 × ctx/3 + 0.40 × acc/3)
Then exposure is capability, scaled to 100, discounted by that protection:
exposure = (capability / 4 × 100) × (1 − protection)
Bands:
- 65 and above: automatable now. The tools exist and the protections are thin.
- 35 to 64: augmented. A person still does the task, but with AI doing a large share of it, and the number of people needed changes.
- Below 35: durable. Either the capability isn't there yet or the protection is strong enough that it doesn't matter.
A worked example. A task with capability 3, verification cost 1, context dependency 2 and accountability 1:
protection = 0.75 × (0.25 × 0.33 + 0.35 × 0.67 + 0.40 × 0.33) = 0.75 × 0.45 = 0.34
exposure = 75 × (1 − 0.34) = 49.7 → augmented
About the 0.75. That constant caps how much protection can ever discount exposure. With it, a task the model can fully do (capability 4) but which is maximally protected on all three dimensions still scores 25, not 0. Set it to 1.0 and the same task scores 0, which would say that accountability and context make a task untouchable. Set it to 0.5 and it scores 50, which would say protection barely matters. We chose 0.75 because we think protection is real but not permanent. This single number sets how alarming the entire system is. It is a judgement, not a measurement, and that is exactly why it is printed here rather than buried. If you think it should be different, you now know precisely what you disagree with.
The weights inside the bracket (0.25, 0.35, 0.40) reflect our view that accountability protects more durably than context, which protects more durably than verification cost. Verification gets the smallest weight because it is the protection that improves fastest as tools get better at checking their own work.
Why prediction only enters through the dials
Capability is scored against what exists today. That is deliberate. Baking a forecast into the base score would mean that every disagreement about the future becomes a disagreement about the data, and nobody could tell which was which.
Instead, the person reading the report sets three dials: how far AI agents get in three years, how fast their employer adopts, and how much pressure lands on entry-level hiring. Those settings adjust the capability score and the weight of the hiring signal. The protective dimensions do not move with the dials, because unwritten context and legal accountability do not change on a model release schedule.
This separation is what lets us re-score honestly later. When capability moves, we update the capability column and say so. The dials stay the user's. Nobody has to wonder whether a changed score reflects a changed world or a changed opinion.
The second score: rung erosion
Task exposure answers "how much of this job can be done by AI." For someone in their first five years, that is not quite the right question. The right question is whether the job gets hired for at all.
Junior roles exist because there is a set of tasks that are worth paying a new person to do while they learn. Those tasks are the rung. If they leave, the job does not disappear; the people already in it keep it, and the next hire does not happen. The senior analyst is fine. The rung stops being backfilled.
Rung erosion measures how much of what a junior is specifically hired for is leaving. It uses the same task scores but weights them differently: not by share-of-week, but by how much each task figures in the reason to hire someone at entry level.
This is why it diverges from raw exposure. A customer service representative, by task exposure, is more automatable than a recruiting coordinator. But the rep's rung erosion is lower. Much of what remains in a service role after the routine tickets go is the live, unpredictable handling that still needs a person on the floor, and that is what juniors get hired for. A recruiting coordinator's entry-level rationale, on the other hand, is concentrated in scheduling, screening and candidate communication, which are exactly the tasks going first. The overall job is safer; the rung is not.
For early-career readers, the rung score matters more than the exposure score. We show both, and we explain which one is driving the recommendations.
Accountability is held or borrowed
Within the accountability dimension, we make a distinction that turns out to decide a lot.
Held accountability belongs to the person. A licence, a professional signature, a legal duty that attaches to a named individual. A notary's stamp. A pharmacist's verification. Nobody can reassign it without reassigning the person.
Borrowed accountability belongs to the organisation and is lent to whoever currently does the task. The bookkeeper is answerable for the reconciliation today, but the firm is answerable for the books, and the firm can hand that task to a vendor with a service-level agreement in a single quarter. The protection was real; it just was never theirs.
In our data, of roughly 690 task rows, 640 carry borrowed accountability. Only 16 are held. That ratio is the most sobering thing in the dataset, and it is why the reports push so hard toward moves that build held protection rather than resting on borrowed protection. When a report says a task is "protected", we tell you which kind.
Non-AI risk is scored separately
For 26 of our 50 roles, the dominant threat to the job over the next three years is not AI. It is an interest-rate cycle, store closures, branch consolidation, enrolment decline, or a federal budget. A mortgage processor's 2027 depends far more on the Fed than on any model. A role tied to a college's headcount depends on enrolment more than on any model.
We score this as its own line, with its own signals to watch, and we keep it out of the AI exposure number. A report that blames AI for a branch closure is simply wrong, and it would send someone to learn prompt engineering when they should be watching the regional bank's earnings call. Keeping the two apart is a correctness issue, not a presentation choice.
Every recommendation is verified
When a report recommends a credential, a course or a certification, we checked it against the provider's own page, on a date, and that date is printed on the row.
This turned out to be necessary. Of the credentials we checked, about a third had a problem: retired, renamed, price hidden behind a login, or in one case an issuing body that had gone bankrupt while its credential page still loaded. A recommendation that points to a dead credential is worse than none. We re-check on a schedule and remove what fails.
We take no money from any provider we recommend. That is a business rule, not a preference, and it is the only way the verification means anything.
What we don't know
The limits, stated plainly:
- The scores are judgement, not measurement. Each capability and protection score is a considered estimate, calibrated against published research on task-level AI exposure, not the result of an experiment. Two careful people could score a task differently. We have tried to be consistent; we have not proved we are right.
- We have not tracked outcomes. Nobody has yet followed a Pathcast plan for thirteen weeks and reported back. We cannot tell you that the moves work. We can tell you why we think they do.
- The task shares come from a conversation, not observation. Your share-of-week weights come from your résumé and your answers, not from watching you work. People are unreliable narrators of their own weeks. The report is only as accurate as that input, which is why we show you the parse and ask you to correct it.
- Fifty roles is fifty roles. We do not generate reports for roles we have not researched. If yours is not on the list, the honest answer is a waitlist, not a guess.
- The 0.75 is a choice. So are the internal weights and the band thresholds. We have explained our reasoning. We have not proved it.
What we calibrated against
Our approach draws on published work on task-level AI exposure, principally Eloundou, Manning, Mishkin and Rock, "GPTs are GPTs: Labor market impact potential of LLMs" (Science, 2024), which established the practice of scoring exposure at the task rather than occupation level, and on Anthropic's research on observed AI usage across occupational tasks. We used these to check that our capability scores were not wildly out of line with the published picture. Our protective dimensions, the formula, the rung erosion score and the held/borrowed distinction are our own, and none of the authors or organisations above have reviewed or endorsed this work.
If you find an error in a score, or think a task should be weighted differently, we want to hear it: [email protected].