peopleanalyst

magazine · Methodology · units & aggregation

Forty-one of our jobs stand behind one federal occupation code, from entry level to principal, sharing one median. Occupation-level data answers occupation-level questions — and we keep asking it job-level ones.

By Mike West

July 21, 2026

An Occupation Is Not a Job

Somewhere in the last month, someone in your company priced a job by looking up a code.

It goes like this. A req opens. The recruiter needs a range, the finance partner needs a number, and the compensation analyst does the thing the whole apparatus is built to do: matches the job to a market survey, which is matched in turn to a federal occupation code, and reads the number off the row. For a lot of people reading this, the row was 13-1071, Human Resources Specialists. Median pay, $72,910.

I want to tell you what else is standing in that row.

In our own job canon — the structured map of jobs we maintain, built from the corpus rather than from a survey — forty-one distinct jobs resolve to 13-1071. Not forty-one job titles. Forty-one jobs, each defined by a family, a focus, and a level. HR Business Partner. HR Generalist. HR Transformation. HR and People Operations. Talent Acquisition. Global Mobility. And People Analytics and Workforce Planning — which is to say, several of you, reading this. They run from P1, the first job a person holds, to P6, the one a person spends a career reaching.

One code. One median. An entry-level generalist and a principal people-analytics leader, standing in the same row, wearing the same number.

That is not a data-quality problem. Nothing is broken. The classification is doing precisely what it was built to do, and the trouble starts the moment we ask it to do something else.

The standard says so out loud

The Standard Occupational Classification is not hiding this. Its own definition of an occupation is a group of jobs that are similar with respect to the work performed and the skills possessed by workers.1 A group of jobs. The container is the unit; the jobs inside it are what got contained. And the purpose the classification states for itself is equally plain — it exists so that federal agencies can classify workers for the purpose of collecting, calculating, or disseminating data.1 It was designed for counting. It is very good at counting. Eight hundred sixty-seven detailed occupations is a magnificent instrument for answering how many people in this country do roughly this kind of work, which is a question a republic genuinely needs answered.

It was not designed to answer what should we pay this person, or does the market value this skill in this role, or is our job architecture right. Those are job-level questions. Every one of them is now routinely answered with occupation-level data, by people who would never dream of answering a question about individuals with a statistic about states — which is the same mistake, and it already has a name.

Robinson, 1950

In 1950 a sociologist named W. S. Robinson published six pages in the American Sociological Review that should be taped inside the lid of every analytics laptop.2

He was looking at 1930 census data on literacy. Ask whether foreign-born residents were more often illiterate than native-born ones, at the level of individual people, and the answer is yes — a correlation of +.118. Modest, but the sign is what you'd expect and the direction is real.

Then Robinson asked the same question of the same data, aggregated. Group the country into its nine census divisions, correlate percent foreign-born against percent illiterate, and the correlation is −.619. Do it by the forty-eight states and it's −.526.

The sign flipped. Same population, same year, same two variables. The individual answer and the group answer aren't just different in magnitude; they point in opposite directions, and the aggregate one tells you a confident, clean, entirely false story about people. He found the same shape everywhere he looked: color and illiteracy correlated .203 among persons and .946 among divisions — nearly five times larger.

And the part that ought to make an analyst's stomach turn: the coarser the grouping, the bigger the number gets. States gave .773; the larger divisions gave .946. Robinson traced this to work by Gehlke and Biehl from 1934, who had already noticed that the coefficient grows with the size of the areas it's computed over.2 Aggregate harder and the finding looks stronger. Nothing about the world changed. You just stopped being able to see the variation you averaged away, and the residue reads as signal.

Robinson's conclusion runs two words, and he does not soften them. Can ecological correlations validly be used as substitutes for individual ones? "They cannot."

An occupation is an ecological unit. A job is the individual. Everything Robinson said about states and persons applies, without amendment, to occupations and jobs — and we have spent seventy-five years building an entire market-pricing infrastructure on the substitution he told us not to make.

What we found when we counted our own

Here is the part I'd rather not write, which is usually the sign it's the part worth writing.

We measured the collapse in our own canon, because we were about to build on top of it. Across the jobs whose mappings resolve to a specific occupation rather than a whole occupational family: 956 jobs land on 84 detailed occupations. The mean is 11.4 jobs behind one occupational number. The median is 6. The worst case is 49 — that's 11-2021, Marketing Managers, which holds Marketing, Brand Management, Demand Generation, Growth Marketing, Commercial Strategy, and Product Management, the last of those running out to P7. A staff product manager and a first-year demand-gen marketer, same row, same median.3

And then the second collapse, which is the one people miss. It is not just that eleven jobs sit behind one number. It is that the number itself was never a number. The published wage data for 13-1071 has a 10th percentile of $45,440 and a 75th of $97,270 — a 2.1× spread inside the occupation, before we collapsed anything into it. Marketing Managers runs $81,900 to $211,080, a factor of 2.6.3 We take a distribution that wide, hand it a median as its representative, and then park eleven distinguishable jobs behind that median and call the result a market rate.

That's the machine. Two aggregations stacked, each individually defensible, and the product of them is a number that cannot be wrong because it isn't specific enough to be wrong about anything.

Reliability can't see this

Now the part that connects this to everything else I've argued in these pages.

We had built an instrument to compare what a field's own literature says a role is for against what the occupational description says it's for — an ensemble of raters, measured properly, agreement coefficients and all. It worked. The agreement was real. The reliability was fine.

The reliability was fine and the unit was wrong, and no agreement coefficient can detect that, because every rater in the ensemble was handed the same over-collapsed unit. They agreed with each other beautifully about a question that had eleven jobs jammed into its subject. Reliability tells you your raters are consistent. It is silent — structurally, permanently silent — on whether the thing they were consistent about was the thing you meant to ask. I've written before that a measurement can be perfectly reliable and measure the wrong construct. This is that failure with a federal code number on it.

I'll go one worse, because it's true and it's mine. When we went to write these figures down, the first version was wrong. It said 974 jobs across 86 occupations. Two of those 86 weren't occupations at all — they were two-digit major-group fallbacks, the classification's own coarse bucket, quietly counted as though they were specific. And the 974 was one of two halves of a canon that totals 2,158, quoted without saying so, because the other half's mappings are all major-group fallbacks and had to be excluded. Excluding them was right. Not saying we'd excluded them was the same error as the one this whole essay is about, committed by the people writing the essay, inside a week.

The number was in a document. Documents do not recompute. So the fix wasn't to correct the sentence — it was to make the figure resolve to something that regenerates, with the excluded rows left visible in the output so anyone can see what was dropped and argue with the decision.3 Every number above came out of that. If the canon changes tomorrow, the number changes with it, and the essay is the thing that goes stale rather than the truth.

What the corrected shape looks like

The repair is not exotic, and it is not "get better occupation data." It is a change of unit, and it costs you the comfortable part.

Make the job the unit of identity — family, focus, level, the way an actual organization experiences it. Then treat the occupation as what it honestly is: one source of evidence about that job, pooled across whatever occupational codes the job legitimately touches, rather than an identity the job is forced to wear. And carry the pooling ratio forward as a confidence input, so a job standing behind a code shared with forty others arrives on the desk with a wider interval than one standing nearly alone. The collapse doesn't disappear. It becomes visible, and it becomes a number you can reason with instead of an assumption you can't see.

Some of what this returns is uncomfortable, which is how you know it's working. Occupations carry no level at all — the SOC has no dimension for it — so any apparent seniority signal you were reading off occupational data was the aggregation absorbing a mismatch and handing it back to you as a feature. And a job that pools across three codes, none of them close, is a job the market data cannot price at the grain you asked. The honest output there is a wider band and a stated reason, not a tighter number produced by picking the friendliest code.

That is a worse-looking product and a better-behaved one. The error bar is the product; I've made that argument at length. This is what it looks like when the error bar comes from the unit rather than from the sample.

The row you were standing in

Go back to the req.

The recruiter still needs a range and the finance partner still needs a number, and nothing I've written makes 13-1071 stop existing or stop being useful for what it was built for. If you want to know how many people in the United States do broadly HR-shaped specialist work, that code is one of the better instruments a country has ever built for finding out. Ask it that and it will answer honestly.

Ask it what to pay the people-analytics lead and it will also answer, in the same confident voice, using a median drawn across forty-one jobs and six levels and a wage band that already spans two-to-one before you started. It will not flag the problem. It cannot flag the problem. A container has no way of telling you that you've mistaken it for its contents.

That is the same thing I've said about themes in open text, and about a reliability coefficient, and it keeps being the same thing: a measure that cannot disagree with you is not evidence. The theme list can't tell you your theory is wrong. The agreement statistic can't tell you the question was malformed. And the occupation code can't tell you it's holding forty-one jobs, because from where it sits, it's holding one — the group. That was always its job.

Ours is to stop asking it about individuals.


This is a companion in the Content Pricing program — the argument that you can price what a job's content is worth against the market's own revealed prices, and that the first requirement for doing so is a unit that survives contact with a real organization. Its siblings in method are The Reliability Problem, Themes Aren't Evidence, and The Error Bar Is the Product; The Benchmark Trap takes up the neighboring case where the comparison group, not the unit, is doing the damage. Every figure about our own canon in this piece resolves to a committed, regenerable artifact rather than to this essay — docs/artifacts/occupation-grain.json, regenerated by npm run build:occupation-grain. The wage figures are OEWS national, 2024. Where the first published version of a number was wrong, this piece says so rather than quietly restating it.

Footnotes

  1. U.S. Bureau of Labor Statistics, 2018 Standard Occupational Classification System. The SOC is the federal statistical standard used by agencies to classify workers into occupational categories for the purpose of collecting, calculating, or disseminating data; the 2018 revision defines 867 detailed occupations across 23 major groups. Its classification principle defines an occupation as a group of jobs similar in the work performed and the skills workers possess — i.e. the standard itself states that the unit is a group of jobs, not a job. https://www.bls.gov/soc/ 2

  2. W. S. Robinson, "Ecological Correlations and the Behavior of Individuals," American Sociological Review 15, no. 3 (1950): 351–357. Nativity and illiteracy, U.S. 1930, population 10 and over: individual correlation +.118; ecological correlation −.619 across the nine census divisions and −.526 across the 48 states. Color and illiteracy: individual .203; ecological .946 (divisions) and .773 (states). Robinson credits Gehlke and Biehl (1934) for the observation that the ecological coefficient rises with the size of the sub-areas it is computed over. His conclusion on whether ecological correlations may substitute for individual ones: "They cannot" (p. 357). 2

  3. Job-collapse figures computed over the JobFrame canon by scripts/build-occupation-grain.mjsdocs/artifacts/occupation-grain.json (people-analytics-toolbox). Detailed six-digit SOC mappings only: 956 jobs → 84 occupations, mean 11.38, median 6, worst case 49 (11-2021). Two-digit major-group fallbacks — mappings inferred from super-function and flagged as needing review — are excluded, and are retained inside the artifact as a labeled excluded stratum so the exclusion is auditable rather than assumed. Wage bands are OEWS national 2024 as carried in the same artifact: 13-1071 median $72,910 (p10 $45,440, p75 $97,270); 11-2021 median $161,030 (p10 $81,900, p75 $211,080). 2 3

Was this useful?

Anchored in

Keep going

New issues, oriented to your goals — methodology-first, source-anchored, not a firehose.

The store

The data and tools behind the writing

Compensation and compliance datasets, drop-in code packs, and the deep guides — posted prices, buy now.

Browse the store →
← All magazine pieces