Moneyball Never Made It to Work
Someone in your company has proposed doing Moneyball for the workforce. Possibly you. I've proposed versions of it myself, which is part of why I want to be careful here.
The pitch is always the same and it is always appealing. Baseball had a market full of people paying for the wrong things. A small-budget team worked out what actually produced wins, noticed the market had mispriced it, bought the underpriced thing, and won. Surely a company — sitting on performance data, engagement data, an HRIS, a comp structure — can do that with people.
Twenty years of trying says no. Not because the people trying were unserious; some of them were very good. It fails at a specific place, every time, and the place is worth naming precisely, because the imprecise version of the diagnosis is one I've seen repeated so often it's become received wisdom, and the received version is wrong.
The condition nobody states
Here's the received version: enterprise Moneyball fails because there's no outcome variable.
That's not right, and I need to say so carefully because I've argued the opposite in these pages. I've made the case that leaders should be judged against the measurable state of the systems they're responsible for rather than against how they present in a room. I've put my name on a metric — net activated value — built specifically to put human capital and dollar outcomes on the same axis. If I now tell you there's no outcome variable, I'm arguing against myself, and you should stop reading.
The honest statement has three conditions, and it's the third one that does the damage.
There is no agreed-upon, consistent outcome variable framed in dollars.
Agreed-upon: the field does not concur on what organizational success even is. Consistent: whatever a given company picks, it isn't measured the same way across contexts or across years. And framed in dollars — this is the binding one, and the one that gets skipped. The whole point of the exercise is to compare what a person's or a job's content is worth against what it's paid. That comparison requires the outcome to be denominated in the same currency as the pay. Runs become wins, and wins have a dollar price because a market bids for them. Nobody has to be persuaded that a win is worth something; the market states the number every winter.
Organizational performance has no such denomination. It has proxies, and the proxies are in different units, and reasonable people rank them differently.
Have I proposed a candidate? Yes. Do I think it's right? Yes. Is it proven? No. And until it is, I am not going to build a pricing model on top of it and hand you the output as though the foundation were settled. That constraint isn't modesty. It turns out to be the most useful thing in this essay, because working around it produces a better method than waiting for it would have.
The criterion problem is older than the computers
The measurement people have had a name for this since before anyone had a laptop to put a dashboard on. They call it the criterion problem, and Austin and Villanova's history of it runs from 1917 to 1992 — seventy-five years of the same difficulty, catalogued.1
The shape of it: you can build a predictor and check its properties, but the thing you want to predict — actual job performance — is the part you can never quite pin down. Thorndike, in 1949, called the ideal version the ultimate criterion: "the complete final goal of a particular type of selection or training."2 A complete final goal. Not something you measure on a Tuesday.
I want to be precise about how far this goes, because the fashionable version overshoots. Performance is not unmeasurable. Austin and Villanova report predictor reliabilities averaging around .80 against criterion reliabilities averaging around .60 — worse, meaningfully worse, but nowhere near noise.1 Sturman and colleagues put the test-retest reliability of performance ratings between .83 for subjective ratings in low-complexity jobs and .50 for objective measures in complex ones, with one-year stability from .85 down to .67.3 Those are moderate-to-good numbers. Anyone telling you performance can't be measured is overselling a real problem.
The real problem is narrower and harder to escape: performance is multidimensional, it drifts with time and job complexity, and every operational measure you actually collect is a partial stand-in for something you agreed to stop arguing about. That is survivable if you want to predict performance. It is fatal if you want to decompose it — to say this much of the outcome traces to this input, priced at this many dollars. Decomposition needs the criterion at the individual grain, denominated in money, stable enough to regress on. That's the version we don't have.
Then it turns out baseball didn't quite have it either
Here's where I'd hoped to write the clean contrast, and where the sources refused to cooperate — which is the more interesting result, so it's the one you get.
The canonical academic treatment is Hakes and Sauer, in the Journal of Economic Perspectives in 2006.4 They did the honest thing: regressed team winning percentage on on-base percentage and slugging, found OBP's coefficient roughly twice slugging's, then regressed player salaries on the same inputs and found the market paying handsomely for slugging while OBP came back statistically insignificant. The market was buying the wrong thing. Then in 2004 — after the book — the ordering flips, and OBP finally gets paid. Mispricing found, mispricing exploited, mispricing corrected. It's a beautiful result and it's the one everybody cites.
The follow-up is less famous. In their own 2007 paper, Hakes and Sauer ran the years forward: OBP is significant in 2004 and 2005, and insignificant again by 2006. The correction didn't hold.
And the core comparison has been challenged on grounds that should make any analyst reading this uncomfortable. Yashiki and Nakazono replicated the result, then pointed out that the OBP-beats-slugging conclusion comes from comparing unstandardized regression coefficients across two variables with different variances — an apples-to-oranges move that first-year econometrics warns about. Standardize them, and slugging contributes more to winning than on-base does, before 2003 and after, with the gap widening post-Moneyball.5 On that reading, the market didn't get more efficient after the book. It may have gotten less.
There's one more detail I can't skip, because it's about to be our problem. OBP and slugging correlate at about .75.5 They move together. Hakes and Sauer knew it, and by 2007 were decomposing hitting into separate "eye," "bat," and "power" components precisely to get around it — because when two inputs travel together, the data cannot cleanly tell you what either one is worth on its own.
So: the cleanest case anyone has ever had — one outcome, universally agreed, measured every night, decomposable to the individual, denominated in a currency a market publicly bids — is still contested twenty years on, on collinearity and on standardization.
Sit with that before proposing to do it with engagement scores.
What actually transfers
The wrong lesson is that it's hopeless. The right lesson is that we spent two decades trying to import the wrong half.
The half everyone tried to import is the outcome model — find the thing that causes success, weight the inputs, re-price accordingly. That half requires everything the enterprise doesn't have, and, on the evidence above, is fragile even where the conditions are perfect.
The half that transfers is quieter. It's the willingness to ask whether the market's price and the market's own evidence agree — and to treat a gap between them as a finding rather than an error in the data.
You don't need an outcome variable to do that. You need prices, and you need to know what's inside the things being priced. Sherwin Rosen laid the groundwork in 1974: a differentiated good is a bundle of characteristics, and the prices at which such goods trade reveal a set of implicit prices for the characteristics inside them.6 A house isn't priced; square footage, school district, and a garage are priced, and the house is where you read the sum. A job isn't priced either. Its content is.
I'll flag my own limit here, because Rosen's method has a well-known seam. Recovering implicit prices from observed prices — his first stage — is the robust, widely used part. Going on to recover what a characteristic is worth to a particular buyer — the second stage — is notoriously fragile and remains unresolved.6 The work below lives in the first stage. When it starts drifting into the second, that's the moment to stop and say so.
And this isn't hypothetical: labor economics has been pricing job content against wages for two decades without ever needing a performance outcome. Autor, Levy, and Murnane built the task framework in 2003, sorting work into routine and nonroutine, cognitive and manual, and tracking what happened to each as computing spread. Their model accounts for 60 percent of the estimated relative demand shift favoring college labor from 1970 to 1998, with task changes inside nominally unchanged occupations doing almost half the work.7 Deming carried it into social skill: jobs requiring high social interaction grew nearly 12 percentage points as a share of the U.S. labor force between 1980 and 2012, while math-intensive but less social jobs — many of them STEM — shrank by 3.3 points.8
Neither paper ever asked whether a task caused an outcome. They asked what the market paid for it. That is the whole move.
Where the money actually is
Which points at the question worth asking, and it isn't which content produces success. It's this:
Does identical content fetch different prices depending on the wrapper it comes in?
Title, function, employer, industry, geography, pedigree. If the same work commands materially different pay depending on what it's called and where it sits, that's not a causal claim about performance. It's a violation of the law of one price — an arbitrage, sitting in plain sight, requiring no theory of organizational success whatsoever.
And we have known for a long time that labor markets are full of these. Krueger and Summers found that after controlling for education, age, occupation, sex, race, union status, and more, the weighted standard deviation of inter-industry wage differentials was still .146 — down from .240 raw, but the ranking barely moved, correlating .95 with the uncontrolled version. Petroleum paid 38 percent above average for comparable workers; eating and drinking places paid 19 percent below.9 The modern firm-level work says the same thing with better methods: bias-corrected estimates put the standard deviation of firm wage effects around .155 against .334 for worker effects — firms and workers of broadly comparable importance in setting what someone earns.10
Two cautions, both of which I'd rather state than have pointed out to me. First: the most-cited result in this literature — Abowd, Kramarz, and Margolis's 1999 finding that industry differentials are mostly person effects — turned out to be an artifact of the computational method used to approximate the least-squares solution, not a finding about labor markets.10 It is still cited constantly. Don't.
Second, the collinearity problem from baseball is waiting for us at a far worse scale. Run a principal-components analysis on O*NET's descriptors — the roughly two hundred skills, abilities, knowledge areas, and work activities that constitute the public map of what jobs contain — and the first ten components account for 96 percent of the variation.11 Two hundred measures, ten dimensions of actual information.
And the sharpest statement of it comes from a source I find hard to ignore: when the National Academies reviewed O*NET, David Autor — the same Autor whose task framework I just leaned on — co-authored a dissent appended to the report. Regress any single ability or skill rating on the set of work-activity and work-context ratings, he and Juan Sanchez found, and you get multiple correlations from .65 to .98. A single factor accounts for 43 percent of the variance across all 52 ability ratings.11
Content co-occurs. Nearly everything correlates with nearly everything, and a descriptor that never varies independently of its neighbors has no independent price for a regression to find. That's my inference rather than a published finding, and I'd rather label it than smuggle it: only differentiating content — what varies where the rest doesn't — can carry a price at all. A decomposition that forgets this will hand you confident numbers that are pure artifact, and they'll be confident in exactly the way the batting-average era was confident.
That's a real constraint. It isn't the same as needing an outcome variable, and it has a known remedy, which is more than the criterion problem can say.
Arbitrage, not sabermetrics
So Moneyball never made it to work, and the reason is not that people in companies are harder to measure than people in baseball, though they are. It's that the trick everyone tried to copy was the outcome decomposition, and the outcome decomposition rests on a foundation the enterprise has never had and — on the honest evidence — cannot presently manufacture. I'd like to be the one who manufactures it. I've proposed a candidate. I'll tell you when it's proven, and I'm not telling you that today.
The version that works needs less. It asks whether the market prices the same work differently depending on the wrapper, and it answers using the market's own revealed prices as the standard rather than a theory of success as the standard. It cannot tell you what makes a team win. It can tell you where a price and the work behind it have come apart — and that is a finding you can act on, in both directions, without knowing why the world is the way it is.
Beane wasn't right because he understood baseball more deeply than anyone else. He was right — for about two seasons, on the current evidence — because he checked a price against the evidence and found they disagreed. That's not sabermetrics. That's just refusing to accept a number because everyone else already had.
That part travels.
This is a companion in the Content Pricing program. Its predecessor, An Occupation Is Not a Job, establishes the unit any of this has to run on; the piece that follows takes up what the dispersion actually looks like, and does not publish until the number exists. Adjacent in method: Borrowed Validity on treating a published estimate as a prior rather than a verdict, and The Error Bar Is the Product. Every source below was read rather than cited from memory, and where the popular version of a finding diverges from the paper's — the Moneyball correction, the 60 percent figure, the Abowd-Kramarz-Margolis result — this piece follows the paper.
Footnotes
-
John T. Austin & Peter Villanova, "The Criterion Problem: 1917–1992," Journal of Applied Psychology 77, no. 6 (1992): 836–874. Reports predictor reliabilities averaging ≈.80 against criterion reliabilities averaging ≈.60. See also Hubert E. Brogden & Erwin K. Taylor, "The Theory and Classification of Criterion Bias," Educational and Psychological Measurement 10 (1950): 159–186, for the deficiency/contamination/distortion taxonomy, and John P. Campbell's eight-factor performance model in Handbook of Industrial and Organizational Psychology, 2nd ed. (1990), vol. 1, 687–732. ↩ ↩2
-
Robert L. Thorndike, Personnel Selection: Test and Measurement Techniques (New York: Wiley, 1949), 121, defining the ultimate criterion as "the complete final goal of a particular type of selection or training" — an abstraction far removed from any actual criterion measure. The stronger formulation, that the ultimate criterion cannot be measured or observed at all, is a later elaboration (Binning & Barrett, Journal of Applied Psychology 74, no. 3 [1989]: 478–494), not Thorndike's own wording. ↩
-
Michael C. Sturman, Robin A. Cheramie & Luke H. Cashen, "The Impact of Job Complexity and Performance Measurement on the Temporal Consistency, Stability, and Test-Retest Reliability of Employee Job Performance Ratings," Journal of Applied Psychology 90, no. 2 (2005): 269–283. Test-retest reliability from .83 (subjective ratings, low-complexity jobs) to .50 (objective measures, high-complexity); one-year stability .85 to .67. ↩
-
Jahn K. Hakes & Raymond D. Sauer, "An Economic Evaluation of the Moneyball Hypothesis," Journal of Economic Perspectives 20, no. 3 (2006): 173–186. Yearly salary regressions 2000–2004 on prior-season OBP and slugging, non-pitchers with ≥130 at-bats, controlling for plate appearances, arbitration/free-agency status and position. Their follow-up — Journal of Sports Economics (2007) — reports OBP significant in 2004 and 2005 and insignificant again in 2006, and decomposes hitting into "eye," "bat," and "power" components to address collinearity between the inputs. ↩
-
Yashiki & Nakazono, ESRI Discussion Paper No. 362 (2021). Successfully replicate Hakes & Sauer, then show the OBP-over-slugging conclusion depends on comparing unstandardized coefficients across variables with unequal variances; with standardized coefficients, slugging contributes more to winning both before and after 2003, with the gap widening after 2004. Reported correlation between OBP and slugging in their data: .746. See also Deli (2013). ↩ ↩2
-
Sherwin Rosen, "Hedonic Prices and Implicit Markets: Product Differentiation in Pure Competition," Journal of Political Economy 82, no. 1 (1974): 34–55. The first stage — recovering implicit characteristic prices from observed product prices — is the robust and widely applied part; the second stage, recovering structural willingness-to-pay from those implicit prices, faces a well-documented and unresolved endogeneity problem. The argument here stays in the first stage deliberately. ↩ ↩2
-
David H. Autor, Frank Levy & Richard J. Murnane, "The Skill Content of Recent Technological Change: An Empirical Exploration," Quarterly Journal of Economics 118, no. 4 (2003): 1279–1333. The published paper reports 60 percent of the estimated relative demand shift favoring college labor, 1970–1998 — note that the freely circulating 2001 NBER working paper says "thirty to forty percent," and the published figure is the one to cite. Task measures are DOT-based, not ONET; the authors note ONET provides no time series on within-occupation job content. The "routine work has been declining since 1960" summary is wrong in both routine categories: in centiles of the 1960 task distribution, routine cognitive input rose from 50.0 to 51.8 by 1980 before falling to 44.4 in 1998, and routine manual rose to 53.8 by 1980 before returning to 49.2 — essentially flat across the full period. Nonroutine manual (50.0 → 41.3) is the only measure declining across every decade. ↩
-
David J. Deming, "The Growing Importance of Social Skills in the Labor Market," Quarterly Journal of Economics 132, no. 4 (2017): 1593–1640. High-social-interaction jobs grew 11.8 percentage points as a share of the U.S. labor force 1980–2012; math-intensive, less-social jobs shrank 3.3 points. The load-bearing result is the interaction — jobs high in both math and social skill grew and paid most; cognitive skill still returns roughly three to four times more per standard deviation than social skill, and what changed is the trend, not the ranking. ↩
-
Alan B. Krueger & Lawrence H. Summers, "Efficiency Wages and the Inter-Industry Wage Structure," Econometrica 56, no. 2 (1988): 259–293; magnitudes as reported in the companion NBER Working Paper 1968 (1986), Table 1, 1984 CPS, n=10,289. Weighted standard deviation of industry differentials .240 unadjusted, .146 with full controls; correlation between adjusted and unadjusted differentials .95. Adjusted premia include petroleum +38.2%, public utilities +28.7%, apparel −15.6%, eating and drinking −18.8%. Krueger and Summers themselves flag the dissenting reading (Murphy & Topel) that much of this is unobserved individual heterogeneity. ↩
-
Patrick Kline, "Firm Wage Effects," NBER Working Paper 33084 (2024, rev. 2025), forthcoming in the Handbook of Labor Economics. Bias-corrected estimates across nine international studies place firm and worker effect magnitudes broadly comparable; the Veneto benchmark gives a bias-corrected firm-effect SD of .155 against a person-effect SD of .334. Kline also documents that the headline finding of Abowd, Kramarz & Margolis, "High Wage Workers and High Wage Firms," Econometrica 67, no. 2 (1999): 251–333 — that industry differentials are largely person effects — was driven by the computational method used to approximate the least-squares solution in their largest samples (Abowd, Creecy & Kramarz, 2002), and should not be cited for that claim. ↩ ↩2
-
Principal-components result: Morgan R. Frank, Yong-Yeol Ahn & Esteban Moro, "AI exposure predicts unemployment risk" (arXiv:2308.02624), supplementary information §1.4 — across roughly 120–230 skill, knowledge, ability and work-activity descriptors per year over ~700 SOC occupations, 2010–2020, the first ten principal components account for 96% of total variation in occupations' skill requirements. Multiple-R and single-factor results: Juan I. Sanchez & David H. Autor, Appendix A to National Research Council, A Database for a Changing Economy: Review of the Occupational Information Network (O*NET), ed. Nancy T. Tippins & Margaret L. Hilton (Washington, DC: National Academies Press, 2010). That appendix is titled "Dissent" — these are Sanchez and Autor's statements, not the consensus finding of the NRC panel, and should not be attributed to the National Academies. The panel's own chapter 2 separately reports field-test factor analyses yielding fewer factors than investigators expected and recommends research to reduce descriptor redundancy. See also Michael J. Handel, "The O*NET content model: strengths and limitations," Journal for Labour Market Research 49, no. 2 (2016): 157–176. The further step — that this collinearity leaves individual descriptors' implicit prices econometrically unidentified — is my inference from the evidence above, not a published result, and is labelled as such in the text. ↩ ↩2