Everything Is Already Known, But Almost Nobody Who Needs to Know, Knows It
You are not behind because the answer hasn't been worked out. You're behind because it was worked out by someone whose work you will never read.
Somewhere this quarter a company is going to redesign how it pays people, and it will do it worse than the field already knows how to do it.
Not through laziness. The person running it will be capable and will care. They'll read a few articles, talk to a peer at another company, look at what a vendor recommends, and make a defensible-sounding call. And the specific thing they get wrong will be a thing that was settled — genuinely settled, with evidence — sometime between 1985 and about four years ago, in material they have no realistic path to.
This is the ordinary condition of professional work, and it is worth being precise about the diagnosis, because the obvious diagnosis is wrong. The problem is not that nobody knows. On almost any question an organization faces, somebody knows, and often they have known for thirty years. The problem is that you don't know what other people already know, and there has never been a good way to fix that in the time you actually have.
What we got, and what it didn't solve
Give credit where it's enormous. The internet, and then Google, moved humanity forward by light years on exactly this problem. The distance between a question and someone who has already answered it collapsed from a library trip and a week to about four seconds. Nothing since has matched that, including what came next.
But it left two things unsolved, and both of them bite hardest on the questions that matter most.
The first: search requires you to already know the question. That sounds like a small constraint until you notice that the expensive errors are never the ones you're searching about. The compensation lead searching how to build a salary structure will find a great deal of competent material on building salary structures. What they will not find, because they don't know to ask, is the argument that their whole premise — a structure priced at the market median across every job — quietly makes a decision about internal equity that they never chose and can't defend. You cannot search your way to a question you don't have.
The second: what comes back is a list of documents you now have to read. Google didn't answer you. It routed you. And the routing points at the open web, which is a genuinely poor sample of what's known — because the best material on most professional subjects is in books, behind publishers, and in journals, behind paywalls and behind a dialect most people can't parse. The most reliable knowledge is the least reachable. That is close to the exact opposite of what you'd design.
Then AI, which is the real thing and also a new problem
I don't think it's overheated to say that AI has changed how a person can engage with knowledge more than anything since the search engine. Five years ago you could not ask a question in your own words, about your own situation, and get a coherent synthesis back. Now you can, and it is genuinely useful, and I use it constantly.
Here's the part nobody has solved. The answer comes back confident. It comes back fluent, well-organized, appropriately hedged in the places a careful writer would hedge — and you have no way to determine what it rests on. How much of that paragraph came from a real literature? Which part is a faithful summary of a real finding, and which part is the model doing what models do, producing the sentence that should come next given everything it has read?
You cannot tell. Not by reading harder, not by asking it to cite — because it can produce a citation the same way it produced the claim. And the failure isn't loud. A fabricated policy reads exactly like a real one; a slightly-wrong effect size reads exactly like a right one. The trouble with a confident answer is that it doesn't announce which kind it is.
The field is working on this. It is not solved, and I'd be careful of anyone who tells you it is. So the honest question for a working professional right now is: while that gets worked out, what do you actually do?
The fallback, and why it doesn't scale
The fallback is to go back to the primary material — books written by people credible enough to assemble a body of knowledge and get a publisher to stand behind it. That's a real quality filter. It's also where you immediately hit arithmetic.
How many do you need to read to have a complete view of a subject? Not one — one book is one person's argument, and you can't tell from inside it which parts the field shares and which are that author's particular hobby-horse. Five gets you overlap. Ten gets you a real sense of consensus. Thirty gets you the shape of the whole field, including who the outliers are.
And then, on the claims that turn out to be novel or fringe: how does the author actually support this? Is the evidence compelling, or is it an anecdote deployed with great confidence? That's a second reading pass, on the parts that are hardest to read.
Meanwhile there's a second shelf entirely. For anything a science can address — and most questions about people at work can be — there's the peer-reviewed literature. It's more rigorous by construction: it requires the math, it requires the method, and it has to survive people whose professional interest is in finding the hole. It is also frequently in tension with the practitioner books. Sometimes it is years ahead of them. Sometimes, less often than academics like to think, it's behind — a lab result that doesn't survive contact with a real organization.
And the two shelves never meet. That's the part I find genuinely maddening. It would be enormously valuable to a practitioner to vet the popular claim against the research, or to take a robust research finding and try it where the work happens. Almost nobody does. Not because they don't care — because parsing a journal article is a skill, it takes an hour, and they have a job.
So the arithmetic finishes itself. Nobody reads thirty books on a subject. Most people read about one book a year, and if we're being honest about it, retain a fraction. That's not a character flaw. It's a time budget, and any solution that ignores it isn't a solution.
So we built the thing that reads them
That's the whole origin of what we call a Bicycle Guide: the whole of a subject, in about an hour.
We read the canon of a field — the books the field wrote to explain itself — and build one model out of them rather than thirty summaries. The guide names where the field broadly agrees, and it names the outliers instead of averaging them away; on an outlier it takes a position, weighed by the quality of the evidence behind it. Every claim traces back to the books it came from, so you can go read the source when it matters. And it's ordered as a route — foundations to summit — rather than as an encyclopedia, because the point is to get you somewhere, not to be comprehensive at you.
What it explicitly is not: a summary. A free summary is one voice with no receipts. The receipts are the product.
Once a subject exists as a corpus and a model rather than a shelf, things become possible that weren't before — and this is where it stops being a reading aid and starts being an instrument. You can take any claim and test it against the corpus: does the field actually agree with this, or is it one author? You can take that same claim to the research literature and ask a harder question: is there evidence, and what does it say? We've started doing exactly that across a range of subjects. It is the thing I most wanted to exist for twenty years and could not have built even five years ago.
The part where it gets hard
I've been at this, with successive generations of AI, for the better part of two years. What exists now is the accumulation of that — and of a lot of failure, most of it the failure modes the whole field keeps writing about. I want to walk some of the problems honestly, because if you're building anything in this direction you will hit every one of them, and because a description of a system that only reports its successes should make you suspicious.
Getting the material out of the book at all. A book is not a database. It's a long argument in one person's organizing logic, with the load-bearing claim sometimes stated once, in a subordinate clause, three chapters after the section that would make you look for it. Getting a usable structure out of that — the claims, the conditions, the mechanism, what the author actually rests it on — is the first hard problem and it stays hard. Do it too shallowly and you get a summary, which is what everyone else has. Do it too eagerly and the machine finds structure that isn't there, which is worse than a summary, because it looks like rigor.
Doing it as a process rather than a heroic run. One book, done carefully with a good model and a patient human, is a demo. A field is thirty books, and thirty is not one thing thirty times — it's a multi-stage pipeline with cost at every stage, failures that have to be caught before you pay for the next step, and an absolute requirement that a rerun doesn't quietly produce a different answer. Most of the engineering here isn't clever. It's the unglamorous business of making an expensive multi-step process idempotent, resumable, and honest about what it already did.
Encoding it so one extraction serves many uses. The same underlying claim has to work in a guide a person reads, in a model that gets compared across books, in a service some software calls, and in a check that runs against research. That means the extraction can't be shaped for the surface that happens to need it first. It has to be tagged and structured for uses that don't exist yet — which is a design bet you make early and pay for repeatedly if you get it wrong.
Assembling the right corpus, for a particular reader with a particular problem. A subject is not a shelf. A guide for a founder bringing a regulated product to market and a guide for an operator running a factory floor may share books and share almost no useful claims. Deciding what belongs in a corpus — for this persona, for this problem — is an editorial judgment that has to be made explicitly, because if you don't make it, the corpus makes it for you by whatever you happened to own.
The question that was dumb, and what it exposed
I want to take one failure in public, because it's the most instructive thing that happened to us this year.
Our claim-checking stage takes a guide's claims to the peer-reviewed literature. On a guide about how to get a patent, it came back unable to support a single claim. My first instinct was that something was broken. Nothing was broken. Retrieval worked, the papers it found were real, and the judge was right to reject every one of them.
The question was dumb. We had, in effect, asked what does the organizational-science literature say about patent prosecution? — and the honest answer to that is nothing, and why on earth are you asking me. A worthless question, competently executed.
But strip the embarrassment off and there's something underneath worth more than the guide was. To check a claim you need a corpus of credible sources capable of adjudicating that claim — and that corpus is not free. It has to be acquired, which costs money. The relevant material has to be extracted, which costs work. It has to be tagged so it can be found by a claim rather than by a keyword, which costs design. We do that work, and we'll keep doing it, and it is a permanent line item rather than a phase.
What that means practically is that our guide ambitions run ahead of our claim-checking ambitions, and they always will. We can build a corpus on how to get a patent — there are books, written by people who have done it, and they are the real thing. That does not mean we have a body of peer-reviewed science standing behind patent practice, and for some subjects one may not meaningfully exist.
So the fix isn't to acquire everything. The fix is to be honest about which universe a claim lives in, and to check it against the corpus we actually have — do the field's own books agree on this, is it one author's position, what does he rest it on — rather than routing it to a literature that was never going to answer. Where a scientific literature genuinely bears on the subject, that check is enormously valuable and we run it. Where it doesn't, saying so is the correct output. Then, separately and deliberately, we decide whether that gap is worth the acquisition. Sometimes it will be. That's a budget decision, not a bug report.
The copyright question, which I'd rather answer than dodge
There's an obvious shortcut sitting next to everything I've described: put all the books in a model and let it answer questions about them.
I won't pretend to settle the law here — it's unsettled, it's being litigated, and anyone giving you a confident answer about where it lands is doing the thing this whole essay is about. But you don't need a ruling to see that the shortcut is uncomfortable. What it produces is a thing that stands in for the books, generated from the authors' expression, delivered without attribution, to a reader who now has less reason to buy any of them. Whatever a court eventually says, that is a substitute for the work, built out of the work.
What we do is a different act, and the difference is in the design rather than in the disclaimer.
We put an independent set of questions to each book — questions derived from what a practitioner needs to decide, not from the author's table of contents. The answers come back organized by our structure, which is deliberately not the author's structure, because the selection and arrangement of a book is itself the author's creative work and inheriting it is how you end up reproducing it. The result is expressed in new language rather than lifted. Every claim carries a citation back to the book it came from, by name — so the reader is pointed at the author, and the honest end state is that they go buy the book we told them was worth reading.
One approach makes the books unnecessary. The other makes them findable, and says who wrote them. That distinction was a design constraint from the start, not a compliance review afterward, and it shaped the pipeline: the question set, the schema, the citation trail, the refusal to render source text verbatim.
And the ones that just bite
A confident number can be measuring something other than its name. The patent result had a second lesson in it: the score that reported zero claims supported was, read correctly, reporting the limits of our own collection while wearing a name that implied a verdict on the guide. Any system that grades itself will eventually produce a number like that. The useful question is never the number — it's what's the denominator, and who declared it.
The trust layer is the part with no user complaining. The evidence panel — the citations, the honest gaps, the thing that makes a grounded guide different from generated text — silently stopped rendering for a stretch. A field carried through the pipeline correctly, typed correctly, drawn by nothing. No error, no alarm, pages that looked completely fine. Nobody files a bug about receipts they never knew they were owed. Prose failures announce themselves; evidence failures go quiet, and quiet failures survive.
Identity is harder than it sounds and it's load-bearing. Is this the same book? is genuinely difficult at scale — editions, translations, subtitles, a title differing by one article — and every downstream claim depends on getting it right. Attach a finding to the wrong work and you've produced a citation that is confidently, invisibly false. Which is precisely the disease we set out to treat.
None of this is exotic. It's the specific, boring shape that the general problem — how do you know the confident answer is right? — takes when you try to build something honest instead of something impressive. We've solved some of it, we're actively working the rest, and the workflow that addresses it is the real product underneath the guides.
The knowledge was already there. It was there the whole time, in books nobody has time for and journals nobody outside the academy reads, and the distance between what is known and what gets used has been the quiet tax on professional work for as long as there has been professional work.
What's changed is that the distance is finally closable. Whether it actually closes turns on something duller than the technology — whether the thing handing you an answer will also show you what it rests on, and say plainly which parts it can't support. That is slower to build, more expensive to run, and much harder to demo than a system that simply answers everything.
Anyway. That's the one I'd trust with a decision that mattered, so that's the one we're building.