Platform · Shipped
Known-answer validity controls for judgment pipelines
Positive and negative controls with known answers, run against the DETERMINISTIC stages of a judgment pipeline (retrieval, shortlisting, filtering). Reliability statistics answer 'do the raters agree'; they cannot answer 'was the pipeline capable of finding the right answer'. Free, repeatable, every build.
- Cluster
- Platform
- Type
- gate
- Status
- Shipped
- Used by
- people-analytics-toolbox, peopleanalyst-site, canonicai, devplane
Maturity evidence
- Functioning
- Proven against a pre-fix fixture: exit 1 with 4 failures on the pre-fix artifact, exit 0 on the corrected one. A gate never shown to fail is not evidence. EXTENDED 2026-07-22 (MEJ-11a): known-answer retrieval controls on the wage-compliance ordinance extractor — the deterministic retrieval heuristics extracted to a dependency-free module the CI gate executes directly (one implementation, no copy), 3 hand-verified fixtures (FLSA §206 + real ordinance page positives; SPA nav-shell negative = the pre-fix failure class), an askability RetrievalRecord on every extractor output (never-shown vs shown-and-silent durably distinguishable), and check:ordinance-controls proven to exit 1 on a broken manifest in CI. EXTENDED 2026-07-22 (MEJ-11d): known-answer controls on the JFM canon factory's corpus-retrieval stage (scripts/jfm-harness/lib/grounding.mjs — discipline routing + apex-IC grounding lookup, the exact module factory.mjs imports): 4 hand-verified controls against the open synthesis data (data-science and software-routing positives whose discriminator lines must reach the generation prompt; sub-P5-envelope and Management-track known-absent negatives), a per-unit grounding askability record stamped into every factory output (ran-and-found vs ran-and-corpus-silent vs never-attempted durably distinguishable — MEJ-17 says this provenance cannot be backfilled), and check:jfm-retrieval-controls proven to exit 1 on a broken manifest in CI. EXTENDED 2026-07-22 (MEJ-11c): known-answer controls on the PAT-176 metric↔strategy value-mapping (anycomp) — the deterministic mapping stage (rules + signals) extracted to dependency-free value-mapping-heuristics.mjs the gate executes directly (one implementation), 6 hand-verified controls against the real catalog data (pay-equity-gap 0.08 = 4x the analyses-catalog materiality line → negative equity signal; deliberately-unmapped midpoint-progression → no-rule, never fabricated), an askability ValueMappingRecord on every metric-strategy response (0 candidates = never-asked, 1 = leading-question — an empty posture or a score-0 recommendation sheet under never-asked is a vocabulary miss, not a finding), and check:value-mapping-controls proven to exit 1 on a broken manifest in CI. EXTENDED 2026-07-22 (MEJ-11e): known-answer retrieval controls on the niche-discover campaign — the gate executes the pipeline's REAL deterministic stage (scripts/jfm-harness/lib/promotion.mjs, imported never copied) over hand-verified fixtures ('perception' traced to the committed robotics sweep must promote; 'glassblowing' verified absent from all 44 committed manifests must not), a three-way SweepRetrievalRecord verdict (searched-and-found / searched-and-none / retrieval-not-proven — 'we did not look' can never publish as 'no niches found'), an askability sweep proving every frontier vertical marked covered carries retrieval evidence (44/44), and check:niche-controls proven to exit 1 on a broken manifest in CI. EXTENDED 2026-07-22 (MEJ-11b): same treatment on the jurisdiction-discovery pipeline — deterministic stages (curated-source selection, scoreDomain trust ladder, wage-content markers) extracted to dependency-free discovery-heuristics.mjs (marker heuristics imported from the MEJ-11a module, never copied), 7 hand-verified controls (CA known-present / TX known-absent source-selection probes; 3 trust-ladder rungs; live-fetched CA DLSE positive + SPA nav-shell negative), per-source DiscoveryRetrievalRecord + per-scan DiscoverySourceSelection askability records on scan summaries, and check:jurisdiction-controls proven to exit 1 on a broken manifest in CI.
- Valuable
- Catches the exact class of failure that produced a wrong published headline: a flawless ensemble judging an empty candidate set.
- Understood
- docs/CAPABILITY-MANIFEST.md#judgment-validity-controls; module header states the reliability/validity distinction.
- Integrated
- MEJ-11 SERIES COMPLETE 2026-07-22: known-answer retrieval controls + askability records wired on ALL SIX judgment pipelines — canon-occupation pairing, ordinance extractor (11a), jurisdiction discovery (11b), metric↔strategy value-mapping (11c), JFM factory retrieval (11d), niche-discover (11e) — each with a CI gate proven to exit 1 on a broken manifest.
How to use it
Not surfaced with a callable endpoint yet — still building. Tell us you want Known-answer validity controls for judgment pipelinesand we’ll prioritize it.
Request Known-answer validity controls for judgment pipelinesOther Platform capabilities
- Comp audit self-serve — CSV intake, free validate, commerce-gated watermarked run (HO-155)
- Gift pass — gifted ownership on the entitlement store (magic-link claim)
- JobFrame canonical Job Family × Focus × Level taxonomy + classifier
- SMB pay-system wizard — chocolate/vanilla over the pay-system catalog
- Support channel — portfolio-wide /support form → ticket store → operator email
- Ask-your-data — generalizable NL→rank→tier→cite spoke (HO-378)
- Diagnostic acquisition — next-best-ask over a live hypothesis set
- Executive-Priorities Elicitation Instruments (static catalog)
Capability detail sourced from the People Analytics Toolbox capability feed (source of truth), snapshot retrieved 2026-07-23.