All True Fact

Rubric log

Every change to a prompt or to the arithmetic, with the version it produced, why, and what it did to the fixture. Newest at the top of each section. The brief's rule: if the number does not match intuition, the rubric is wrong. This file is where that gets fixed in the open.

Versions live in code: RUBRIC_VERSION in atf/rubric.py, *_VERSION strings in atf/judge.py, atf/retrieve.py, atf/decompose.py. Every dataset records all of them and the parameter snapshot, so any published number can be reproduced from its own JSON.

Human gate (Phase 3, task 4)

Status: closed 2026-09-17. Curtis picked pair 3, read the plain-language summary, asked for the one-sided-pool fix (rubric 1.2.0 below) and said to close Phase 3 once it was in. It is in; both claims and the fixture are re-scored under it. Curtis's standing direction for Phase 4, in his words: find out "if things deviate and if they do, if it really makes sense why they do", on post-cutoff claims where we can get them and on the rubric-versus-raw audit regardless.

Curtis picked pair 3 on 2026-09-17. Both claims ran on the public profile with the defaults (DuckDuckGo pools; Google's quota was gone for the day). Reports:

Results (2026-09-17)

A: the surge reduced violence B: the Recovery Act reduced unemployment
pool 35 items (3 PDFs), 60 URLs 42 items (14 PDFs), 67 URLs
consensus Supported, p 0.909 (L +2.30) Supported, p 0.873 (L +1.93)
per model Haiku 0.93, GPT mini 0.91, Gemini 0.90, DeepSeek Flash 0.81 (leans), Grok 0.87 Haiku 0.90, GPT mini 0.84 (leans), Gemini 0.97, DeepSeek Flash 0.64 (unclear), Grok 0.84 (leans)
outcome leaf median L +4.94, 22 items in 13 clusters median L +3.02, 25 items in 11 clusters
cause leaf median L +2.38, 14 items in 13 clusters median L +2.38, 19 items in 16 clusters
mean pairwise disagreement 1.22 1.36
cost $0.70 $1.10 (fresh retrieval and roots)

Asymmetry (A − B), root log-odds: consensus +0.37; per model +0.44, +0.72, −1.21, +0.87, +0.22; mean +0.21. Both verdicts are Supported. The models split on which side they favour (Gemini leans B, DeepSeek Flash leans A), and the consensus difference is a third of a log-odds unit. On this pair, the rubric does not show a systematic lean.

Rubric versus raw opinion (added 2026-09-17, baseline_v1)

Each model was also asked cold, no evidence, for its probability. Root p:

model fixture raw → rubric surge raw → rubric stimulus raw → rubric
Haiku 4.5 0.05 → 0.056 0.45 → 0.930 0.72 → 0.895
GPT-5.4 mini 0.08 → 0.143 0.95 → 0.914 0.42 → 0.839
Gemini 3.5 Flash Lite 0.00 → 0.001 0.75 → 0.896 0.85 → 0.966
DeepSeek V4.1 Flash 0.01 → 0.011 0.75 → 0.806 0.80 → 0.636
Grok 4.3 0.05 → 0.015 0.68 → 0.869 0.68 → 0.842
median 0.05 → 0.015 0.75 → 0.909 (+1.2 log-odds) 0.72 → 0.873 (+1.0)

Three things this says:

  1. On the fixture the rubric and the raw opinion agree, which is the leakage caveat in one row: every model already knows Iraq had no WMD.
  2. On both live claims the rubric pushed the number up by about one log-odds, and every raw rationale hedges the causal leaf ("contested", "debated", "identification remains debated") harder than the rubric's 0.92 on that leaf. The models cold are more sceptical of causation than the models plus the pool. Given that the "against" queries produced thin, low-basis documents, the raw opinion may be the better-calibrated one here, and the rubric's job is to earn its number from evidence, not to be more confident than the model. This is the central question for the approval.
  3. The rubric compresses cross-model disagreement. Raw asymmetry on the pair ranged −1.15 to +3.27 log-odds per model (GPT mini cold: surge 0.95, stimulus 0.42); under the rubric −1.21 to +0.87. Mean asymmetry is the same either way (+0.24 raw, +0.21 rubric): no systematic lean cold or scored.

docs/mirror-pair-3.md carries both columns.

What to stare at

  1. Both cause leaves land at 0.92 and the literature is more divided than that. On A the top item is Biddle, Friedman and Shapiro's "Testing the Surge" (quality 0.9, finding, +0.8), whose actual thesis is synergy of surge and Awakening, read as support; the one strong contrary item is a JMSS article (−0.7). The "against" queries produced 14 usable documents, mostly Wikipedia pages scored as assertions at ×0.4 with direction 0. The number is high because contrary retrieval is thin, not because a model leaned. A reasonable human prior for "major cause, beyond the Awakening and ceasefire alone" is 0.7–0.8.
  2. Basis treats macroeconomics as assertion. On B's cause leaf 18 of 19 items are assertion ×0.4: CBO's ARRA estimates, the CEA reports, Blinder–Zandi, the NBER paper. Only an FRBSF working paper was read as a finding. That is the prompt applied consistently (a model-based estimate presents no direct observation), but it means counterfactual claims in economics can only ever be supported by discounted evidence, while a quasi-experimental study in security studies counts in full. Whether that is right is a rubric question for Curtis: options are a fourth basis value for "estimate from a stated model with data" at 0.7, or leaving it and saying so on the methodology page.
  3. Primary-root clustering barely collapses anything on these pools (13 of 14, 16 of 19 on the cause leaves), the opposite of the fixture. The roots really are different sources here, so it may simply be right, but it is the other end of the same dial and worth a look.
  4. DeepSeek V4.1 Flash is the outlier both times (0.81 and 0.64). On B's outcome leaf it read the BLS monthly-payroll pages as direction 0 where the other four read them as support. Cheapest model, most literal reading.
  5. Alias merging worked without drama under roots_canon_v2: on B it collapsed nine spellings of the BLS Current Employment Statistics and five of Blinder–Zandi, at 161 output tokens and no reasoning; on A, six groups at 1,856 reasoning tokens. One cent total against 32 cents burned by v1 on the same pool.

Candidates offered (2026-09-16)

Candidates offered (2026-09-16)

All are 2001–2012 policies with a stated effect that can be checked against the record, settled for 10–20 years, with a politically mirrored twin of the same tree shape. The pair is what makes bias measurable (task 5), so candidates come in pairs.

# claim (A) mirror (B) shape why
1 The 2001 and 2003 Bush tax cuts paid for themselves through higher growth. The 2009 Recovery Act paid for itself through higher growth. AND(growth effect, revenue offset ≥ 100%) Classic stated-versus-actual. CBO, JCT, Treasury 2006 dynamic analysis, CEA reports, academic literature on both. Both proponents made the claim. Strong video.
2 Medicare Part D (2003) cost less than the CBO projected at enactment. The Affordable Care Act's coverage provisions (2010) cost less than the CBO projected at enactment. AND(projection at enactment, actual outlays below it) Both are actually true-ish, which surprises partisans of both sides. Almost pure numbers (CBO, CMS trustees). The cleanest bias probe: same evidence type, opposite ownership.
3 The 2007 Iraq troop surge reduced violence in Iraq. The 2009 Recovery Act reduced unemployment below what it would otherwise have been. AND(outcome moved, policy caused it) Both mainstream-supported with a counterfactual leaf that is genuinely contested. Tests how the rubric handles causal claims.
4 "If you like your plan you can keep it" (2010) held for people with individual-market plans. "No senior will lose their existing drug coverage" (Medicare Part D, 2003) held for retirees with employer drug coverage. single empirical leaf each Narrow, checkable promises. Smaller evidence pools; less of a video.

Recommendation offered, not decided: pair 1 for the video (widest audience, richest evidence, both sides said it), pair 2 as the bias probe if only one pair gets run. Claim files for pairs 1 and 2 are written with hand trees and topics; they have not been run.

Approval

Curtis, 2026-09-17 (by chat, recorded here): the numbers were acceptable except that the system was more confident about the causal leaves than the models cold and than the literature, for a bad reason (a one-sided evidence pool read as a strong verdict). Instruction: "do the one-sided pool fix, then we'll close phase three." Done as rubric 1.2.0. After the fix, on the same judgments:

surge (A) stimulus (B) fixture
consensus before → after 0.909 → 0.851 0.873 → 0.852 0.015 → 0.027
raw opinion (median) 0.75 0.72 0.05
labels 1 Supported, 4 Leans 2 Supported, 2 Leans, 1 Unclear Refuted ×4, Leans refuted ×1
asymmetry A − B, rubric consensus −0.01, per-model mean +0.06
asymmetry A − B, raw per-model mean +0.24, range −1.15 to +3.27

The causal leaves moved toward the models' own hedged view without a retrieval change, because the mixed high-directness sources (CRS, GAO, Fed reports, the Princeton abstract of Biddle et al.) now count as evidence of contest. Every leaf on both claims is flagged one-sided in the report, which is the honest state of those pools and the first item on Phase 4's list.

Open questions for the approval: (a) is 0.92 on the cause leaves acceptable when the same models cold say 0.45–0.95 and every rationale calls causation contested; should the rubric be harder on causal leaves with thin contrary retrieval (more "against" queries, or a mass floor per side); (b) the basis rule on model-based economic estimates, item 2 above; (c) whether to re-run the pair on the union pool once Google is back (about 70 cents a claim, everything else cached).

Defaults

2026-09-17, later — the backtest runs on four models (Curtis)

Curtis, on the balance: "how much do you really need to spend to finish this? I thought we decided to do this on the cheap", then "go with lower end LLMs". So the backtest scores with the four cheapest pinned models (Haiku 4.5, GPT-5.4 mini, Gemini 3.5 Flash Lite, DeepSeek V4.1 Flash) and drops Grok 4.3, which reasons on every call and was about 40 percent of each run. Locked parameters unchanged; the public profile for videos still has five models. About 45 cents a single-leaf claim. Eight post-cutoff claims run under a $3.00 total cap on 2026-09-17; the set grows when the key has money.

2026-09-17 — locked-2026-09-17 (Phase 4)

LOCKED_DEFAULTS and LOCKED_VERSION now live in atf/params.py; every dataset records params_version (locked-2026-09-17 when the run used exactly the locked set, custom otherwise), so a published verdict says which locked set it used. The values are the 2026-09-16 cheap configuration plus one new retrieval dial:

No scoring dial moved at this version: the tuning grid needs the backtest set run, and on 2026-09-17 the OpenRouter balance ($5.47 on a key shared with other projects) allowed two live claims, not fifty. The set (51 claims), the tuning code and the evaluation are in place; the numbers that would justify moving a dial are not. When the set has run, the procedure is: make backtest-tune, read data/backtest/tuning.md, and move a dial only if the held-out (cross-validated) Brier improves, then bump LOCKED_VERSION and record it here.

2026-09-16 evening — the cheap configuration (Curtis's call)

Arithmetic unchanged (still 1.1.0); three defaults moved so a claim costs a dollar or two instead of nineteen. Curtis: a brand-new channel cannot justify $40 a video.

Not taken yet, the next lever if needed: fold directness into the direction call (both already see the sub-claim; only quality must stay blind), which would take a claim to roughly 60¢.

Fixture on the new defaults (data/runs/iraq-wmd-2003/20260917T022244Z, 55 items, $1.26): all five public models Refuted, consensus 0.015; the frontier reference (DeepSeek V4 Pro alone, 88 items) was 0.012. Per model root p: Haiku 0.056, GPT-5.4 mini 0.143, Gemini 3.5 Flash Lite 0.001, DeepSeek V4.1 Flash 0.011, Grok 4.3 0.015. Widest leaf bio, log-odds spread 6.1. Largest item disagreements: the CIA reading-room landing page (quality 0.0 from four models, 0.9 from GPT mini) and Powell's UN transcript on nuke (direction −0.7 to +0.9 across models).

Arithmetic

1.2.0 — 2026-09-17 (Phase 3, the one-sided-pool fix)

Two changes, arithmetic only; every existing dataset re-scores with zero calls.

  1. Contested evidence pulls toward the prior. An item with direction 0 and directness ≥ mixed_directness_min (0.5) is a source that addresses the leaf and finds it mixed. Its discounted weight is mass_mixed. The leaf's summed log-odds are multiplied by 1 / (1 + contested_pull × mass_mixed / (mass_for + mass_against)), contested_pull default 1.0. Why: on both pair-3 cause leaves, several high-quality reviews that said "contested" sat at direction 0 and contributed nothing, while a thin set of supporting items drove the leaf to 0.92. Asked cold, every model called causation contested. The prompt's direction 0 conflates "mixed" with "does not address"; directness separates them without a prompt change. contested_pull=0 reproduces 1.1.0.
  2. One-sided flag. mass_for, mass_against, mass_mixed are recorded per leaf; a leaf whose weaker side has directional mass below side_mass_min (0.5) is flagged one_sided in the dataset, the leaf line, the verdict line and a call-out under the leaf. Reported, never applied to the number: a settled fact legitimately has nothing against it, but the reader must be able to see that the pool did not test the other side. Every leaf on pair 3 is flagged; on the fixture, chem is flagged for every model.

Effect on the same judgments: see the approval table above. Not done, and first on Phase 4's list: a second retrieval pass for the weaker side when a leaf comes out one-sided, so the flag can clear itself where contrary material exists.

1.1.0 — 2026-09-16 (Phase 3)

Four changes, all visible in the fixture report and all re-scorable on old datasets with zero model calls (atf rescore).

  1. Basis factor. weight = quality^a × directness^b × basis_factor, with basis_weights = (finding 1.0, mixed 0.7, assertion 0.4). Why: the Phase 2 report scored the October 2002 National Intelligence Estimate (quality 0.90, "official primary document") and Powell's UN transcript above the Iraq Survey Group's nuclear volume, and the nuke leaf came out leans supported on the DuckDuckGo pool. The blind quality prompt rates the source, correctly: an NIE is an official record. But its content is an estimate, not an observation, and the rubric had no way to say so. Basis is a separate atomic judgment (returned by direction_v2, see below), so it is auditable per item and the discount is a dial. Unknown basis (old judgments) gets 1.0, never guessed. The 0.4 default is a judgment call recorded here for Curtis to move at the gate.
  2. Root keys ignore type. primary_source: iraq survey group and author_group: iraq survey group were two clusters. Now one. The type is kept for display.
  3. Root aliases merged once per pool (roots_canon_v1, see prompts). cia and central intelligence agency, iraq survey group and duelfer report 2004, were separate clusters. The alias map is stored in the dataset (roots.aliases).
  4. cluster_roots dial: primary (default; union on each item's first, most load-bearing root only) or all (transitive union on every root, the 1.0.0 behaviour). Why: chem already had 19 items collapsed under unmovic by transitivity in Phase 2. Once aliases merged, all put 39 of 40 bio items and 35 of 40 nuke items of the fixture into one cluster, because nearly every article mentions the ISG, Powell or the NIE somewhere; items ranked 8 and below then contribute nothing and the verdict rests on three or four items. primary models "rests on" rather than "mentions" and gave 20–26 clusters per 40 items while still collapsing the articles that relay the Duelfer Report. Default set to primary on 2026-09-16 after the first union-pool run, before anything was published under 1.1.0.
  5. Leave-one-backend-out spread per leaf, measured when the pool came from --search both. Why: Phase 2 measured retrieval luck moving the root verdict by 0.72 between Google and DuckDuckGo pools. The pool is now treated as a source of uncertainty next to the cluster structure.

Reproducing 1.0.0 numbers: --set basis_weights=1,1,1 on a single-backend pool, except where two roots that differed only by type now merge.

Fixture effect (DeepSeek V4 Pro, union pool of 88 items, 40 per leaf, run data/runs/iraq-wmd-2003/20260916T150110Z, 81 cents), root p by dial setting:

basis_weights cluster_roots chem bio nuke root clusters per 40
(1, .7, .4) primary 0.001 0.006 0.005 0.012 refuted 20 / 26 / 24
(1, .7, .4) all 0.063 0.047 0.014 0.119 refuted 4 / 2 / 6
(1, 1, 1) primary 0.000 0.012 0.003 0.015 refuted 20 / 26 / 24
(1, 1, 1) all 0.274 0.364 0.053 0.563 unclear 4 / 2 / 6

The basis factor matters most where clustering is coarse (the NIE and Powell's transcript are the strongest pro items and are both assertions); with fine clustering the sheer count of independent findings already decides it. Under the defaults the nuke leaf that Phase 2 got wrong on the DuckDuckGo pool is now refuted at 0.005. Phase 2's Google pool rescored under 1.1.0 with all (no aliases, no basis available) went from Refuted to Unclear on over-clustering alone; with primary it stayed Refuted at 0.003. Leave-one-backend-out spread on the union pool: 0.04–0.11 per leaf, against the 0.72 root swing Phase 2 saw between the two single-backend pools.

1.0.0 — 2026-09-16 (Phase 2)

Initial arithmetic. See docs/phases/phase-2-engine.md.

Prompts

queries_side_v1 — 2026-09-17 (Phase 4)

The second retrieval pass. Given the leaf, the topic, which side came out thin and the queries already tried, write exactly queries_per_side new queries for that side: "primary documents, official statistics, peer-reviewed or expert analysis, investigative reporting, or the best case made by those who hold that position", and "think about who would have an interest in documenting the [supporting or contradicting] case". Cached like every other call, so a re-run of a finished claim makes no second-pass calls either.

direction_v2 — 2026-09-16

Adds basis to the direction reply (finding / mixed / assertion) with the definition above and the explicit line: "A document can be an official primary record and still be an assertion: an estimate made before the fact is an assertion even if it was later published by a government." Direction and strength wording unchanged. The flipped variant returns basis too; the plain one is used.

roots_canon_v2 — 2026-09-17

v1 sent every distinct root id in one call. On the fixture that was 150 ids and DeepSeek V4 Pro needed 33k reasoning tokens; on the surge claim's pool it burned 16k then 65k tokens ($0.32) without answering and was heading for a 262k-token attempt when killed. v2 bounds it three ways:

The prompt also says "never group across candidate sets; decide quickly; when unsure, do not group."

roots_canon_v1 — 2026-09-16

New. Once per pool, the retrieval model sees the sorted list of distinct root ids (with the types they were seen as) and returns groups that name the same organisation, document, dataset or study. Explicitly not grouped: an agency and a body it oversaw, a funder and the study it funded, two reports by one author, a newspaper and a story it ran. Canonical id in a group = named by the most items, then shortest, then lexical. Validator drops unknown ids, singletons and any id that appears twice. Cost to watch: on the fixture's 150-odd root ids DeepSeek V4 Pro spent 16k hidden reasoning tokens and returned nothing on the first try, then 33k tokens and 25 good groups on the retry at a larger cap (17 cents for both). Once per pool and cached, so tolerable; if it recurs on the hand-run claim, shrink the list by pre-grouping ids that share a token before asking.

quality_v1, directness_v1, roots_v1, queries_v1, decompose_v1

Unchanged since Phase 2.

Retrieval

2026-09-17 (Phase 4): the second pass for one-sided leaves

Rubric 1.2.0 flagged every leaf on pair 3 one-sided. The flag now triggers action: when a majority of the models that scored a leaf flag it, the weaker side (majority of the flagging models' weaker sides; ties to "against") gets queries_per_side more queries (queries_side_v1), the results are fetched and deduplicated against the pool, up to second_pass_items new items per leaf are attached, roots are extracted for the new items (the alias pass re-runs over the whole id set, a cent), every model judges the new (item, leaf) pairs, and the rubric and comparison are recomputed. The dataset keeps the first-pass numbers per model per leaf in second_pass.before and the after-state in second_pass.after, plus cleared, the leaves no longer one-sided by majority. Three outcomes are recorded as such: cleared, new material judged but still one-sided, and no new material found (the first-pass numbers stand untouched). A failure in the pass (search, budget, HTTP 402) is recorded and the first-pass state is kept.

Why once and bounded: the point is to test the other side properly, not to keep searching until the number moves. Cost is at most second_pass_items items per flagged leaf across the five public models, about 30 cents a leaf.

"What moved it" (report section, Phase 4)

Per model, the five items with the largest push (llr times the leaf's contested shrink), with direction, strength, basis, blind quality, directness, cluster rank and source, then per leaf: total push from the prior, the share the listed items account for, and the model's own cold number with the rubric-minus-raw delta. Pure view over the dataset (atf/audit.py). Curtis's test: a reader should be able to say "real study, the move makes sense" or "three Wikipedia pages, it doesn't" from the table alone.

2026-09-16 (Phase 3)

Rendered from docs/RUBRIC_LOG.md in the engine repository. The page and the file are one text.