kestrel.markets
Version:
A typed, token-efficient language + runtime for agentic trading: agents author bounded plans, the runtime fires them at the tick. CLI + typed library + MCP server.
247 lines (214 loc) • 18.9 kB
Markdown
# Percept-reading comprehension is a first-class, measured property at every layer
**Status:** Accepted (2026-07-15 — reviewed, amended (A1-A4), and accepted by the benchmark/
orchestrator session on owner relay; drafted in the owner percept hill-climb session; A5/A6 amended
2026-07-17 per kestrel-wa0j.55). Critical-discovery ADR. Elevates a grade axis that ADR-0009 already
named in passing ("comprehension probes") to a first-class objective and a gate. Extends
ADR-0009 (**the screen is measured, not designed** — renderings are graded hypotheses),
ADR-0041 (**percept inflection + template-as-hypothesis**), and ADR-0043 (**coordinate
ascent over five knob families** — this ADR sets the objective the F2/F4 knobs climb toward
and the F0 instrument that scores them). Sibling in method to ADR-0042 (**season validity**):
where 0042 designs the abstention trap out of the *judgment* benchmark, this ADR establishes
a *comprehension* instrument that has **no abstention trap by construction** (a factual quiz
is right-or-wrong; there is no "stood down" answer to reward). Follows the gate pattern of
ADR-0040 (a session that cannot measure a property can never ground a claim about it).
Companion beads: `kestrel-4gl.14` (comprehension gate + frozen readability-probe — JudgeCell,
Oracle-passes/Blind-fails teeth, versioned frozen QA probe, Grade-inadmissible-until-clear),
`kestrel-wa0j.17` (in-tree comprehension-gate runner), `kestrel-wa0j.21` (percept
property-test harness), `kestrel-wa0j.27` (tokenizer arm), `kestrel-wa0j.28` (interface-effect
de-confound).
## Context
**The founding thesis, verbatim.** Percept-reading ability is the entire reason kestrel was
created. The owner's account: *"I was annoyed that Claude Code couldn't read or understand the
IBKR or Robinhood JSON MCP data — and it took 10 tool calls to get all of the data it should
have in a single view."* The percept **is** the product. Its whole claim is that one dense,
model-readable view replaces N tool calls of unreadable JSON — that a model presented with the
percept *understands the situation it is in* where the same model handed raw broker JSON does
not. Everything downstream (judgment, training, P&L) presupposes that the model can **read** the
screen. That presupposition has, until now, been assumed rather than measured.
**The discovery chain that forced this ADR.** Four committed results converged on the same gap —
comprehension is a distinct, prior property from judgment, and we were not measuring it:
1. **Frame-blindness is real and it is fatal to training** (`inner-loop-1c`, doc
`inner-loop-1c-frameshuffle.md`; predecessors `inner-loop-1`, `inner-loop-1b`). A pre-registered
frame-shuffle control trained the 0.5B on *correctly paired* percepts vs percepts *deliberately
mispaired within oracle-class* (class prior held exactly; only the percept→advantage binding
destroyed), from **bit-identical LoRA inits** on an identical schedule. Result: **NORMAL ≈
SHUFFLED on every VAL metric** (max probability-metric delta 0.034 ≈ 3 wakes/100). The kill shot:
**TRAIN end-loss separates hard — 1.51 (NORMAL) vs 2.95 (SHUFFLED)** — so the frames provably
carry fittable signal and the shuffle was provably disruptive, yet **zero of it transferred to
the policy.** A model that cannot read the percept cannot be trained to judgment on it: the
gradient teaches the class prior and one serialization per class, and the percept is unread. The
0.5B curriculum lane was closed on this evidence — *capability (reading) is the binding
constraint, not data or loss shape.*
2. **Even frontier results may be measuring render fit, not model capability** (the benchmark
session's render-handicap / "Fable-tie" hypothesis, PM-1b — in flight, not yet a LEDGER row).
The renderer was authored by and for one model family; a leaderboard that ranks models on that
renderer partly ranks *how well each tokenizer happens to fit the render the author's model
liked.* Comprehension and render are confounded in the very number we use to pick models. This
is exactly ADR-0043's F4 warning (canonical rendering must stay the control while
tokenizer-optimal renders compete as arms) — but it also means we need a comprehension instrument
*independent of downstream P&L* to separate the two.
3. **Panes are not additive; more render is not more comprehension** (`knob-sweep-s1w2a`, doc
`wave2-2026-07-14/wave2a-analysis.md`; the pane A/B intervention later NULL at n=8 in
`wave2b-orb-1`). Non-additivity was the wave-2a discovery: for both candidate models `both` panes
scored **below** `max(single pane)` on the discretion arm (qwen +1666/+1600 → +814; gemma +556 →
−1198), and the bundled arm *suppressed* gemma's own default-pane de-arm (+2627 → 0). Wave-1's
bundled "trap arm" had been measuring pane **overload**, not pane **information**. Adding render
can *reduce* comprehension — so "how much does the model actually understand from this screen"
must be measured directly, not assumed monotone in render.
4. **The direct instrument exists and is cheap** (the percept-comprehension quiz lane — in flight;
`kestrel-4gl.14` / `kestrel-wa0j.17`). ~6–8 auto-generated factual Q/A per percept, derived from
**kernel ground truth** (the typed Frame the renderer serialized), asked against the rendered
screen. Zero labeling cost (the kernel already holds every answer), verifiable reward
(right-or-wrong), and — unlike the judgment benchmark — **no abstention trap** (there is no
passive "stand down" answer that a degenerate constant policy can farm; cf. ADR-0042). This is
the missing measurement: it scores *reading* directly, before and independent of *judging*.
The chain closes a decomposition the program had been eliding: **reading** (can the model recover
the facts the percept encodes?) is a distinct, testable, and *prior* property to **judging** (given
the facts, does it act well?). We had a rich instrument for judgment (ADR-0042 matched sets) and
none for reading.
## Decision
**Percept-reading comprehension is a first-class, measured property at every layer. It is scored
directly by an auto-generated factual quiz over kernel ground truth, and it gates renderer
selection, model selection, and training curricula before any judgment claim is admitted.**
### 1. Comprehension-per-token is the renderer's primary grade axis (joining, not yet replacing, P&L — see Open-2)
Every tokenizer/model-specific renderer variant (ADR-0043's F2 encoding and F4 reader/tokenizer
families) is scored by **quiz accuracy per token** — machine comprehension per unit of budget —
*not* by downstream P&L alone. P&L is noisy, slow, entangled with judgment, and available only on
the acting pole; the quiz is dense, fast, and available on every percept. The renderer tournament
of ADR-0009 gains a primary, taste-free grade axis: a render that costs fewer tokens for the same
recovered facts, or recovers more facts at the same budget, wins on comprehension-per-token. This
makes ADR-0009's already-named "comprehension probes" the objective the screen is optimized *for*,
and it directly measures the pane-overload finding (§Context.3): a bundled pane that lowers quiz
accuracy at higher token cost is now *scored down*, not argued about.
### 2. Model selection gates on comprehension BEFORE judgment
A base model that cannot read the percept is **disqualified from judgment claims and from training
spend** — the reading/judging decomposition made operational. Concretely: a candidate base must
clear a comprehension threshold on the canonical-render quiz before its judgment numbers are
admissible or before it earns curriculum/RL budget. This is the ADR-0040 pattern applied to a new
property: as a latency-blind session can never ground a latency claim, a *comprehension-blind* base
can never ground a *judgment* claim. `inner-loop-1c` is the founding precedent — the 0.5B was
frame-blind, so no data or loss fix could matter, and spend on it was correctly halted. The gate
makes that reasoning a standing rule rather than a post-mortem, and it fires *before* dollars, not
after.
### 3. Training curricula get a stage-0: comprehension SFT/RL
Curricula gain a **stage-0** — comprehension SFT/RL on the auto-generated Q/A — run *before or
alongside* judgment training. Its reward is verifiable and right-or-wrong, so it carries **none of
the degeneracy** that has defeated the judgment objective: no flat-advantage collapse
(`inner-loop-1`), no null-policy-wins constant (Phase-0), no class-prior memorization
(`inner-loop-1b`). Stage-0 both *teaches* reading where the base is weak-but-not-blind and
*verifies*, at the roster scale where the frame-shuffle control is re-run as the first gate
(`inner-loop-1c` §What-this-implies), that the model reads the percept at all before any
GRPO/search-and-distill spend is committed to judgment.
### 4. The product claim, now measurable: comprehension per token per dollar
The percept's value is **machine comprehension per token per dollar**, and it is now a measurable
quantity rather than a founding intuition. The founding annoyance — N tool calls of unreadable
broker JSON — becomes a **benchmarkable delta**:
> **The Reading Delta (canonical demo experiment).** Same model, same kernel state, same quiz, at
> **equal token budgets**: the kestrel percept vs the raw broker-JSON MCP dump (the IBKR/Robinhood
> payloads that motivated the project). Comprehension accuracy of (percept) minus (raw JSON) at
> matched tokens is the founding thesis expressed as a number. It is the demo, the marketing claim,
> and a standing regression — the "10 tool calls → one view" story with a measured y-axis.
## Consequences — what changes at each layer
- **Renderer.** The tournament (ADR-0009) adds comprehension-per-token as a primary grade axis;
F2/F4 variants (ADR-0043) are scored on it directly — F2/F4 *encoding* arms are evaluated for
whether they *raise recovered facts per token* (encoding varies, information held fixed), killing
the pane-overload confound (§Context.3) with a number instead of an argument. Pane-SET *membership*
is F1 (ADR-0043) and is admitted on ADR-0041 §2 both-poles economics, never on
recovered-facts-per-token alone — the quiz is at most a pre-screen there (A5). Canonical render
stays the leaderboard control (ADR-0043 F4); the quiz is what separates render-fit from capability
that PM-1b flagged (§Context.2).
- **Bench.** A comprehension instrument sits *beside* the judgment instrument (ADR-0042), scoring a
distinct property. It is an **F0 measurement concern** in ADR-0043's taxonomy — it is *audited,
not climbed for advantage*: the quiz generator, its ground-truth binding, and its threshold are
versioned, and a change to them mints a new comprehension regime.
- **Training.** Stage-0 comprehension SFT/RL enters every curriculum; the roster-scale
frame-shuffle re-run becomes the go/no-go before judgment spend; comprehension-blind bases are
cut before budget, not after.
- **Product.** The Reading Delta is the canonical demo and a standing regression. The founding
thesis is now something we can put a number on and defend.
## Falsifiability — what would overturn this ADR
This ADR is wrong, and comprehension should be demoted, if any of the following holds on real data:
1. **Quiz accuracy is uncorrelated with judgment** across models and renders — i.e. models that
read the percept better do *not* judge better once capability is otherwise controlled. Then the
quiz measures a property that does not matter, and gating on it is pure cost. (The frame-shuffle
result is evidence *for* correlation — reading was necessary for any transfer — but correlation
must be shown at roster scale, not assumed.)
2. **Raw JSON comprehends equally well at equal token budget** — the Reading Delta is ≈ 0 or
negative on the canonical demo. Then the percept's *reading* advantage is illusory and its value
must be re-argued on other grounds (density, latency, cost) — the founding thesis itself would be
falsified as stated.
3. **The quiz is gameable** — high quiz accuracy is achievable by surface pattern-matching or
answer-leakage without genuine situation comprehension (the Blind-model foil passes). Then the
instrument is invalid until hardened, and no gate may rest on it.
## The quiz is itself an instrument — it obeys the honest-measurement invariants
A gate is only as trustworthy as the instrument under it. The comprehension quiz is subject to the
program's measurement law, not exempt from it:
- **Through-the-real-driver.** Q/A are generated from the **real kernel Frame** and asked against
the **real rendered screen** through the actual render path — never a mock or a hand-authored
probe that could drift from what ships (the rule proven in `SCAN-KILLTEST-DESIGN.md`; the
frame-shuffle control is the same discipline — bit-identical inits, provably-live shuffle).
- **Guards need failing fixtures** (`kestrel-4gl.14`'s Oracle-passes/Blind-fails teeth). The quiz
ships with fixtures that make it **fire red on purpose**: an **Oracle** foil (a probe that can
read the ground truth) MUST pass, and a **Blind** foil (a model or render denied the answering
facts) MUST fail. A quiz on which the Blind foil scores well is answer-leaking and does not ship.
Anti-answer-leak (the render must not echo the quiz's answers verbatim) is a first-class fixture,
not a review note.
- **Inadmissible until the probe clears** (`kestrel-4gl.14`). Comprehension is a JudgeCell with **no
P&L and never blended into the P&L grade** (touches ADR-0006). A Grade is inadmissible until the
frozen readability-probe clears — reading is a *precondition* of a scored judgment, kept on its
own axis so it can never be laundered into or out of the money number.
## Acceptance amendments (2026-07-15, benchmark/orchestrator session)
- **A1 — Forward-only transition.** The gates in the Decision bind new measurements, seasons,
renderer selections, and training spend from acceptance forward. Already-published leaderboard
rows remain valid under the pre-0044 regime and are labeled with it; nothing shipped is
retroactively invalidated. Every new row/season stamps its **comprehension-regime id**
(quiz-generator version + threshold), consistent with the F0 versioning rule above.
- **A2 — Refusals score wrong.** A non-answer, refusal, or "cannot determine" on a quiz item scores
as incorrect. The quiz admits no abstention channel — otherwise the no-abstention-trap property
claimed in the header silently re-opens through the back door.
- **A3 — Comprehension is a PUBLISHABLE benchmark tier (owner, same day).** Reading is not only an
internal gate — it is publishable as its own leaderboard axis: per-model x renderer-arm
**comprehension-per-token** rows (tokenizer-optimized arms included; canonical render as the
control column), labeled READING as a distinct property and **never blended into or presented as
confusable with the alpha/judgment rows** (per the no-P&L JudgeCell rule above). Rows are
admissible only from the hardened instrument (Oracle-passes/Blind-fails teeth clear;
anti-answer-leak fixture green). The **Reading Delta** (Decision 4) is the flagship row of that
tier. It is also contamination-favorable: items auto-generate from kernel ground truth, so the
bank is effectively infinite and refreshable.
- **A4 — Citation status.** The founding evidence docs (`inner-loop-1c-frameshuffle.md`,
`wave2-2026-07-14/wave2a-analysis.md`, `SCAN-KILLTEST-DESIGN.md`) currently live in the
watcher-pareto session's worktree (`kestrel-wt/watcher-pareto/docs/research/watcher-pareto/`) and
land on main with that session's next merge; until then, this ADR is their citation of record and
that merge is owed.
- **A5 — Quiz authority stops at F1's border (2026-07-17, owner; kestrel-wa0j.55).** The
comprehension quiz is the OBJECTIVE for F2/F4 (encoding varies, information held fixed — the render
that recovers more facts per token wins) and at most a cheap PRE-SCREEN for F1 (an unreadable pane
is dead on arrival). F1 pane-SET admission remains ADR-0041 §2 both-poles matched-set ECONOMICS,
always: a pane can raise recovered-facts-per-token and still net negative EV — the passivity trap
is invisible to a comprehension quiz **by construction** (the range-velocity strategist read the
pane at ~100% comprehension and forfeited the action pole). Decisive evidence (2026-07-15 RL
reading-x-judging matrix): gemma reads 97% of the percept and judges ZERO, and the 30B judges
*better* than it reads — necessary-not-sufficient is now empirical, so a comprehension score cannot
be an admission/selection grade for anything downstream of reading. This scopes the header's own
rule and controls where the §Consequences renderer bullet earlier read as F1 admission (now fixed).
ADR-0048 names the identical diagnostic-only discipline for embedding geometry ("the A5 pattern").
- **A6 — Train/gate split (2026-07-17, owner; kestrel-wa0j.55).** The model-selection gate quiz
(§Decision.2) is a HELD-OUT generator version / disjoint question-template family from anything
stage-0 trained on (§Decision.3), with the split recorded in the **comprehension-regime id**. A1's
regime id stamps the generator version but does not by itself split train from gate; without the
split the gate can measure memorization of the generator's question templates (the Blind foil tests
render leakage, not train/gate contamination). The s6ng RECALL/DERIVE/PATTERN tier ladder supplies
the natural held-out structure (disjoint template families / tiers from the gate bank).
## Open questions
1. The comprehension threshold for the model-selection gate (§Decision.2) is TBD — set it from the
roster-scale frame-shuffle re-run (the first empirical calibration point) rather than by fiat.
2. Does comprehension-per-token replace, or merely join, P&L as the renderer's primary objective?
Leaning *join* until §Falsifiability.1 (quiz↔judgment correlation) is measured at roster scale;
if the correlation is tight, the cheaper quiz could largely stand in for the expensive P&L in the
F2/F4 climb.
3. Should the Reading Delta (percept vs raw JSON) become an L-gate in the platform's
`docs/LAUNCH-GATES.md`, beside the no-latency-blind (ADR-0040) and season-validity (ADR-0042)
gates — i.e. "no launch claims percept value without a measured, positive Reading Delta"?
4. Per-tokenizer comprehension floors: does each supported model family get its own quiz-accuracy
floor on the canonical render, or one absolute floor? (Interacts with ADR-0043 F4 — the tokenizer
arm.)