UNPKG

kestrel.markets

Version:

A typed, token-efficient language + runtime for agentic trading: agents author bounded plans, the runtime fires them at the tick. CLI + typed library + MCP server.

247 lines (214 loc) 18.9 kB
# Percept-reading comprehension is a first-class, measured property at every layer **Status:** Accepted (2026-07-15 — reviewed, amended (A1-A4), and accepted by the benchmark/ orchestrator session on owner relay; drafted in the owner percept hill-climb session; A5/A6 amended 2026-07-17 per kestrel-wa0j.55). Critical-discovery ADR. Elevates a grade axis that ADR-0009 already named in passing ("comprehension probes") to a first-class objective and a gate. Extends ADR-0009 (**the screen is measured, not designed** — renderings are graded hypotheses), ADR-0041 (**percept inflection + template-as-hypothesis**), and ADR-0043 (**coordinate ascent over five knob families** — this ADR sets the objective the F2/F4 knobs climb toward and the F0 instrument that scores them). Sibling in method to ADR-0042 (**season validity**): where 0042 designs the abstention trap out of the *judgment* benchmark, this ADR establishes a *comprehension* instrument that has **no abstention trap by construction** (a factual quiz is right-or-wrong; there is no "stood down" answer to reward). Follows the gate pattern of ADR-0040 (a session that cannot measure a property can never ground a claim about it). Companion beads: `kestrel-4gl.14` (comprehension gate + frozen readability-probe — JudgeCell, Oracle-passes/Blind-fails teeth, versioned frozen QA probe, Grade-inadmissible-until-clear), `kestrel-wa0j.17` (in-tree comprehension-gate runner), `kestrel-wa0j.21` (percept property-test harness), `kestrel-wa0j.27` (tokenizer arm), `kestrel-wa0j.28` (interface-effect de-confound). ## Context **The founding thesis, verbatim.** Percept-reading ability is the entire reason kestrel was created. The owner's account: *"I was annoyed that Claude Code couldn't read or understand the IBKR or Robinhood JSON MCP data — and it took 10 tool calls to get all of the data it should have in a single view."* The percept **is** the product. Its whole claim is that one dense, model-readable view replaces N tool calls of unreadable JSON — that a model presented with the percept *understands the situation it is in* where the same model handed raw broker JSON does not. Everything downstream (judgment, training, P&L) presupposes that the model can **read** the screen. That presupposition has, until now, been assumed rather than measured. **The discovery chain that forced this ADR.** Four committed results converged on the same gap — comprehension is a distinct, prior property from judgment, and we were not measuring it: 1. **Frame-blindness is real and it is fatal to training** (`inner-loop-1c`, doc `inner-loop-1c-frameshuffle.md`; predecessors `inner-loop-1`, `inner-loop-1b`). A pre-registered frame-shuffle control trained the 0.5B on *correctly paired* percepts vs percepts *deliberately mispaired within oracle-class* (class prior held exactly; only the percept→advantage binding destroyed), from **bit-identical LoRA inits** on an identical schedule. Result: **NORMAL ≈ SHUFFLED on every VAL metric** (max probability-metric delta 0.034 ≈ 3 wakes/100). The kill shot: **TRAIN end-loss separates hard — 1.51 (NORMAL) vs 2.95 (SHUFFLED)** — so the frames provably carry fittable signal and the shuffle was provably disruptive, yet **zero of it transferred to the policy.** A model that cannot read the percept cannot be trained to judgment on it: the gradient teaches the class prior and one serialization per class, and the percept is unread. The 0.5B curriculum lane was closed on this evidence — *capability (reading) is the binding constraint, not data or loss shape.* 2. **Even frontier results may be measuring render fit, not model capability** (the benchmark session's render-handicap / "Fable-tie" hypothesis, PM-1b — in flight, not yet a LEDGER row). The renderer was authored by and for one model family; a leaderboard that ranks models on that renderer partly ranks *how well each tokenizer happens to fit the render the author's model liked.* Comprehension and render are confounded in the very number we use to pick models. This is exactly ADR-0043's F4 warning (canonical rendering must stay the control while tokenizer-optimal renders compete as arms) — but it also means we need a comprehension instrument *independent of downstream P&L* to separate the two. 3. **Panes are not additive; more render is not more comprehension** (`knob-sweep-s1w2a`, doc `wave2-2026-07-14/wave2a-analysis.md`; the pane A/B intervention later NULL at n=8 in `wave2b-orb-1`). Non-additivity was the wave-2a discovery: for both candidate models `both` panes scored **below** `max(single pane)` on the discretion arm (qwen +1666/+1600 → +814; gemma +556 → −1198), and the bundled arm *suppressed* gemma's own default-pane de-arm (+2627 → 0). Wave-1's bundled "trap arm" had been measuring pane **overload**, not pane **information**. Adding render can *reduce* comprehension — so "how much does the model actually understand from this screen" must be measured directly, not assumed monotone in render. 4. **The direct instrument exists and is cheap** (the percept-comprehension quiz lane — in flight; `kestrel-4gl.14` / `kestrel-wa0j.17`). ~6–8 auto-generated factual Q/A per percept, derived from **kernel ground truth** (the typed Frame the renderer serialized), asked against the rendered screen. Zero labeling cost (the kernel already holds every answer), verifiable reward (right-or-wrong), and — unlike the judgment benchmark — **no abstention trap** (there is no passive "stand down" answer that a degenerate constant policy can farm; cf. ADR-0042). This is the missing measurement: it scores *reading* directly, before and independent of *judging*. The chain closes a decomposition the program had been eliding: **reading** (can the model recover the facts the percept encodes?) is a distinct, testable, and *prior* property to **judging** (given the facts, does it act well?). We had a rich instrument for judgment (ADR-0042 matched sets) and none for reading. ## Decision **Percept-reading comprehension is a first-class, measured property at every layer. It is scored directly by an auto-generated factual quiz over kernel ground truth, and it gates renderer selection, model selection, and training curricula before any judgment claim is admitted.** ### 1. Comprehension-per-token is the renderer's primary grade axis (joining, not yet replacing, P&L — see Open-2) Every tokenizer/model-specific renderer variant (ADR-0043's F2 encoding and F4 reader/tokenizer families) is scored by **quiz accuracy per token** — machine comprehension per unit of budget — *not* by downstream P&L alone. P&L is noisy, slow, entangled with judgment, and available only on the acting pole; the quiz is dense, fast, and available on every percept. The renderer tournament of ADR-0009 gains a primary, taste-free grade axis: a render that costs fewer tokens for the same recovered facts, or recovers more facts at the same budget, wins on comprehension-per-token. This makes ADR-0009's already-named "comprehension probes" the objective the screen is optimized *for*, and it directly measures the pane-overload finding (§Context.3): a bundled pane that lowers quiz accuracy at higher token cost is now *scored down*, not argued about. ### 2. Model selection gates on comprehension BEFORE judgment A base model that cannot read the percept is **disqualified from judgment claims and from training spend** — the reading/judging decomposition made operational. Concretely: a candidate base must clear a comprehension threshold on the canonical-render quiz before its judgment numbers are admissible or before it earns curriculum/RL budget. This is the ADR-0040 pattern applied to a new property: as a latency-blind session can never ground a latency claim, a *comprehension-blind* base can never ground a *judgment* claim. `inner-loop-1c` is the founding precedent — the 0.5B was frame-blind, so no data or loss fix could matter, and spend on it was correctly halted. The gate makes that reasoning a standing rule rather than a post-mortem, and it fires *before* dollars, not after. ### 3. Training curricula get a stage-0: comprehension SFT/RL Curricula gain a **stage-0** — comprehension SFT/RL on the auto-generated Q/A — run *before or alongside* judgment training. Its reward is verifiable and right-or-wrong, so it carries **none of the degeneracy** that has defeated the judgment objective: no flat-advantage collapse (`inner-loop-1`), no null-policy-wins constant (Phase-0), no class-prior memorization (`inner-loop-1b`). Stage-0 both *teaches* reading where the base is weak-but-not-blind and *verifies*, at the roster scale where the frame-shuffle control is re-run as the first gate (`inner-loop-1c` §What-this-implies), that the model reads the percept at all before any GRPO/search-and-distill spend is committed to judgment. ### 4. The product claim, now measurable: comprehension per token per dollar The percept's value is **machine comprehension per token per dollar**, and it is now a measurable quantity rather than a founding intuition. The founding annoyance — N tool calls of unreadable broker JSON — becomes a **benchmarkable delta**: > **The Reading Delta (canonical demo experiment).** Same model, same kernel state, same quiz, at > **equal token budgets**: the kestrel percept vs the raw broker-JSON MCP dump (the IBKR/Robinhood > payloads that motivated the project). Comprehension accuracy of (percept) minus (raw JSON) at > matched tokens is the founding thesis expressed as a number. It is the demo, the marketing claim, > and a standing regression — the "10 tool calls → one view" story with a measured y-axis. ## Consequences — what changes at each layer - **Renderer.** The tournament (ADR-0009) adds comprehension-per-token as a primary grade axis; F2/F4 variants (ADR-0043) are scored on it directly — F2/F4 *encoding* arms are evaluated for whether they *raise recovered facts per token* (encoding varies, information held fixed), killing the pane-overload confound (§Context.3) with a number instead of an argument. Pane-SET *membership* is F1 (ADR-0043) and is admitted on ADR-0041 §2 both-poles economics, never on recovered-facts-per-token alone — the quiz is at most a pre-screen there (A5). Canonical render stays the leaderboard control (ADR-0043 F4); the quiz is what separates render-fit from capability that PM-1b flagged (§Context.2). - **Bench.** A comprehension instrument sits *beside* the judgment instrument (ADR-0042), scoring a distinct property. It is an **F0 measurement concern** in ADR-0043's taxonomy — it is *audited, not climbed for advantage*: the quiz generator, its ground-truth binding, and its threshold are versioned, and a change to them mints a new comprehension regime. - **Training.** Stage-0 comprehension SFT/RL enters every curriculum; the roster-scale frame-shuffle re-run becomes the go/no-go before judgment spend; comprehension-blind bases are cut before budget, not after. - **Product.** The Reading Delta is the canonical demo and a standing regression. The founding thesis is now something we can put a number on and defend. ## Falsifiability — what would overturn this ADR This ADR is wrong, and comprehension should be demoted, if any of the following holds on real data: 1. **Quiz accuracy is uncorrelated with judgment** across models and renders — i.e. models that read the percept better do *not* judge better once capability is otherwise controlled. Then the quiz measures a property that does not matter, and gating on it is pure cost. (The frame-shuffle result is evidence *for* correlation — reading was necessary for any transfer — but correlation must be shown at roster scale, not assumed.) 2. **Raw JSON comprehends equally well at equal token budget** — the Reading Delta is ≈ 0 or negative on the canonical demo. Then the percept's *reading* advantage is illusory and its value must be re-argued on other grounds (density, latency, cost) — the founding thesis itself would be falsified as stated. 3. **The quiz is gameable** — high quiz accuracy is achievable by surface pattern-matching or answer-leakage without genuine situation comprehension (the Blind-model foil passes). Then the instrument is invalid until hardened, and no gate may rest on it. ## The quiz is itself an instrument — it obeys the honest-measurement invariants A gate is only as trustworthy as the instrument under it. The comprehension quiz is subject to the program's measurement law, not exempt from it: - **Through-the-real-driver.** Q/A are generated from the **real kernel Frame** and asked against the **real rendered screen** through the actual render path — never a mock or a hand-authored probe that could drift from what ships (the rule proven in `SCAN-KILLTEST-DESIGN.md`; the frame-shuffle control is the same discipline — bit-identical inits, provably-live shuffle). - **Guards need failing fixtures** (`kestrel-4gl.14`'s Oracle-passes/Blind-fails teeth). The quiz ships with fixtures that make it **fire red on purpose**: an **Oracle** foil (a probe that can read the ground truth) MUST pass, and a **Blind** foil (a model or render denied the answering facts) MUST fail. A quiz on which the Blind foil scores well is answer-leaking and does not ship. Anti-answer-leak (the render must not echo the quiz's answers verbatim) is a first-class fixture, not a review note. - **Inadmissible until the probe clears** (`kestrel-4gl.14`). Comprehension is a JudgeCell with **no P&L and never blended into the P&L grade** (touches ADR-0006). A Grade is inadmissible until the frozen readability-probe clears — reading is a *precondition* of a scored judgment, kept on its own axis so it can never be laundered into or out of the money number. ## Acceptance amendments (2026-07-15, benchmark/orchestrator session) - **A1 — Forward-only transition.** The gates in the Decision bind new measurements, seasons, renderer selections, and training spend from acceptance forward. Already-published leaderboard rows remain valid under the pre-0044 regime and are labeled with it; nothing shipped is retroactively invalidated. Every new row/season stamps its **comprehension-regime id** (quiz-generator version + threshold), consistent with the F0 versioning rule above. - **A2 — Refusals score wrong.** A non-answer, refusal, or "cannot determine" on a quiz item scores as incorrect. The quiz admits no abstention channel — otherwise the no-abstention-trap property claimed in the header silently re-opens through the back door. - **A3 — Comprehension is a PUBLISHABLE benchmark tier (owner, same day).** Reading is not only an internal gate — it is publishable as its own leaderboard axis: per-model x renderer-arm **comprehension-per-token** rows (tokenizer-optimized arms included; canonical render as the control column), labeled READING as a distinct property and **never blended into or presented as confusable with the alpha/judgment rows** (per the no-P&L JudgeCell rule above). Rows are admissible only from the hardened instrument (Oracle-passes/Blind-fails teeth clear; anti-answer-leak fixture green). The **Reading Delta** (Decision 4) is the flagship row of that tier. It is also contamination-favorable: items auto-generate from kernel ground truth, so the bank is effectively infinite and refreshable. - **A4 — Citation status.** The founding evidence docs (`inner-loop-1c-frameshuffle.md`, `wave2-2026-07-14/wave2a-analysis.md`, `SCAN-KILLTEST-DESIGN.md`) currently live in the watcher-pareto session's worktree (`kestrel-wt/watcher-pareto/docs/research/watcher-pareto/`) and land on main with that session's next merge; until then, this ADR is their citation of record and that merge is owed. - **A5 — Quiz authority stops at F1's border (2026-07-17, owner; kestrel-wa0j.55).** The comprehension quiz is the OBJECTIVE for F2/F4 (encoding varies, information held fixed — the render that recovers more facts per token wins) and at most a cheap PRE-SCREEN for F1 (an unreadable pane is dead on arrival). F1 pane-SET admission remains ADR-0041 §2 both-poles matched-set ECONOMICS, always: a pane can raise recovered-facts-per-token and still net negative EV — the passivity trap is invisible to a comprehension quiz **by construction** (the range-velocity strategist read the pane at ~100% comprehension and forfeited the action pole). Decisive evidence (2026-07-15 RL reading-x-judging matrix): gemma reads 97% of the percept and judges ZERO, and the 30B judges *better* than it reads — necessary-not-sufficient is now empirical, so a comprehension score cannot be an admission/selection grade for anything downstream of reading. This scopes the header's own rule and controls where the §Consequences renderer bullet earlier read as F1 admission (now fixed). ADR-0048 names the identical diagnostic-only discipline for embedding geometry ("the A5 pattern"). - **A6 — Train/gate split (2026-07-17, owner; kestrel-wa0j.55).** The model-selection gate quiz (§Decision.2) is a HELD-OUT generator version / disjoint question-template family from anything stage-0 trained on (§Decision.3), with the split recorded in the **comprehension-regime id**. A1's regime id stamps the generator version but does not by itself split train from gate; without the split the gate can measure memorization of the generator's question templates (the Blind foil tests render leakage, not train/gate contamination). The s6ng RECALL/DERIVE/PATTERN tier ladder supplies the natural held-out structure (disjoint template families / tiers from the gate bank). ## Open questions 1. The comprehension threshold for the model-selection gate (§Decision.2) is TBD — set it from the roster-scale frame-shuffle re-run (the first empirical calibration point) rather than by fiat. 2. Does comprehension-per-token replace, or merely join, P&L as the renderer's primary objective? Leaning *join* until §Falsifiability.1 (quiz↔judgment correlation) is measured at roster scale; if the correlation is tight, the cheaper quiz could largely stand in for the expensive P&L in the F2/F4 climb. 3. Should the Reading Delta (percept vs raw JSON) become an L-gate in the platform's `docs/LAUNCH-GATES.md`, beside the no-latency-blind (ADR-0040) and season-validity (ADR-0042) gates — i.e. "no launch claims percept value without a measured, positive Reading Delta"? 4. Per-tokenizer comprehension floors: does each supported model family get its own quiz-accuracy floor on the canonical render, or one absolute floor? (Interacts with ADR-0043 F4 — the tokenizer arm.)