kestrel.markets
Version:
A typed, token-efficient language + runtime for agentic trading: agents author bounded plans, the runtime fires them at the tick. CLI + typed library + MCP server.
63 lines (37 loc) • 9.04 kB
Markdown
# The percept climb is coordinate ascent over five knob families — the search space is a fixed coordinate system, and evidence never pools across coordinates
**Status:** Accepted (owner, 2026-07-16; drafted in the owner percept hill-climb session, 2026-07-15). Dependent Train beads `kestrel-wa0j.24`/`.26`/`.27`/`.28` are merged and the owner ratified acceptance, closing the Accepted-depends-on-Proposed inversion. Extends ADR-0009 (**the screen is measured, not designed** — the tournament mechanics and the frozen invariants), ADR-0029 (the viewshop loop and its budget knobs), and ADR-0041 (**percept inflection + template-as-hypothesis** — the cell key, the admission rules, the channel invariants). Companion beads: `kestrel-wa0j.24` (inflection substrate), `.26` (founder library), `.27` (tokenizer arm), `.28` (2×2 re-measure), `.33` (capacity curves), `.38` (pane-order arm), `.49`–`.53` (pod-level percepts).
## Context
The percept hill-climb ("the Bloomberg terminal for agents") is beginning in earnest, and its search space is documented only in fragments: ADR-0009 carries the tournament mechanics and the invariants renderings compete inside; `docs/rendering-variants.md` enumerates encoding dimensions only; ADR-0041 carries the evidence schema; the delivery, reader, and budget dimensions exist as individual beads with no document stating that they are coordinates of one space. Two failure modes follow from an unwritten taxonomy. First, an experiment moves two coordinates at once (a new pane AND a new cache policy) and the resulting grade is attributed to whichever coordinate the experimenter favored — the merged strategist hill-climb was nearly misread for exactly this reason (the interface-effect confound `wa0j.28` exists to remove). Second, a climber re-litigates a frozen invariant as if it were a knob, or treats a knob as frozen and never searches it.
## Decision
### 1. The rails are not knobs
The following never enter the climb; they are the surface variants compete ON, not dimensions they compete OVER (ADR-0009's platform-scope test, ADR-0041 §1): the kernel leads every frame; honesty markers come from the closed codebook; streaming is append-only/cache-monotone; the renderer invents no value; the legibility floor holds; safety-critical facts render exactly twice, byte-identical; `print(parse(text))` is byte-stable. A proposal that varies a rail is refused as ill-posed, not graded.
### 2. Five knob families, plus the meta-knob
Every climbable dimension belongs to exactly one family:
- **F1 — Selection & addressing (WHAT):** pane membership per cell; pane internals via inflection args (chain `Count`/`ExpiryOrdinal`, tape bucket/window/anchor, band-keyed level sets, session ordinals); pane order (`wa0j.38`). Highest leverage; where the passivity trap operates; every step graded both-poles on economics (ADR-0041 §2).
- **F2 — Encoding (HOW):** within-row tape encoding, format (`ascii|unicode|md`), volume encoding, anchor cadence, precision conventions. The `docs/rendering-variants.md` pool is F2's registry.
- **F3 — Delivery (WHEN):** keyframe vs delta; cache policy (conversation-cached vs stateless-redraw); coalescing/squelch; zoom-on-wake; phase conditioning (OPEN/WAKE/SHOCK/CLOSE templates, SHOCK dwell/hysteresis); the push/pull split (default View vs addressable-on-demand, viewshop `viewRequestCap`/authoring budget).
- **F4 — Reader (WHO):** seat-keyed templates (pm | strategist | watcher, per `wa0j.52`); model-tier matching; tokenizer-optimal renders (`wa0j.27`).
- **F5 — Budget (HOW MUCH):** total percept budget per phase (`VIEW … budget N`), per-pane spend, capacity curves (`wa0j.33` — the marginal-value derivative the climb follows).
- **F0 — the measurement meta-knob:** matched-set corpus composition, grade-axis roster, noise floor, N, pre-committed falsifiers. F0 is not climbed for advantage; it is *audited* — a change to F0 invalidates comparability and mints a new evidence regime, and every F1–F5 result is only as real as the F0 design under it.
### 3. Coordinate-ascent discipline
One family moves per experiment; the other four are frozen at declared values recorded in the ConfigId. The existing machinery is the bookkeeper: the ADR-0041 cell key stops evidence pooling across (seat, band, class, archetype, phase); `RENDERER_REVISION` mints a fresh grid column on any behavioral change; a two-family move is admissible only with a pre-registered PredictionRecord naming the interaction it tests (the forking-paths firewall extends to the search itself).
### 4. Cost-tiering and sequencing
F2 and F5 are cheap (N renderers over one frozen replay corpus; token counting under a declared tokenizer) and run wide. F1 is expensive (live cohorts, both-poles economics) and runs narrow. F3's cache-policy coordinate is settled FIRST (`wa0j.28` precedes new pane grading — until the interface effect is separated, every F1 grade carries its confound). F4's tokenizer coordinate is an ARM forever: canonical rendering stays the leaderboard control (benchmark fairness), and a tokenizer-optimal render can win a cohort without ever becoming the comparison baseline.
### 5. Ordering doctrine
Template correctness by scenario (F1, per seat and phase) precedes token-encoding optimization (F2/F4-tokenizer): a token-optimal rendering of the wrong template is a cheaper wrong answer. The corpus (`tests/golden/accept/views.kestrel`) leads, the catalog follows (ADR-0041 §3); this ADR adds only the *order in which the catalog chases it*.
## Consequences
- A percept experiment's design review reduces to three questions: which family moves, is the F0 design sound for it, and are the other four families' values declared in the ConfigId.
- "Every pane lost" style results become attributable: the family that moved is named in the record, so a null on F1 under a confounded F3 setting can no longer masquerade as a settled null on F1.
- The knob families give the founder-library and viewshop loops a shared vocabulary for what an agent-authored variant is allowed to vary.
- Nothing in ADR-0009/0029/0041 changes; this ADR is their index and sequencing discipline.
## Open questions
1. Should the family/coordinate declaration be a typed field on the experiment envelope (machine-checked) rather than convention? Leaning yes once the first cross-family dispute happens.
2. Interaction effects: F1×F5 (pane value depends on budget) is known-real from the capacity-curve design — when do paired moves graduate from pre-registered exception to a standing two-coordinate protocol?
3. Whether F0 audits deserve their own record kind in the evidence ledger (sibling to TemplateRecord/PredictionRecord).
## Amendment A1 (2026-07-17): the why-audit (adopted from platform canon)
This amendment extends **F0 — the measurement meta-knob** (§2): the why-audit is a mandatory F0 audit step, not a climbable dimension. It adopts, verbatim-faithful, Rule 10 of platform ADR-0043 Amendment 1. Every F1–F5 result that names a capability-attribution verdict is only as real as the why-audit under it.
**Rule 10 — No capability-attribution verdict is believed until the why-audit runs.** Any verdict attributing a scalar delta to capability, skill, or judgment (seat rankings, knob/percept wins, teacher screens, SFT/RL training gains, and every holdout burn) requires transcript-level verification of WHY the trajectories scored before the verdict is recorded: sample the highest- and lowest-scoring trajectories (≥10 each, plus a reference arm), read them against their inputs, and answer whether the delta rides the intended competence or a grader/parser/convention/surface artifact. An artifact-shaped why is recorded as **CONTAMINATED-LIFT**, never as the win. Pre-registrations declare the audit sample plan. Scope bound: model-free mechanical measurements (miner outputs, determinism checks) and self-labeled diagnostics are exempt.
**Two companions adopted with it:**
- **One-bit surface sweep.** Two-sided surface probes plus a set-level composition gate against monotone one-bit classifiers: a curated item set must pass the composition gate before it is graded (a classifier that separates the set on a single surface bit fails it).
- **Rendered-difference fixture rule.** For any randomization or augmentation (shuffles, twin construction), assert the RENDERED example differs between orderings/augmentations — never the internal structure. Verify shuffles at the rendered-example level.
**Provenance:** platform ADR-0043 Amendment 1 (Rule 10) @ `1e5adf6`; origin bead `kestrel-9n4q`; operational form in the watcher `PROGRAM.md` §9 A4 @ `639d325`. Origin evidence: the bench lane's wave-7 conformance signal (a `+0.5` "care signal" that was 100% grader-formula conformance, visible only in transcripts) and this program's repeated artifact catches that scalars missed (grammar-off, transcript-prefix, units, parser-death, emission-failure-as-reading-loss, the SFT shuffled-twin collapse).