UNPKG

kestrel.markets

Version:

A typed, token-efficient language + runtime for agentic trading: agents author bounded plans, the runtime fires them at the tick. CLI + typed library + MCP server.

58 lines (44 loc) 3.04 kB
# The two-part evaluation protocol (agent-day + armory counterfactual) > The standing protocol for grading agent judgment on Kestrel. Extends ADR-0006 > ("everything authored is gradable") into the experiment that answers the platform's > own hinge question. Strategy *content* produced under this protocol is application > data: it lives in the registry (SQLite) and gitignored run directories, never in > Kestrel source. ## Part 1 — the agent-day grade (supervised) One session, lived end to end by an agent in sim: 1. **T-5m briefing.** Before the open the agent receives the OPEN keyframe: macro / long-horizon tape, overnight and futures context, the session's instruments and book state. Date-blind (relative time only). 2. **Authoring.** The agent writes the day's plan document *and its wake sequence* — when-to-look is part of the strategy and is graded with it. 3. **The day.** The session runs on the recorded bus. At each wake — authored triggers plus the structural cadence into the close (T-2h, T-1h, T-30m, T-15m, T-10m, T-5m: the phase where theta is spent and gamma peaks) — the agent gets a delta frame (tape since last vantage, positions, fills, budget) and may revise: supersede plans, re-arm, stand down. 4. **The grade.** Realized (floor) and expected (calibrated p_fill) P&L, plus the decision trace: wakes spent, revisions made, each revision's marginal effect. ## Part 2 — the armory counterfactual (`VS unsupervised`) The same authored plans, **frozen at their first version** — no wakes delivered, no revisions — replayed pure-algo across the corpus under the same fill model. Then, per lineage and per tape: management_alpha = EV(supervised day) − EV(frozen armory, same day) Three outcomes, all informative: - **The armory alone carries edge** → the plans are a product; agents are authors. - **The delta carries the edge** → supervision is the product: the value of Kestrel is that an agent manages the plan as the portfolio and market evolve. - **Neither** → an honest null on this authoring cohort; iterate the authoring, the screen, or the wake taxonomy (each independently gradable). Aggregate across days/lineages in the registry (`leaderboard`, `lineage`) — a single day answers nothing; the protocol is designed to accumulate. ## The contamination fence (binding, per ADR-0006) An LLM's grade is **honest-tier** only on sessions after its training cutoff, date-blinded. Pre-cutoff sessions (however dramatic) are practice tier: mechanics, harness debugging, SFT material — never headline grades. The honest tier replenishes automatically: tomorrow is always unseen. ## Boundary rules - Authored plan documents, wake sets, revision transcripts, reports: registry + gitignored `data/runs/` (or the private application repo). Kestrel commits the *protocol*, never the *strategies*. - Committed results docs describe behavior qualitatively with illustrative numbers; no bulk market data, no strategy dumps, no calibration constants.