UNPKG

pi-decider

Version:

Decision backends for pi and omp — TypeSafe Jev, OpenRouter's Decisions API, or an OpenAI-compatible chat proxy — exposed as one typed tool (noul / choice / score)

298 lines (237 loc) • 15.9 kB
# pi-decider Decision backends as a tool call, for **pi** and **omp**. The model asks a backend for typed judgments over state you hand it and gets back answers with probabilities instead of prose — so code (or the model) can threshold them. ``` decide(state, questions[]) -> answers[] with probabilities + confidence ``` Backends are swappable per call and per question, so one request can combine several of them: | backend | kind | endpoint | credential | |---|---|---|---| | `typesafe` | decisions | `POST https://api.typesafe.ai/v1/systemone` | `TYPESAFE_API_KEY` | | `openrouter` | decisions | `POST https://openrouter.ai/api/alpha/decisions` | `OPENROUTER_API_KEY` | | `llm` | chat | `POST <llm.baseUrl>/chat/completions` | `DECIDER_LLM_API_KEY` | `typesafe` and `openrouter` both serve **Jev** (OpenRouter through its Decisions API, model `~typesafe/jev-latest` — see [Jev through OpenRouter](#jev-through-openrouter)). `llm` is a stand-in: it asks an ordinary chat model and enforces the same answer contract around it. ## Install The plugin is a **package**: `package.json` declares its extension entry (`pi.extensions`) and its bundled skill (`pi.skills`), so a harness that loads the package directory gets both. `typebox` and the harness API come from the harness itself — there is nothing to `npm install`. **pi — project mount** (`.pi/settings.json`; paths resolve relative to `.pi`): ```json { "extensions": ["../plugins/pi-decider"], "skills": ["../plugins/pi-decider/skills"] } ``` pi mounts extensions by path, so the bundled skill needs the explicit `skills` line. Project settings load only after the project is trusted: pi asks once, or run `pi --approve` / save it with `/trust`. **omp — package mount** (omp loads the package, including `skills/`): ```bash omp config set extensions '["<abs path>/plugins/pi-decider"]' # persistent omp -e "<abs path>/plugins/pi-decider" # per run ``` omp's directory scan of `<cwd>/.omp/extensions/` loads *extension files* only — a package that way does **not** bring its `skills/`, which is why the mount points at the package directory itself. Global alternative: `cp -r plugins/pi-decider ~/.omp/agent/extensions/pi-decider` (then the skill needs a root omp scans, e.g. `.omp/skills/decide/` or `skills.customDirectories`). > Load the plugin **once per harness**. Two copies registering the same tool name make the second one fail to load > (`Tool "decide" conflicts with …`), so drop the global copy when you use a project mount. Then `/reload` in a running pi session (`omp` has no `/reload` — start a new session). ## Skill `skills/decide/SKILL.md` ships inside the package and teaches the model when and how to use the tool: the three primitives, the four shapes (single / fan-out / gate / composite), question-design rules, per-question backend routing, threshold and weight handling, configuration commands, and limits. Declared as `"skills": ["./skills"]` under `pi` in `package.json`, so any harness that loads the package directory discovers it: pi via the `skills` settings line above, omp directly from the package mount. pi exposes it as `/skill:decide` and lists it in the system prompt; omp additionally serves it as `skill://decide`. ## Configure `/decide setup` walks through it inside the harness: pick the default backend, choose whether a key is an environment reference (recommended) or a literal value, confirm model/base URL from their defaults, then write the file. `/decide status` shows what resolved and probes the selected backend. Everything lives in `<agentDir>/decider.json` and is re-read on every call, so edits need no `/reload`. The agent directory comes from the harness itself (`getAgentDir()`), which means: | harness | config file | override | |---|---|---| | pi | `~/.pi/agent/decider.json` | `PI_CODING_AGENT_DIR` | | omp | `~/.omp/agent/decider.json` | `OMP_CODING_AGENT_DIR` | Each backend block is self-contained — its own credential, endpoint, path, model, prices, limits: ```json { "backend": "auto", "typesafe": { "apiKey": "$TYPESAFE_API_KEY", "baseUrl": "https://api.typesafe.ai", "path": "/v1/systemone", "model": "jev-latest" }, "openrouter": { "apiKey": "$OPENROUTER_API_KEY", "baseUrl": "https://openrouter.ai/api", "path": "/alpha/decisions", "model": "~typesafe/jev-latest" }, "llm": { "apiKey": "$DECIDER_LLM_API_KEY", "baseUrl": "https://openrouter.ai/api/v1", "path": "/chat/completions", "model": "", "temperature": 0, "maxTokens": 2048, "jsonMode": "json_object", "repairAttempts": 1 }, "modelsPath": "/v1/models" } ``` `"$VAR"` (or `"${VAR}"`) resolves the environment at request time, so keys stay out of the file; a literal value is used as-is. Per-backend `timeoutMs`, `maxRetries`, `costPerMTokInput`, and `costPerMTokOutput` follow their defaults (60s, 2 retries for decisions backends, 0 for chat; Jev is $0.042/Mtok in and free out). | env var | effect | |---|---| | `DECIDER_BACKEND` | `auto` \| `typesafe` \| `openrouter` \| `llm` | | `TYPESAFE_API_KEY` / `OPENROUTER_API_KEY` / `DECIDER_LLM_API_KEY` | credentials, used when a block references nothing | | `DECIDER_<ID>_API_KEY` / `_BASE_URL` / `_PATH` / `_MODEL` | per-backend overrides (`ID` is upper-cased) | | `DECIDER_MODELS_PATH` | catalogue path (default `/v1/models`) | **`auto` resolution:** explicit call argument → the `backend` field → the first usable backend, preferring real Jev (`typesafe`, then `openrouter`) over the `llm` proxy. A backend is usable when it has a credential and, for chat, a model id. ## Tool ``` decide({ state: "Help! My payouts have been failing for 3 days.", questions: [ { id: "is_urgent", type: "noul", instructions: "Does this convey urgency?", criteria: { true: "Explicitly time-sensitive", false: "No urgency expressed" } }, { id: "department", type: "choice", instructions: "Which team should handle this?", criteria: { billing: "Payments, invoicing, refunds", technical: "Bugs, outages", none: null } }, { id: "frustration", type: "score", instructions: "How frustrated is the customer?", criteria: ["Calm", "Frustrated", "Very angry"] } ] }) ``` | parameter | notes | |---|---| | `state` | string, structured JSON, or a JSON-encoded string | | `questions[].id` | answer key, echoed back | | `questions[].type` | `noul` (P(yes)), `choice` (one option + distribution), `score` (position across ordered levels) | | `questions[].instructions` | one narrow judgment; string or structured JSON | | `questions[].criteria` | noul `{true?, false?}`; choice `{option: rubric\|null}` or an array (≥2); score ordered levels (≥2) | | `questions[].backend` / `questions[].model` | route this question to a specific backend/model | | `backend` / `model` | call-level default for every question that does not override it | Questions sharing a backend+model are sent as one request; different batches run in parallel. Failures are per-question: one failing batch only marks its own questions `ERROR`. Answer text carries one header per batch, then one block per answer: ``` openrouter · OpenRouter Decisions API · model ~typesafe/jev-latest · provider TypeSafe · 640 ms · tokens 321 in / 57 out · cost $0.000013 is_urgent [noul] 0.87 department [choice] billing · confidence 0.91 probabilities: billing 0.8 · technical 0.1 · none 0.1 frustration [score] 1.5 · confidence 0.76 levels: 0 Calm · 1 Frustrated · 2 Very angry probabilities: 0 0.33 · 1 0.33 · 2 0.33 ``` Structured results (per-batch endpoint, provider, tokens, cost, plus every question/answer/issue/note) are persisted on the tool result `details`, and tokens/cost flow into pi's usage totals. ## Commands ``` /decide status: resolved config, per-backend state, live probe (default) /decide add [backend] guided: key, model, base URL for one backend /decide set [t] [f] [v] one validated field; interactive when a part is missing /decide unset [t] [f] drop a field, or a whole backend block (confirms first) /decide question [text] ask your own question (interactive; Enter through = smoke test) /decide models [b] [text] models a backend catalogue reports, optionally filtered /decide setup guided setup incl. the default-backend choice /decide help /jev alias for /decide ``` Aliases: `setup` = `init`, `unset` = `remove` = `rm`, `question` = `ask`. `set`/`unset` targets are `root` (plus the shorthand `set backend llm`) and the backend ids. Writable fields: `apiKey`, `baseUrl`, `path`, `model`, `timeoutMs`, `maxRetries`, `costPerMTokInput`, `costPerMTokOutput`, and for `llm` also `temperature`, `maxTokens`, `jsonMode`, `repairAttempts`. Arguments are optional everywhere — a bare `/decide set` walks you through target, field (showing current values), and value, and literal credentials are masked in the report. Everything is also scriptable: `/decide set openrouter apiKey $OPENROUTER_API_KEY`. ### Question shapes `/decide question` covers the documented usage shapes rather than a single canned probe. Enter through the prefills to get a smoke test; answer the dialogs to ask anything. | shape | what it does | |---|---| | `single` | one typed question (noul / choice / score) | | `fanout` | many independent questions in one request (speculative fan-out) | | `gate` | confidence-gated routing: noul gates on its probability, choice/score on confidence, and the verdict is printed per question | | `composite` | composite scoring: each score answer is normalized over its own levels, then combined by weight in code | Intent routing needs no separate shape: it is a `choice` question whose answer selects a handler in the code that calls the tool. ## Jev through OpenRouter OpenRouter serves Jev on its **Decisions API**, not chat completions: ``` POST https://openrouter.ai/api/alpha/decisions Authorization: Bearer $OPENROUTER_API_KEY { "model": "~typesafe/jev-latest", "state": "...", "questions": { "is_urgent": { "type": "noul", "instructions": "..." } } } → { "model": "~typesafe/jev-latest", "provider": "TypeSafe", "answers": { ... }, "usage": { "inputTokens": ..., "outputTokens": ... } } ``` Verified against the live service on 2026-09-18: - `GET https://openrouter.ai/api/v1/models/~typesafe/jev-latest/endpoints` → `modality: "text->decisions"`, `output_modalities: ["decisions"]`; `POST /api/alpha/decisions` without a key returns 401 (the route exists). - Pricing matches TypeSafe direct: $0.042/Mtok input, $0/Mtok output. - Only the `~typesafe/jev-latest` family alias exists (pinned ids such as `~typesafe/jev-1.13.0` return 404), and it is absent from the public `/v1/models` listing — `/decide status` falls back to the model-route lookup above. - Usage keys are camelCase on OpenRouter and snake_case on TypeSafe; both are accepted, and `provider` is reported when the response carries it. ## Contract enforcement The plugin owns the input and output shapes for every backend: - Questions are validated and canonicalized before any call (`normalizeQuestions`); aliases (`yes_no`, `rating`, `classify`), option arrays, and level maps are normalized, while duplicate ids, single-option choices, and single-level scores are rejected. - Answers are validated against the question they belong to (`conformAnswers`): `noul` in `[0,1]`, `choice` must be one of the criteria keys, distributions must cover every option/level and are renormalized to sum 1, `score` must lie inside the level range. - The `llm` backend additionally demands a single JSON object (`response_format: json_object`, retried without it if the endpoint rejects the parameter), recovers fenced or prose-wrapped output, and retries once with the validation errors fed back before reporting a per-question `ERROR`. - A question that names an unusable backend (no credential, or a chat backend without a model id) is reported as that question's `ERROR`; the rest of the call still runs. Only an unusable call-level default is fatal, because then no question has a backend left. - A chat model that spends its entire output budget before writing JSON (reasoning models do) is retried once with a budget raised 4× (capped at 16k). A batch that still cannot answer fails with the budget and the reasoning tokens in the message, and the tokens/cost of every attempt — including the rejected ones — are counted in the report. - `confidence` is only ever relayed, never invented; `cost` appears per backend: computed from configured Jev prices, taken from the response when a chat endpoint reports it, otherwise omitted. ## Benchmark `test/jev-suite/` is a scored benchmark for the backends, not a smoke test: **100 items** over **8 self-contained state documents**, on two axes — domain (`reasoning`, `arithmetic`, `news`, `law`, `medicine`, `engineering`, `meta`) × category (`judgment`, `calibration`, `abstention`, `arithmetic`, `architecture`, `meta`, `no-answer`). Every item carries ground truth — a computed value, a normative best option, a forbidden assertion — or an invariant, so a run reports **pass/fail per item, Brier over the `noul` items, and pair assertions**: language / option-order items must return the same pick, and a changed constraint (`w1` vs `w2`, `l1` vs `l2`) must flip it. The law and medicine documents are fictional and only ever test applying the clauses/protocol stated in the document. ```bash bun run test/jev-suite/run.ts # live; backend from decider.json bun run test/jev-suite/run.ts --backend openrouter --repeat 2 # + per-pass stability bun run test/jev-suite/run.ts --json out.json # record a run as a reference bun run test/jev-suite/run.ts --replay out.json # re-score offline, no network bun test test/jev-suite # suite integrity + scorer only ``` `test/jev-suite/reference/jev-1.13.json` is a recorded run (`~typesafe/jev-latest`, 2026-09-18): 100/100 answered, 15.3k in / 4.1k out, **$0.00064**, Brier **0.079** over 37 `noul` items. | domain | pass | what it says | | --- | --- | --- | | reasoning | 20/20 | determinate claims, routing, severity, language + option-order invariance | | law | 9/9 | clause application, deadline arithmetic, "the document does not say" | | medicine | 10/10 | protocol application, weight-based dose, contraindications, scope limits | | news | 10/10 | source quality, headline-vs-body, missing-source detection, unknown follow-up | | engineering | 24/26 | architecture calls, quorum/capacity/cache arithmetic, one latency + one retry error | | arithmetic | 13/22 | **the weak spot**: multi-step and 6+-digit evaluation, plus the framing effect below | | meta | 2/2 | self-contradicting requirement, absolute claim | Two measured properties the suite pins down: * **Framing matters.** The same fact scores very differently when the candidate answer sits in `state`: `num_span_claim` 0.21 ("is 1098.58 h correct?") versus `num_span_value` picking `1098.58` outright. Put candidates in `criteria`, keep `state` to raw facts. * **No-answer options only help when the model is unsure.** With four wrong options and an escape hatch, `opt_span_escape` and `opt_ledger_escape` take the hatch, while `opt_comb_escape` walks past it at confidence 0.92 (its own wrong value looks right). A confidence threshold is the only reliable guard. ## Limits Jev takes text only (convert images/audio first), has a 64k request budget (32k for `state` plus the longest question), streams nothing, and returns decisions rather than text or tool calls — which is why this is a tool and not a pi model provider. A `llm` backend answer is a general model's best effort, not a calibrated System One answer.