pi-decider
Version:
Decision backends for pi and omp — TypeSafe Jev, OpenRouter's Decisions API, or an OpenAI-compatible chat proxy — exposed as one typed tool (noul / choice / score)
298 lines (237 loc) • 15.9 kB
Markdown
# pi-decider
Decision backends as a tool call, for **pi** and **omp**. The model asks a backend for typed judgments over
state you hand it and gets back answers with probabilities instead of prose — so code (or the model) can
threshold them.
```
decide(state, questions[]) -> answers[] with probabilities + confidence
```
Backends are swappable per call and per question, so one request can combine several of them:
| backend | kind | endpoint | credential |
|---|---|---|---|
| `typesafe` | decisions | `POST https://api.typesafe.ai/v1/systemone` | `TYPESAFE_API_KEY` |
| `openrouter` | decisions | `POST https://openrouter.ai/api/alpha/decisions` | `OPENROUTER_API_KEY` |
| `llm` | chat | `POST <llm.baseUrl>/chat/completions` | `DECIDER_LLM_API_KEY` |
`typesafe` and `openrouter` both serve **Jev** (OpenRouter through its Decisions API, model `~typesafe/jev-latest`
— see [Jev through OpenRouter](#jev-through-openrouter)). `llm` is a stand-in: it asks an ordinary chat model and
enforces the same answer contract around it.
## Install
The plugin is a **package**: `package.json` declares its extension entry (`pi.extensions`) and its bundled skill
(`pi.skills`), so a harness that loads the package directory gets both. `typebox` and the harness API come from the
harness itself — there is nothing to `npm install`.
**pi — project mount** (`.pi/settings.json`; paths resolve relative to `.pi`):
```json
{
"extensions": ["../plugins/pi-decider"],
"skills": ["../plugins/pi-decider/skills"]
}
```
pi mounts extensions by path, so the bundled skill needs the explicit `skills` line. Project settings load only after
the project is trusted: pi asks once, or run `pi --approve` / save it with `/trust`.
**omp — package mount** (omp loads the package, including `skills/`):
```bash
omp config set extensions '["<abs path>/plugins/pi-decider"]' # persistent
omp -e "<abs path>/plugins/pi-decider" # per run
```
omp's directory scan of `<cwd>/.omp/extensions/` loads *extension files* only — a package that way does **not** bring
its `skills/`, which is why the mount points at the package directory itself. Global alternative:
`cp -r plugins/pi-decider ~/.omp/agent/extensions/pi-decider` (then the skill needs a root omp scans, e.g.
`.omp/skills/decide/` or `skills.customDirectories`).
> Load the plugin **once per harness**. Two copies registering the same tool name make the second one fail to load
> (`Tool "decide" conflicts with …`), so drop the global copy when you use a project mount.
Then `/reload` in a running pi session (`omp` has no `/reload` — start a new session).
## Skill
`skills/decide/SKILL.md` ships inside the package and teaches the model when and how to use the tool: the three
primitives, the four shapes (single / fan-out / gate / composite), question-design rules, per-question backend
routing, threshold and weight handling, configuration commands, and limits.
Declared as `"skills": ["./skills"]` under `pi` in `package.json`, so any harness that loads the package directory
discovers it: pi via the `skills` settings line above, omp directly from the package mount. pi exposes it as
`/skill:decide` and lists it in the system prompt; omp additionally serves it as `skill://decide`.
## Configure
`/decide setup` walks through it inside the harness: pick the default backend, choose whether a key is an environment
reference (recommended) or a literal value, confirm model/base URL from their defaults, then write the file.
`/decide status` shows what resolved and probes the selected backend.
Everything lives in `<agentDir>/decider.json` and is re-read on every call, so edits need no `/reload`. The agent
directory comes from the harness itself (`getAgentDir()`), which means:
| harness | config file | override |
|---|---|---|
| pi | `~/.pi/agent/decider.json` | `PI_CODING_AGENT_DIR` |
| omp | `~/.omp/agent/decider.json` | `OMP_CODING_AGENT_DIR` |
Each backend block is self-contained — its own credential, endpoint, path, model, prices, limits:
```json
{
"backend": "auto",
"typesafe": {
"apiKey": "$TYPESAFE_API_KEY",
"baseUrl": "https://api.typesafe.ai",
"path": "/v1/systemone",
"model": "jev-latest"
},
"openrouter": {
"apiKey": "$OPENROUTER_API_KEY",
"baseUrl": "https://openrouter.ai/api",
"path": "/alpha/decisions",
"model": "~typesafe/jev-latest"
},
"llm": {
"apiKey": "$DECIDER_LLM_API_KEY",
"baseUrl": "https://openrouter.ai/api/v1",
"path": "/chat/completions",
"model": "",
"temperature": 0,
"maxTokens": 2048,
"jsonMode": "json_object",
"repairAttempts": 1
},
"modelsPath": "/v1/models"
}
```
`"$VAR"` (or `"${VAR}"`) resolves the environment at request time, so keys stay out of the file; a literal value is
used as-is. Per-backend `timeoutMs`, `maxRetries`, `costPerMTokInput`, and `costPerMTokOutput` follow their defaults
(60s, 2 retries for decisions backends, 0 for chat; Jev is $0.042/Mtok in and free out).
| env var | effect |
|---|---|
| `DECIDER_BACKEND` | `auto` \| `typesafe` \| `openrouter` \| `llm` |
| `TYPESAFE_API_KEY` / `OPENROUTER_API_KEY` / `DECIDER_LLM_API_KEY` | credentials, used when a block references nothing |
| `DECIDER_<ID>_API_KEY` / `_BASE_URL` / `_PATH` / `_MODEL` | per-backend overrides (`ID` is upper-cased) |
| `DECIDER_MODELS_PATH` | catalogue path (default `/v1/models`) |
**`auto` resolution:** explicit call argument → the `backend` field → the first usable backend, preferring real Jev
(`typesafe`, then `openrouter`) over the `llm` proxy. A backend is usable when it has a credential and, for chat, a
model id.
## Tool
```
decide({
state: "Help! My payouts have been failing for 3 days.",
questions: [
{ id: "is_urgent", type: "noul", instructions: "Does this convey urgency?",
criteria: { true: "Explicitly time-sensitive", false: "No urgency expressed" } },
{ id: "department", type: "choice", instructions: "Which team should handle this?",
criteria: { billing: "Payments, invoicing, refunds", technical: "Bugs, outages", none: null } },
{ id: "frustration", type: "score", instructions: "How frustrated is the customer?",
criteria: ["Calm", "Frustrated", "Very angry"] }
]
})
```
| parameter | notes |
|---|---|
| `state` | string, structured JSON, or a JSON-encoded string |
| `questions[].id` | answer key, echoed back |
| `questions[].type` | `noul` (P(yes)), `choice` (one option + distribution), `score` (position across ordered levels) |
| `questions[].instructions` | one narrow judgment; string or structured JSON |
| `questions[].criteria` | noul `{true?, false?}`; choice `{option: rubric\|null}` or an array (≥2); score ordered levels (≥2) |
| `questions[].backend` / `questions[].model` | route this question to a specific backend/model |
| `backend` / `model` | call-level default for every question that does not override it |
Questions sharing a backend+model are sent as one request; different batches run in parallel. Failures are
per-question: one failing batch only marks its own questions `ERROR`.
Answer text carries one header per batch, then one block per answer:
```
openrouter · OpenRouter Decisions API · model ~typesafe/jev-latest · provider TypeSafe · 640 ms · tokens 321 in / 57 out · cost $0.000013
is_urgent [noul] 0.87
department [choice] billing · confidence 0.91
probabilities: billing 0.8 · technical 0.1 · none 0.1
frustration [score] 1.5 · confidence 0.76
levels: 0 Calm · 1 Frustrated · 2 Very angry
probabilities: 0 0.33 · 1 0.33 · 2 0.33
```
Structured results (per-batch endpoint, provider, tokens, cost, plus every question/answer/issue/note) are
persisted on the tool result `details`, and tokens/cost flow into pi's usage totals.
## Commands
```
/decide status: resolved config, per-backend state, live probe (default)
/decide add [backend] guided: key, model, base URL for one backend
/decide set [t] [f] [v] one validated field; interactive when a part is missing
/decide unset [t] [f] drop a field, or a whole backend block (confirms first)
/decide question [text] ask your own question (interactive; Enter through = smoke test)
/decide models [b] [text] models a backend catalogue reports, optionally filtered
/decide setup guided setup incl. the default-backend choice
/decide help
/jev alias for /decide
```
Aliases: `setup` = `init`, `unset` = `remove` = `rm`, `question` = `ask`.
`set`/`unset` targets are `root` (plus the shorthand `set backend llm`) and the backend ids. Writable fields:
`apiKey`, `baseUrl`, `path`, `model`, `timeoutMs`, `maxRetries`, `costPerMTokInput`, `costPerMTokOutput`, and for `llm`
also `temperature`, `maxTokens`, `jsonMode`, `repairAttempts`. Arguments are optional everywhere — a bare `/decide set`
walks you through target, field (showing current values), and value, and literal credentials are masked in the report.
Everything is also scriptable: `/decide set openrouter apiKey $OPENROUTER_API_KEY`.
### Question shapes
`/decide question` covers the documented usage shapes rather than a single canned probe. Enter through the prefills
to get a smoke test; answer the dialogs to ask anything.
| shape | what it does |
|---|---|
| `single` | one typed question (noul / choice / score) |
| `fanout` | many independent questions in one request (speculative fan-out) |
| `gate` | confidence-gated routing: noul gates on its probability, choice/score on confidence, and the verdict is printed per question |
| `composite` | composite scoring: each score answer is normalized over its own levels, then combined by weight in code |
Intent routing needs no separate shape: it is a `choice` question whose answer selects a handler in the code that
calls the tool.
## Jev through OpenRouter
OpenRouter serves Jev on its **Decisions API**, not chat completions:
```
POST https://openrouter.ai/api/alpha/decisions
Authorization: Bearer $OPENROUTER_API_KEY
{ "model": "~typesafe/jev-latest", "state": "...", "questions": { "is_urgent": { "type": "noul", "instructions": "..." } } }
→ { "model": "~typesafe/jev-latest", "provider": "TypeSafe", "answers": { ... }, "usage": { "inputTokens": ..., "outputTokens": ... } }
```
Verified against the live service on 2026-09-18:
- `GET https://openrouter.ai/api/v1/models/~typesafe/jev-latest/endpoints` → `modality: "text->decisions"`,
`output_modalities: ["decisions"]`; `POST /api/alpha/decisions` without a key returns 401 (the route exists).
- Pricing matches TypeSafe direct: $0.042/Mtok input, $0/Mtok output.
- Only the `~typesafe/jev-latest` family alias exists (pinned ids such as `~typesafe/jev-1.13.0` return 404), and it
is absent from the public `/v1/models` listing — `/decide status` falls back to the model-route lookup above.
- Usage keys are camelCase on OpenRouter and snake_case on TypeSafe; both are accepted, and `provider` is reported
when the response carries it.
## Contract enforcement
The plugin owns the input and output shapes for every backend:
- Questions are validated and canonicalized before any call (`normalizeQuestions`); aliases (`yes_no`, `rating`,
`classify`), option arrays, and level maps are normalized, while duplicate ids, single-option choices, and
single-level scores are rejected.
- Answers are validated against the question they belong to (`conformAnswers`): `noul` in `[0,1]`, `choice` must be
one of the criteria keys, distributions must cover every option/level and are renormalized to sum 1, `score` must
lie inside the level range.
- The `llm` backend additionally demands a single JSON object (`response_format: json_object`, retried without it if
the endpoint rejects the parameter), recovers fenced or prose-wrapped output, and retries once with the validation
errors fed back before reporting a per-question `ERROR`.
- A question that names an unusable backend (no credential, or a chat backend without a model id) is reported as
that question's `ERROR`; the rest of the call still runs. Only an unusable call-level default is fatal, because
then no question has a backend left.
- A chat model that spends its entire output budget before writing JSON (reasoning models do) is retried once with a
budget raised 4× (capped at 16k). A batch that still cannot answer fails with the budget and the reasoning tokens in
the message, and the tokens/cost of every attempt — including the rejected ones — are counted in the report.
- `confidence` is only ever relayed, never invented; `cost` appears per backend: computed from configured Jev prices,
taken from the response when a chat endpoint reports it, otherwise omitted.
## Benchmark
`test/jev-suite/` is a scored benchmark for the backends, not a smoke test: **100 items** over **8 self-contained
state documents**, on two axes — domain (`reasoning`, `arithmetic`, `news`, `law`, `medicine`, `engineering`, `meta`)
× category (`judgment`, `calibration`, `abstention`, `arithmetic`, `architecture`, `meta`, `no-answer`).
Every item carries ground truth — a computed value, a normative best option, a forbidden assertion — or an invariant,
so a run reports **pass/fail per item, Brier over the `noul` items, and pair assertions**: language / option-order
items must return the same pick, and a changed constraint (`w1` vs `w2`, `l1` vs `l2`) must flip it. The law and
medicine documents are fictional and only ever test applying the clauses/protocol stated in the document.
```bash
bun run test/jev-suite/run.ts # live; backend from decider.json
bun run test/jev-suite/run.ts --backend openrouter --repeat 2 # + per-pass stability
bun run test/jev-suite/run.ts --json out.json # record a run as a reference
bun run test/jev-suite/run.ts --replay out.json # re-score offline, no network
bun test test/jev-suite # suite integrity + scorer only
```
`test/jev-suite/reference/jev-1.13.json` is a recorded run (`~typesafe/jev-latest`, 2026-09-18): 100/100 answered,
15.3k in / 4.1k out, **$0.00064**, Brier **0.079** over 37 `noul` items.
| domain | pass | what it says |
| --- | --- | --- |
| reasoning | 20/20 | determinate claims, routing, severity, language + option-order invariance |
| law | 9/9 | clause application, deadline arithmetic, "the document does not say" |
| medicine | 10/10 | protocol application, weight-based dose, contraindications, scope limits |
| news | 10/10 | source quality, headline-vs-body, missing-source detection, unknown follow-up |
| engineering | 24/26 | architecture calls, quorum/capacity/cache arithmetic, one latency + one retry error |
| arithmetic | 13/22 | **the weak spot**: multi-step and 6+-digit evaluation, plus the framing effect below |
| meta | 2/2 | self-contradicting requirement, absolute claim |
Two measured properties the suite pins down:
* **Framing matters.** The same fact scores very differently when the candidate answer sits in `state`:
`num_span_claim` 0.21 ("is 1098.58 h correct?") versus `num_span_value` picking `1098.58` outright. Put candidates
in `criteria`, keep `state` to raw facts.
* **No-answer options only help when the model is unsure.** With four wrong options and an escape hatch,
`opt_span_escape` and `opt_ledger_escape` take the hatch, while `opt_comb_escape` walks past it at confidence 0.92
(its own wrong value looks right). A confidence threshold is the only reliable guard.
## Limits
Jev takes text only (convert images/audio first), has a 64k request budget (32k for `state` plus the longest
question), streams nothing, and returns decisions rather than text or tool calls — which is why this is a tool and
not a pi model provider. A `llm` backend answer is a general model's best effort, not a calibrated System One answer.