pi-decider
Version:
Decision backends for pi and omp — TypeSafe Jev, OpenRouter's Decisions API, or an OpenAI-compatible chat proxy — exposed as one typed tool (noul / choice / score)
130 lines (101 loc) • 6.88 kB
Markdown
---
name: decide
description: Use the `decide` tool for calibrated typed judgments (Jev / TypeSafe System One, OpenRouter Decisions API, or a chat-model proxy) instead of asking a general model for prose. Covers the three question primitives, the four usage shapes (single, fan-out, confidence gate, composite score), backend routing per call and per question, question-design rules, thresholds, and configuration. Load when a task needs a probability, a pick among named options, a graded score, routing/ranking, moderation, or verification.
---
# decide — typed decisions as a tool call
`decide` returns **answers, not text**: a probability, a chosen option with its full distribution, or a position on
an ordered scale. Code (or you) thresholds the numbers. Reach for it whenever the right output is a decision rather
than a paragraph.
```
decide(state, questions[]) -> answers with probabilities + confidence
```
## Primitives (the three question types)
| type | answers | use for |
|---|---|---|
| `noul` | `noul`: P(yes) in [0,1] | one yes/no condition; → probability is the answer |
| `choice` | `choice` + `probabilities` + `confidence` | pick one of a named set (routing, intent, entity matching) |
| `score` | `score` + `probabilities` + `legend` + `confidence` | graded position across ordered levels |
- A `noul` near 0.5 means yes and no are equally likely — **not** "medium intensity". For intensity use `score`.
- `choice` probabilities compare the listed options with each other; `confidence` says how concentrated that
distribution is. Two decent options can both be plausible — read the distribution, not just the winner.
- `score` may land between levels (probability-weighted). Levels are 0-indexed.
## Question design
- **One narrow, coherent judgment per question.** Split independent dimensions; don't destroy the relationship you
are judging.
- **Criteria carry the definitions.** `choice` needs ≥2 options (`{option: rubric|null}` or an array of options);
give every `choice` a no-match option when nothing may fit. `score` needs ≥2 ordered level descriptions, lowest
first. `noul` criteria (`{true?, false?}`) are optional but sharpen the boundary.
- **Put everything the judgment needs in `state`** (source text, records, policies, current facts) — a string,
JSON object/array, or a JSON-encoded string.
- **Batch independent questions into one call.** They run in parallel in a single request, which is far cheaper and
faster than one call per question. Add a second call only when an earlier answer is needed to build the next
state or to choose the next options.
- **Give each question an `id`** you will recognise in code; ids are not sent to the model.
## Shapes (compose them in code)
| shape | what it is | what your code does |
|---|---|---|
| single | one question | act on the number |
| fan-out | many independent questions, one request | consume only the answers that apply; keep speculative ones cheap |
| gate | confidence-gated routing | act when P/confidence ≥ threshold; otherwise hold, escalate, or ask a human |
| composite | several `score` questions | normalize each over its own levels, weight, combine — weights are code, not prompts |
Intent routing needs no special shape: it is a `choice` question whose answer selects a handler.
Thresholds belong to the caller. Validate them on real data, and remember that a low-confidence answer is often
still usable for a harmless preference while it must not silently drive a consequential action.
## Backends
| backend | kind | serves |
|---|---|---|
| `typesafe` | decisions | TypeSafe System One (`/v1/systemone`) |
| `openrouter` | decisions | OpenRouter Decisions API (`/api/alpha/decisions`, Jev) |
| `llm` | chat | any OpenAI-compatible chat model, wrapped in the same contract |
- `auto` (default) prefers real Jev when a credential exists, then the chat proxy.
- Override per call with `backend` / `model`, and **per question** with `questions[i].backend` / `.model`. Questions
sharing a backend+model are batched into one request; different batches run in parallel — so one call can combine a
cheap calibration question with a second opinion from another model.
- An `llm` answer is a general model's best effort, not a calibrated System One answer: treat it as lower trust, and
note that `confidence` is only present when that model returned one.
## Configure
```
/decide status: resolved config, per-backend state, live probe
/decide add [backend] guided: key, model, base URL
/decide set [t] [f] [v] one field, e.g. set openrouter model ~typesafe/jev-latest
/decide unset [t] [f] drop a field or a whole backend block
/decide question [text] ask your own question (shapes: single | fanout | gate | composite)
/decide models [b] [text] catalogue of a backend
/decide setup guided setup incl. the default backend
```
Aliases: `/jev` = `/decide`; `setup` = `init`, `unset` = `remove` = `rm`, `question` = `ask`.
Credentials come from the config file (`<agentDir>/decider.json`: `~/.pi/agent/` under pi, `~/.omp/agent/` under
omp) or the environment — `TYPESAFE_API_KEY`, `OPENROUTER_API_KEY`, `DECIDER_LLM_API_KEY`. Prefer `$ENV_VAR`
references over literal keys in the file.
## Limits
Jev is text-only (convert images/audio first), has a 64k request budget (≈32k for `state` plus the longest question),
streams nothing, and cannot call tools — it answers, your code acts. Long states cost input tokens on every call
(Jev does not bill output).
## Examples
Route a support message (one request, two independent judgments):
```jsonc
{
"state": "Help! My payouts have been failing for 3 days.",
"questions": [
{ "id": "urgent", "type": "noul", "instructions": "Does this message convey urgency?",
"criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" } },
{ "id": "team", "type": "choice", "instructions": "Which team should handle this?",
"criteria": { "billing": "Payments, invoicing, refunds", "technical": "Bugs, outages", "none": null } }
]
}
```
Then: `if (urgent.noul > 0.8 && team.choice !== "none") route(team.choice)` — the threshold is yours.
Grade a candidate with a gate and a composite:
```jsonc
{
"state": { "candidate": "…", "requirements": "…" },
"questions": [
{ "id": "meets_bar", "type": "noul", "instructions": "Does this candidate meet the stated bar?" },
{ "id": "depth", "type": "score", "instructions": "How deep is the relevant experience?",
"criteria": ["None", "Adjacent", "Direct", "Deep"] },
{ "id": "evidence", "type": "score", "instructions": "How well is it evidenced?",
"criteria": ["Claim only", "Some evidence", "Strong evidence"] }
]
}
```
Gate on `meets_bar.noul`, then take `0.7 × depth/(levels-1) + 0.3 × evidence/(levels-1)`.