UNPKG

pi-decider

Version:

Decision backends for pi and omp — TypeSafe Jev, OpenRouter's Decisions API, or an OpenAI-compatible chat proxy — exposed as one typed tool (noul / choice / score)

130 lines (101 loc) • 6.88 kB
--- name: decide description: Use the `decide` tool for calibrated typed judgments (Jev / TypeSafe System One, OpenRouter Decisions API, or a chat-model proxy) instead of asking a general model for prose. Covers the three question primitives, the four usage shapes (single, fan-out, confidence gate, composite score), backend routing per call and per question, question-design rules, thresholds, and configuration. Load when a task needs a probability, a pick among named options, a graded score, routing/ranking, moderation, or verification. --- # decide — typed decisions as a tool call `decide` returns **answers, not text**: a probability, a chosen option with its full distribution, or a position on an ordered scale. Code (or you) thresholds the numbers. Reach for it whenever the right output is a decision rather than a paragraph. ``` decide(state, questions[]) -> answers with probabilities + confidence ``` ## Primitives (the three question types) | type | answers | use for | |---|---|---| | `noul` | `noul`: P(yes) in [0,1] | one yes/no condition; → probability is the answer | | `choice` | `choice` + `probabilities` + `confidence` | pick one of a named set (routing, intent, entity matching) | | `score` | `score` + `probabilities` + `legend` + `confidence` | graded position across ordered levels | - A `noul` near 0.5 means yes and no are equally likely — **not** "medium intensity". For intensity use `score`. - `choice` probabilities compare the listed options with each other; `confidence` says how concentrated that distribution is. Two decent options can both be plausible — read the distribution, not just the winner. - `score` may land between levels (probability-weighted). Levels are 0-indexed. ## Question design - **One narrow, coherent judgment per question.** Split independent dimensions; don't destroy the relationship you are judging. - **Criteria carry the definitions.** `choice` needs ≥2 options (`{option: rubric|null}` or an array of options); give every `choice` a no-match option when nothing may fit. `score` needs ≥2 ordered level descriptions, lowest first. `noul` criteria (`{true?, false?}`) are optional but sharpen the boundary. - **Put everything the judgment needs in `state`** (source text, records, policies, current facts) — a string, JSON object/array, or a JSON-encoded string. - **Batch independent questions into one call.** They run in parallel in a single request, which is far cheaper and faster than one call per question. Add a second call only when an earlier answer is needed to build the next state or to choose the next options. - **Give each question an `id`** you will recognise in code; ids are not sent to the model. ## Shapes (compose them in code) | shape | what it is | what your code does | |---|---|---| | single | one question | act on the number | | fan-out | many independent questions, one request | consume only the answers that apply; keep speculative ones cheap | | gate | confidence-gated routing | act when P/confidence ≥ threshold; otherwise hold, escalate, or ask a human | | composite | several `score` questions | normalize each over its own levels, weight, combine — weights are code, not prompts | Intent routing needs no special shape: it is a `choice` question whose answer selects a handler. Thresholds belong to the caller. Validate them on real data, and remember that a low-confidence answer is often still usable for a harmless preference while it must not silently drive a consequential action. ## Backends | backend | kind | serves | |---|---|---| | `typesafe` | decisions | TypeSafe System One (`/v1/systemone`) | | `openrouter` | decisions | OpenRouter Decisions API (`/api/alpha/decisions`, Jev) | | `llm` | chat | any OpenAI-compatible chat model, wrapped in the same contract | - `auto` (default) prefers real Jev when a credential exists, then the chat proxy. - Override per call with `backend` / `model`, and **per question** with `questions[i].backend` / `.model`. Questions sharing a backend+model are batched into one request; different batches run in parallel — so one call can combine a cheap calibration question with a second opinion from another model. - An `llm` answer is a general model's best effort, not a calibrated System One answer: treat it as lower trust, and note that `confidence` is only present when that model returned one. ## Configure ``` /decide status: resolved config, per-backend state, live probe /decide add [backend] guided: key, model, base URL /decide set [t] [f] [v] one field, e.g. set openrouter model ~typesafe/jev-latest /decide unset [t] [f] drop a field or a whole backend block /decide question [text] ask your own question (shapes: single | fanout | gate | composite) /decide models [b] [text] catalogue of a backend /decide setup guided setup incl. the default backend ``` Aliases: `/jev` = `/decide`; `setup` = `init`, `unset` = `remove` = `rm`, `question` = `ask`. Credentials come from the config file (`<agentDir>/decider.json`: `~/.pi/agent/` under pi, `~/.omp/agent/` under omp) or the environment — `TYPESAFE_API_KEY`, `OPENROUTER_API_KEY`, `DECIDER_LLM_API_KEY`. Prefer `$ENV_VAR` references over literal keys in the file. ## Limits Jev is text-only (convert images/audio first), has a 64k request budget (≈32k for `state` plus the longest question), streams nothing, and cannot call tools — it answers, your code acts. Long states cost input tokens on every call (Jev does not bill output). ## Examples Route a support message (one request, two independent judgments): ```jsonc { "state": "Help! My payouts have been failing for 3 days.", "questions": [ { "id": "urgent", "type": "noul", "instructions": "Does this message convey urgency?", "criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" } }, { "id": "team", "type": "choice", "instructions": "Which team should handle this?", "criteria": { "billing": "Payments, invoicing, refunds", "technical": "Bugs, outages", "none": null } } ] } ``` Then: `if (urgent.noul > 0.8 && team.choice !== "none") route(team.choice)` — the threshold is yours. Grade a candidate with a gate and a composite: ```jsonc { "state": { "candidate": "…", "requirements": "…" }, "questions": [ { "id": "meets_bar", "type": "noul", "instructions": "Does this candidate meet the stated bar?" }, { "id": "depth", "type": "score", "instructions": "How deep is the relevant experience?", "criteria": ["None", "Adjacent", "Direct", "Deep"] }, { "id": "evidence", "type": "score", "instructions": "How well is it evidenced?", "criteria": ["Claim only", "Some evidence", "Strong evidence"] } ] } ``` Gate on `meets_bar.noul`, then take `0.7 × depth/(levels-1) + 0.3 × evidence/(levels-1)`.