UNPKG

pi-lens

Version:

Real-time code feedback for pi — LSP, linters, formatters, type-checking, structural analysis & booboo

265 lines (234 loc) • 19.2 kB
# LSP capability matrix — affirmative-clean-signal strategy How pi-lens knows a just-edited file is **clean** (no diagnostics) vs the server simply **hasn't answered yet** (cold/crashed/silent). pi-lens waits *synchronously* for a verdict, so — unlike an editor, which renders asynchronously and never needs to decide — it must have a positive signal. There is no signal for "silence", so we classify each server and pick a per-server strategy. (Background: #240; mechanism confirmed against Neovim's LSP client, which sidesteps this entirely by being async.) Generate/refresh this matrix with `node scripts/characterize-lsp.mjs [--install]` (the `mode` column) and `node scripts/probe-clean-signal.mjs [--install]` (the `clean-behavior` column — 4-way: 2 / 2* / 3 / unknown — and the `first-publish` column, #3310). Both **merge** in place: a server the running host couldn't spawn keeps its prior row, so an ubuntu-poor run can't regress a richer one (#390). The nightly **tool-smoke** workflow runs both (plus `server-capabilities.mjs`) and opens/updates a single auto-PR (`bot/lsp-docs-refresh`, "docs(nightly): refresh LSP capability docs") with the regenerated docs — so this file self-populates from CI without manual copy-paste. `probe-clean-signal.mjs` also runs a **drift check** (#529): it compares each probed server's observed `clean-behavior` against the hand-set `silentOnClean` marker in `clients/lsp/wait-policy/strategies.ts` and writes any mismatch to the `## silentOnClean drift` section below. This is telemetry only — never a CI gate — because the probe is a timing-based negative observation; a mismatch just tells a human the marker may need updating. `unknown` observations are never compared in either direction (a slow/absent server isn't evidence of anything). The native TS7 launch variant (`typescript7`/`typescript7-clean`, #524/#526) is deliberately excluded from comparison against classic's marker — they share a server id but not a verified clean-signal behavior. What IS gated is the pair of COMMITTED sources: `tests/config/lsp-clean-behavior-census.test.ts` (#3347) compares the `clean-behavior` column below against the `silentOnClean` markers in both directions — every measured push row against its marker, and every marker against a push row measured `silent` — so a nightly refresh cannot merge a re-measured server until its marker moves with it, and a marker cannot outlive the measurement that justified it. Push rows the probe has not classified are admitted by name in that file, and an admission reds once its row becomes measured. ## First publish: is the first push an answer? (#3310) A second, ORTHOGONAL axis to the tiers below, measured from the same probe run and written as the `first-publish` column: | value | meaning | wait policy | |---|---|---| | `direct` | the first publish for a dirty file carries the findings | the first publish resolves the wait, as it always did | | `empty-first` | the server answers `didOpen` with an EMPTY set while a one-time index builds, then publishes again once indexing ends | the client HOLDS that first empty publish (`emptyFirstPublish: "indexing"` in `clients/lsp/wait-policy/strategies.ts`); the server's own next publish releases it, and the hold is one-shot per client session | | `empty-only` | every publish on the dirty fixture was empty | not classifiable on this axis — an undiagnosed dirty fixture looks exactly like an unfinished index, so nothing is inferred in either direction | | `n/a (pull)` / `TBD` | pull-mode, or not measured on a host that could reach the server | no behavior change | Why this axis had to exist at all: the wait needs a positive signal, and an empty publish is a positive signal for a Tier 2/2\* server (a genuinely clean file) and a MEANINGLESS one for a server that has not finished indexing. The two are identical on the wire, so the difference is a measurement, not an inference. intelephense is the measured member (2026-09-23, v1.18.5): `[]` at +295ms, `indexingStarted` at +301ms — i.e. AFTER the empty publish, so no work-in-progress signal is available at the moment the decision must be made — `indexingEnded` at +689ms and the real 2-diagnostic set at +696ms. On a genuinely clean file the same session publishes `[]` twice (at +304ms and again at +663ms, right after indexing), which is what lets the held publish be released by the server itself rather than by a timer; a WARM touch publishes the real set first, so the hold costs nothing once the session is warm. pi-lens does not advertise `window.workDoneProgress` (#974: advertising it crash-loops opengrep's `--experimental` LSP mode), so `$/progress` is not an available discriminator either. `tests/config/lsp-first-publish-census.test.ts` compares this column against the `emptyFirstPublish` markers and reds in BOTH directions, which is the expiry check: when the nightly re-measures a server into a different class, the docs refresh cannot merge until the marker is updated with it. ## The strategies | Tier | Signal | Affirmative clean? | Example | |---|---|---|---| | **1 — pull** | `textDocument/diagnostic` returns an authoritative report (empty = clean) | YES, deterministic | rust-analyzer | | **2 — push, publishes-versioned** | `publishDiagnostics([])` **with version** on every scan, incl. clean→clean | YES, currency-proven via version | ast-grep | | **2\* — push, publishes-unversioned** | re-publishes on a clean scan but **version-less** — the wait still early-returns (the client accepts a version-less publish as fresh: it can't be proven stale), but currency is only *temporally correlated*, not proven. Since #3484 the client fences a server measured to answer a `documentSymbol` fence before it publishes (`diagnosticsFence: "reply-first"`: yaml, intelephense): a pre-edit publish is dropped until the fence's reply. Every other 2\* server (docker publishes before answering; prisma, taplo, zls, dart, gleam, clojure unmarked) keeps this behaviour | YES at runtime, with a staleness-risk caveat (not a latency cost) | opengrep | | **3 — push, silent on clean** | server publishes nothing when nothing changed | **NO** — budget-wait floor (safe; a timeout is *not* a false clean). **This tier is #458's learned-deadline target set.** | typescript-language-server | | **Navigation-only — custom, no evidence** | custom `lsp.servers.*` entry has no pull provider and has not published in this session | **NO** — diagnostics are unsupported; skip the wait and report navigation-only | Dexter | Detection is **cached** at `initialize` (`detectWorkspaceDiagnosticsSupport` → `state.workspaceDiagnosticsSupport.mode`, upgraded on `client/registerCapability`), so the tier is free at collection time — no per-edit probe. Custom servers without a pull provider begin with one bounded first-contact probe. The probe uses the smaller of the server's push-wait budget and the live hook deadline. If the hook cuts it off first, the result remains unconfirmed and the next touch probes again. Only silence through the full push budget latches that server id as navigation-only for the service session. A publish upgrades it to the ordinary push-wait policy. ## Matrix (dev box + CI nightly; mode last refreshed 2026-06-17 from run 27713958681, clean-behavior probed in run 35914033696 — #460) `mode` from cached capabilities; `clean-behavior` from the phase-aware publish-trace probe. The probe attributes publishes to two phases — the **dirty touch** (proves the server is live) and the **clean transitions** (the discriminator) — and classifies 4-way: `publishes-versioned` (tier 2: publish WITH version on a clean transition — affirmative + currency-proven), `publishes-unversioned` (tier 2\*: version-less publish on a clean transition — the wait still early-returns at runtime, currency only temporally correlated), `silent` (tier 3: alive on dirty, silent on clean — budget-wait, the #458 target), `unknown` (no publish at all — slow/absent, conservatively not classified). `src` = where a row was measured: **ci** = the nightly steps, **dev** = the dev box (a row measured on both reads `dev+ci`). Merges never blank a prior good value, so a CI non-result leaves the dev classification standing. Two bounded guards keep that preservation from hiding a dead instrument (#3401). A `direct` `first-publish` cell whose axis the nightly probe does not re-observe is stamped once with the date of its first miss and degrades to `unknown` once five calendar days have elapsed since it (skipped nights do not stall it, and a re-observation clears the stamp). `empty-first` cells are never expired: they back the live `emptyFirstPublish` markers (php, terraform) and expiring one would erase the measurement behind a marker. A `clean-behavior`/ `tier` change is written only after two consecutive nightly runs observe the same new value, so one flapping nightly (ast-grep went 2 → 2* → 3 → 2* across four runs) cannot rewrite a cell; a night that measured nothing for the lang resets the hold. A subset probe (`probe-clean-signal.mjs <langs>`) leaves the bookkeeping of the langs it did not probe untouched. The bookkeeping lives in the generated `## Capability matrix refresh state` section at the end of this doc, the only state the refresh persists. The clock is the nightly run, not the bot PR's merge: each nightly starts from the last `bot/lsp-docs-refresh` doc when that branch is ahead of master and was built on master's current doc (`scripts/seed-matrix-from-bot-branch.mjs`), and from master's doc otherwise (branch absent, squash-merged, or master edited the doc since). Closing the bot PR unmerged does not reset the bookkeeping: the branch is kept and keeps seeding, so only deleting `bot/lsp-docs-refresh` resets it. `vue`'s `clean-behavior` was hand-reset to `unknown` (#3390): its `publishes-unversioned` cell came from 58/45 publishes that the shared `extension.log` window had attributed to vue but that belonged to `tinymist`. With the sink scoped per server, nightly 36046209160 measured vue 0/0, and an `unknown` result is never written by the merge above — so the refuted value had to be cleared by hand. It stays `unknown` until a run observes `@vue/language-server` publish; `tests/config/lsp-clean-behavior-census.test.ts` carries the named admission until then. | lang | server | mode | clean-behavior | first-publish | tier | src | |---|---|---|---|---|---|---| | json | vscode-json-language-server | pull | — | n/a (pull) | 1 | dev+ci | | css | vscode-css-language-server | pull | — | n/a (pull) | 1 | dev+ci | | html | vscode-html-language-server | pull | — | n/a (pull) | 1 | dev+ci | | rust | rust-analyzer | pull | — | n/a (pull) | 1 | dev+ci | | svelte | svelte-language-server | pull | — | n/a (pull) | 1 | dev+ci | | deno | deno (alt of typescript) | pull | — | n/a (pull) | 1 | dev+ci | | ruby | ruby-lsp | pull | — | n/a (pull) | 1 | ci | | csharp | csharp-ls | pull | — | n/a (pull) | 1 | ci | | typescript | typescript-language-server | push-only | silent | direct | 3 | dev+ci | | markdown | marksman | push-only | silent | direct | 3 | ci | | lua | lua-language-server | push-only | silent | direct | 3 | dev+ci | | python | pyright | push-only | publishes-versioned | direct | 2 | dev+ci | | jedi | jedi-language-server (alt of python) | push-only | publishes-versioned | direct | 2 | ci | | yaml | yaml-language-server | push-only | publishes-unversioned | direct | 2* | dev+ci | | shell | bash-language-server | push-only | publishes-versioned | direct | 2 | dev+ci | | dockerfile | docker-langserver | push-only | publishes-unversioned | direct | 2* | dev+ci | | toml | taplo | push-only | publishes-unversioned | direct | 2* | dev+ci | | terraform | terraform-ls | push-only | publishes-unversioned | empty-first | 2* | dev+ci | | prisma | @prisma/language-server | push-only | publishes-unversioned | direct | 2* | dev+ci | | php | intelephense | push-only | publishes-unversioned | empty-first | 2* | dev+ci | | zig | zls | push-only | publishes-unversioned | direct | 2* | dev+ci | | vue | @vue/language-server | push-only | unknown | unknown | 2/3? | dev+ci | | dart | dart language-server | push-only | publishes-unversioned | direct | 2* | ci | | gleam | gleam lsp | push-only | publishes-unversioned | direct | 2* | ci | | clojure | clojure-lsp | push-only | publishes-unversioned | direct | 2* | ci | | opengrep | opengrep (aux) | push-only | publishes-unversioned | direct | 2* | dev+ci | | ast-grep | ast-grep (aux) | push-only | publishes-versioned | direct | 2 | dev+ci | | cue | CUE Language Server (cue lsp serve) | push-only | publishes-versioned | direct | 2 | dev+ci | **Unknown — fixture exists, mode not yet captured.** The toolchain-gated family (no auto-install today; tracked in #241) — `go` (gopls), `java` (jdtls), `kotlin`, `swift` (sourcekit-lsp), `cpp` (clangd), `haskell`, `elixir`, `ocaml`, `nix` (nixd), `fsharp`. Their servers don't install in the nightly, so characterize reports `unknown` (a non-failure ⚠). Once #241 lands they'll fill in the same way clojure-lsp/gleam now do (both auto-install via the github strategy and were characterized `push-only` in the run above). ## Key findings - **Mode ≠ tier, and the split needs BOTH axes.** Push-only further splits along latency (does anything publish on a clean transition? — silence is the only budget-wait case, because pi-lens's publish handler emits and early-returns the wait on every publish, versioned or not, with the single #3310 exception of a held empty FIRST publish from an `empty-first` server) and currency-proof (is the publish versioned, i.e. provably about the live edit?). The 4-way `probe-clean-signal.mjs` measurement drives this: ast-grep → 2, yaml/opengrep → 2\*, typescript (clean file) → 3. - **opengrep is 2\*, not 3 and not plain 2.** It *does* re-publish on a clean scan (the wait early-returns at runtime — fast), but every push carries `pubVersion=undefined`, so currency is only temporally correlated, not proven — a staleness-risk note, not a latency cost. An earlier hand-note called it Tier 2 on the "re-publishes" observation alone; the phase-aware probe refines it to 2\*. - **typescript's clean behavior is diagnostic-set-dependent (major probe finding).** On a DIRTY file it re-publishes (version-lessly) after every change — the dirty fixture measures 2\*. On a genuinely CLEAN file (the `typescript-clean` fixture) it publishes nothing on a clean→clean edit — silent, Tier 3. The clean-file behavior is the production case (the observed budget-wait timeouts), so the matrix row records the clean fixture's verdict; the probe prefers `clean: true` fixtures for exactly this reason. Corollary: a 2\* measured only on a dirty fixture may overstate a server whose publishes stop when its set goes empty — langs without a clean fixture carry that caveat. - **The probe's publish capture was dead, and the column looked alive anyway (#3310).** `probe-clean-signal.mjs` intercepted `console.error` for the `[lsp-pub]` trace, but #1333 moved that trace's sink to `extension.log` — so every phase counted zero publishes, every server classified `unknown`, and the #390 merge guard then preserved the July 2026 values verbatim. A dead instrument and a healthy one are indistinguishable when non-results are discarded by design. The probe now reads the sink the client actually writes; `first-publish` measurements on the dev box, 2026-09-23: php `empty-first`, typescript / opengrep / ast-grep / marksman `direct` (so the class does not extend to them on the measurement, whatever their comments suggest). - **#458's learned-deadline target set = the tier-3 rows only.** 2\* rows resolve the wait at runtime and must NOT be given learned deadlines. - **Tier 3 is budget-bound by necessity**, not laziness: a silent server's silence is ambiguous (clean-unchanged vs still-analyzing), so shortening the wait or reusing `lastKnownDiagnostics` would risk a false clean. The wait *is* the safety mechanism. - **ast-grep (Phase 2 / #239) is Tier 2** — it self-signals clean on every scan, so it is not the bottleneck. The cost on a clean with-auxiliary touch is the *silent primary* (typescript), a pre-existing Tier-3 cost independent of ast-grep. ## Completing the matrix The fixtures (`tests/fixtures/tool-smoke/<lang>/`) are durable and cover every registered server. `mode` is read from the server's advertised capabilities at `initialize`, so it is **content-independent** — for languages that already had a tool-layer fixture we point `characterize-lsp.mjs` at the existing (deliberately dirty) `bad.*` source rather than a colliding clean duplicate; new languages get a minimal clean source. Either way the mode reported is the same. The nightly **tool-smoke** workflow runs `characterize-lsp.mjs --install` **and** `probe-clean-signal.mjs --install` (after the LSP handshake layer) on `ubuntu-latest`, then opens/updates an auto-PR with the regenerated docs (#390) — so both the `mode` and `clean-behavior` columns self-populate in CI without manual copy-paste. They fill for servers that either auto-install (npm/pip/github — including clojure-lsp and gleam, both github-strategy as of f263cf3) or whose toolchain the workflow provisions (Ruby/Dart/Zig + .NET→csharp). The remaining `unknown` rows are the toolchain-gated family (#241): until `runtimeInstall` + canonical-bin discovery land, their servers don't install in CI and stay ⚠. The **merge guard** means a CI run that can't reach a server never blanks its dev-measured row. The `clean-behavior` split (4-way: 2 publishes-versioned / 2\* publishes-unversioned / 3 silent / unknown) is now measured per-server by `probe-clean-signal.mjs`, no longer a manual one-off. Locally confirmed: ast-grep + ast-grep-baseline → 2, yaml + opengrep → 2\*, typescript → 3 on its clean fixture (2\* on the dirty fixture — see Key findings); the rest fill in as the nightly reaches them. ## silentOnClean drift (nightly-generated) Telemetry only — never a CI gate. Compares each probed server's observed `clean-behavior` against `clients/lsp/wait-policy/strategies.ts`'s `silentOnClean` marker; a mismatch means the marker may need a human update (#529). `unknown` observations are never compared (a slow/absent server is not evidence either way). _None observed as of the last probe run._ ## Capability matrix refresh state (nightly-generated) Bookkeeping for the date-based `direct` `first-publish` expiry (#3401), the two-run `clean-behavior` hysteresis and the consecutive-night `idle-eviction` counts (#3989). Regenerated every run; never a measurement. ```json {"idle-eviction":{"docker":{"nights":[{"day":"2026-10-07","rssMb":64,"coldMs":566}]},"json":{"nights":[{"day":"2026-10-07","rssMb":67,"coldMs":1116}]},"powershell":{"nights":[{"day":"2026-10-07","rssMb":155,"coldMs":2601}]},"python-jedi":{"nights":[{"day":"2026-10-07","rssMb":51,"coldMs":1796}]},"zizmor":{"nights":[{"day":"2026-10-07","rssMb":58,"coldMs":608}]}}} ```