From 014e63845dfa3d4c9ff66ea3efeb6c78224410ec Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Troed=20S=C3=A5ngberg?= Date: Sun, 14 Jun 2026 12:28:14 +0200 Subject: [PATCH] docs: replace README with factual description of detection and calculation --- README.md | 93 +++++++++++++++++++++++++++++++++++++++++++++++++++---- 1 file changed, 87 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index a42bca7..1cc5f23 100644 --- a/README.md +++ b/README.md @@ -1,16 +1,97 @@ -# oc-tps +# oc-ls-stats -Displays live TPS (tokens per second), average TPS, and average TTFT (time to first token) in the OpenCode session prompt. +A TUI plugin for OpenCode that displays live prefill rate (PP) and generation rate (TG) from a llama.cpp-based llama-server. +## Display Format -![Demo](./assets/demo.gif) +The plugin renders a single line in the session prompt right slot: + +``` +1247 tps (PP) -- during prefill + 25 tps (TG) -- during generation +- tps (TG) -- idle +``` + +## Detection and Calculation + +### Server Discovery + +The plugin discovers the llama-server URL by reading the OpenCode configuration file (parsed as JSONC) and extracting `baseURL`/`base_url` fields from provider options that contain "localhost". Falls back to `http://localhost:8080`. + +### Slot Polling + +Every 500ms, the plugin polls `GET /slots?model=` on each discovered server. The model parameter is required by the `/slots` endpoint. Model discovery is attempted from the current route's session, falling back to an empty model parameter. + +### State Classification + +Each slot is classified as prefill or generation based on the `n_decoded` counter in `next_token[0]`. The plugin tracks a per-slot baseline value: + +1. When a slot first appears as processing, the current `n_decoded` is recorded as the baseline with `hasIncreased = false`. +2. If `n_decoded <= baseline` and `hasIncreased` is false, the slot is classified as prefilling. +3. If `n_decoded > baseline`, `hasIncreased` is set to true and the slot is classified as generating. + +This approach handles the case where `n_decoded` drops when a new request starts on a reused slot, and prevents generation stalls (where `n_decoded` plateaus) from being misclassified as prefill. + +When no slots are processing, all tracked state for those slots is cleared. + +### Prefill Rate (PP) + +During prefill, the plugin calculates the instantaneous prompt processing rate: + +1. On first detection of a prefill slot, the current `n_prompt_tokens` is captured as the baseline. +2. On subsequent polls, the delta in `n_prompt_tokens` is divided by the elapsed time in seconds. +3. The rate is updated only when both `dt > 0` and `delta > 0`. + +The per-slot `n_prompt_tokens` field is used instead of the global `llamacpp:prompt_tokens_total` from `/metrics` because the global counter includes tokens from all slots, producing inflated values when multiple slots are active simultaneously. + +### Generation Rate (TG) + +During generation, the plugin calculates the instantaneous token generation rate: + +1. On first detection of a generation slot, the current `n_decoded` is captured as the baseline. +2. On subsequent polls (same slot ID), the delta in `n_decoded` is divided by the elapsed time in seconds. +3. The rate is updated only when both `dt > 0` and `delta > 0`. + +Slot reuse is tracked via `generateSlotId` to detect when a new generation starts on a different slot. + +## Limitations + +### Progress Percentage + +The plugin cannot display prefill progress percentage. The `/slots` endpoint returns `n_prompt_tokens` (current prompt size) and `n_prompt_tokens_processed` (tokens processed), but not the final prompt size (`task->n_tokens()` from llama.cpp). Progress requires the ratio `n_prompt_tokens_processed / task->n_tokens()`. + +### What Would Improve Compatibility + +The following changes to the `/slots` endpoint would improve the plugin's functionality: + +1. **Expose final prompt size**: Add `n_prompt_tokens_total` (or `n_tokens`) to the `/slots` output, representing `task->n_tokens()` from llama.cpp. This would enable prefill progress percentage calculation as `(n_prompt_tokens_processed / n_prompt_tokens_total) * 100`. + +2. **Per-slot metrics endpoints**: Currently, the `/metrics` endpoint provides only global counters (`llamacpp:prompt_tokens_total`, `llamacpp:prompt_tokens_seconds`). Per-slot metrics would allow independent rate tracking without relying on slot state classification. + +3. **Slot transition notifications**: The plugin polls every 500ms to detect state transitions. A WebSocket or SSE-based notification system for slot state changes would reduce polling overhead and improve detection latency. + +4. **Stall detection**: When generation stalls (e.g., due to context window limits), `n_decoded` remains constant while `n_remain` stops decreasing. The plugin detects this via zero delta but has no way to distinguish a stall from normal generation. An explicit `stalled` flag in the slot output would help. + +5. **Model-agnostic slot data**: The `/slots` endpoint requires a model parameter. Returning all slots without model filtering, or supporting `*` as a wildcard, would simplify discovery when multiple models are loaded. ## Installation -Install from the CLI: - ```bash -opencode plugin oc-tps@latest --global +opencode plugin @troed/oc-ls-stats@latest --global ``` Requires `opencode` `1.3.14` or newer. + +TUI plugins are loaded from `~/.config/opencode/tui.json`: + +```json +{"plugin": ["@troed/oc-ls-stats@latest"]} +``` + +## Debug Logging + +Debug logging is controlled by the `DEBUG_ENABLED` constant in `tui.tsx`. When enabled, full slot state data is written to `/tmp/oc-ls-stats-debug.log` on every poll. + +## License + +Inherited from the original fork.