mirror of
https://git.sync.wtf/troed/oc-ls-stats.git
synced 2026-08-31 09:43:38 +03:00
docs: replace README with factual description of detection and calculation
This commit is contained in:
@@ -1,16 +1,97 @@
|
||||
# oc-tps
|
||||
# oc-ls-stats
|
||||
|
||||
Displays live TPS (tokens per second), average TPS, and average TTFT (time to first token) in the OpenCode session prompt.
|
||||
A TUI plugin for OpenCode that displays live prefill rate (PP) and generation rate (TG) from a llama.cpp-based llama-server.
|
||||
|
||||
## Display Format
|
||||
|
||||

|
||||
The plugin renders a single line in the session prompt right slot:
|
||||
|
||||
```
|
||||
1247 tps (PP) -- during prefill
|
||||
25 tps (TG) -- during generation
|
||||
- tps (TG) -- idle
|
||||
```
|
||||
|
||||
## Detection and Calculation
|
||||
|
||||
### Server Discovery
|
||||
|
||||
The plugin discovers the llama-server URL by reading the OpenCode configuration file (parsed as JSONC) and extracting `baseURL`/`base_url` fields from provider options that contain "localhost". Falls back to `http://localhost:8080`.
|
||||
|
||||
### Slot Polling
|
||||
|
||||
Every 500ms, the plugin polls `GET /slots?model=<model>` on each discovered server. The model parameter is required by the `/slots` endpoint. Model discovery is attempted from the current route's session, falling back to an empty model parameter.
|
||||
|
||||
### State Classification
|
||||
|
||||
Each slot is classified as prefill or generation based on the `n_decoded` counter in `next_token[0]`. The plugin tracks a per-slot baseline value:
|
||||
|
||||
1. When a slot first appears as processing, the current `n_decoded` is recorded as the baseline with `hasIncreased = false`.
|
||||
2. If `n_decoded <= baseline` and `hasIncreased` is false, the slot is classified as prefilling.
|
||||
3. If `n_decoded > baseline`, `hasIncreased` is set to true and the slot is classified as generating.
|
||||
|
||||
This approach handles the case where `n_decoded` drops when a new request starts on a reused slot, and prevents generation stalls (where `n_decoded` plateaus) from being misclassified as prefill.
|
||||
|
||||
When no slots are processing, all tracked state for those slots is cleared.
|
||||
|
||||
### Prefill Rate (PP)
|
||||
|
||||
During prefill, the plugin calculates the instantaneous prompt processing rate:
|
||||
|
||||
1. On first detection of a prefill slot, the current `n_prompt_tokens` is captured as the baseline.
|
||||
2. On subsequent polls, the delta in `n_prompt_tokens` is divided by the elapsed time in seconds.
|
||||
3. The rate is updated only when both `dt > 0` and `delta > 0`.
|
||||
|
||||
The per-slot `n_prompt_tokens` field is used instead of the global `llamacpp:prompt_tokens_total` from `/metrics` because the global counter includes tokens from all slots, producing inflated values when multiple slots are active simultaneously.
|
||||
|
||||
### Generation Rate (TG)
|
||||
|
||||
During generation, the plugin calculates the instantaneous token generation rate:
|
||||
|
||||
1. On first detection of a generation slot, the current `n_decoded` is captured as the baseline.
|
||||
2. On subsequent polls (same slot ID), the delta in `n_decoded` is divided by the elapsed time in seconds.
|
||||
3. The rate is updated only when both `dt > 0` and `delta > 0`.
|
||||
|
||||
Slot reuse is tracked via `generateSlotId` to detect when a new generation starts on a different slot.
|
||||
|
||||
## Limitations
|
||||
|
||||
### Progress Percentage
|
||||
|
||||
The plugin cannot display prefill progress percentage. The `/slots` endpoint returns `n_prompt_tokens` (current prompt size) and `n_prompt_tokens_processed` (tokens processed), but not the final prompt size (`task->n_tokens()` from llama.cpp). Progress requires the ratio `n_prompt_tokens_processed / task->n_tokens()`.
|
||||
|
||||
### What Would Improve Compatibility
|
||||
|
||||
The following changes to the `/slots` endpoint would improve the plugin's functionality:
|
||||
|
||||
1. **Expose final prompt size**: Add `n_prompt_tokens_total` (or `n_tokens`) to the `/slots` output, representing `task->n_tokens()` from llama.cpp. This would enable prefill progress percentage calculation as `(n_prompt_tokens_processed / n_prompt_tokens_total) * 100`.
|
||||
|
||||
2. **Per-slot metrics endpoints**: Currently, the `/metrics` endpoint provides only global counters (`llamacpp:prompt_tokens_total`, `llamacpp:prompt_tokens_seconds`). Per-slot metrics would allow independent rate tracking without relying on slot state classification.
|
||||
|
||||
3. **Slot transition notifications**: The plugin polls every 500ms to detect state transitions. A WebSocket or SSE-based notification system for slot state changes would reduce polling overhead and improve detection latency.
|
||||
|
||||
4. **Stall detection**: When generation stalls (e.g., due to context window limits), `n_decoded` remains constant while `n_remain` stops decreasing. The plugin detects this via zero delta but has no way to distinguish a stall from normal generation. An explicit `stalled` flag in the slot output would help.
|
||||
|
||||
5. **Model-agnostic slot data**: The `/slots` endpoint requires a model parameter. Returning all slots without model filtering, or supporting `*` as a wildcard, would simplify discovery when multiple models are loaded.
|
||||
|
||||
## Installation
|
||||
|
||||
Install from the CLI:
|
||||
|
||||
```bash
|
||||
opencode plugin oc-tps@latest --global
|
||||
opencode plugin @troed/oc-ls-stats@latest --global
|
||||
```
|
||||
|
||||
Requires `opencode` `1.3.14` or newer.
|
||||
|
||||
TUI plugins are loaded from `~/.config/opencode/tui.json`:
|
||||
|
||||
```json
|
||||
{"plugin": ["@troed/oc-ls-stats@latest"]}
|
||||
```
|
||||
|
||||
## Debug Logging
|
||||
|
||||
Debug logging is controlled by the `DEBUG_ENABLED` constant in `tui.tsx`. When enabled, full slot state data is written to `/tmp/oc-ls-stats-debug.log` on every poll.
|
||||
|
||||
## License
|
||||
|
||||
Inherited from the original fork.
|
||||
|
||||
Reference in New Issue
Block a user