dsh-thinking-token-stat
开发工具 活跃维护

dsh-thinking-token-stat

Six6stRINgs/dsh-thinking-token-stat

轻量化插件,可在底部Dock栏与单次对话尾部自动统计展示模型思考token数量,无需额外配置即可快速接入,直观呈现模型思考的资源消耗情况。

1
Stars 标星
0
Forks 分支
1
Watchers 关注
0
Open Issues
JavaScript
主要语言
MIT
开源协议
2.1 MB
仓库大小
13 天前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:Six6stRINgs/dsh-thinking-token-stat
git clone https://github.com/Six6stRINgs/dsh-thinking-token-stat.git
git clone git@github.com:Six6stRINgs/dsh-thinking-token-stat.git
README.md main

dsh-thinking-token-stat

English | 中文

awesome · DSH plugin
dsh
license
npm version
npm downloads
repo
GitHub stars

A lightweight plugin that adds model thinking-token statistics to the bottom
Dock and the end of each conversation. Client-only and zero-cost when idle: it
reads the existing conversation snapshot and renders nothing at all when there
is no reported or visible thinking output.

Thinking statistics depend on the model and provider. Some models do not expose
reasoning tokens or reasoning text to the client; for those replies, the plugin
has no thinking data to count and the readout is intentionally not shown.

thinking-token statistics

It adds two readouts, both of which render nothing when the model has produced
no thinking
for the relevant scope:

Surface Slot Scope
Bottom composer dock conversation.composer.dock Whole session (cumulative)
Each assistant reply (timing row) conversation.chat.assistant-actions That single turn

The per-turn readout is appended to the reply's timing strip, alongside the
built-in usage and timing pills, so it reads as part of that same unit. Click the
thinking pill to open a native-style details panel. The panel follows the active
DSH locale.

The details panel shows the thinking-token count, the all-token denominator and
percentage, and the output-token denominator and percentage. Percentages use one
decimal place; counts use compact units such as K and M when appropriate.

How thinking tokens are counted

For every finalized assistant message, in order:

  1. Provider-reported — if usage.reasoningTokens is present (DeepSeek,
    OpenAI o-series, Anthropic, …), it is used as-is. This is exact.
  2. Otherwise — the reasoning blocks are priced with the harness' fixed
    heuristic (four characters per token), so a provider that returns reasoning
    text but no token count still gets a figure.
  3. Neither present → the message counts as zero thinking and contributes
    nothing.

The two denominators are derived from the same assistant messages, so the
percentages always agree with the count on one consistent scope:

  • share of all tokens = thinking ÷ (input + output + cache reads + cache writes)
  • share of output = thinking ÷ output

Why it's lightweight

  • Client-only, zero host behavior. The node half (lib/index.js) is an empty
    apply that exists only so the package appears in the host Loader. All work is
    done in the browser.
  • Read-only. It consumes the existing conversation snapshot; it adds no
    service, projection, tool, or RPC.
  • Self-contained. It depends on no other plugin.
  • Zero-cost when idle. When nothing is being thought, both readouts render
    null — no element, no layout cost.
  • Theme-aware. Colors come from the --dsw-alias-* tokens, so it follows
    light/dark automatically.

Install

The plugin follows the standard dsh.bundle + dsh.client convention, so it
installs like any DSH plugin.

From GitHub:

dsh plugin add github:Six6stRINgs/dsh-thinking-token-stat

Or install the published npm package:

npm install dsh-thinking-token-stat

After installing, make sure the package is included in your DSH profile bundles,
then restart dsh web and reload the page. The readouts appear only once the
model starts thinking.

Testing

test/harness.mjs stubs the browser/React environment, runs the real factory
and apply, and renders both entries against sample data (provider-reported and
block-estimated paths, the one-decimal formatting, and the hidden-when-empty
cases). Run with:

node test/harness.mjs

Known limitations

  • Figures come from the in-window conversation snapshot. For very long,
    paged sessions the visible window is the counted scope; the shipped stats
    line and projections use whole-log folds instead.
  • The character-per-token estimate is approximate, exactly as the harness'
    fixed heuristic is; provider-reported reasoningTokens is always preferred.