dsh-delegate-router
开发工具 活跃维护

dsh-delegate-router

penguin-oo/dsh-delegate-router

自动完成子代理调用的Flash/Pro智能路由,轻量任务自动走低成本方案,重任务自动匹配高性能方案,支持/delegate参数手动覆盖路由规则,无需复杂配置即可平衡调用成本与处理效果。

0
Stars 标星
0
Forks 分支
0
Watchers 关注
0
Open Issues
JavaScript
主要语言
MIT
开源协议
248 KB
仓库大小
1 个月前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:penguin-oo/dsh-delegate-router
git clone https://github.com/penguin-oo/dsh-delegate-router.git
git clone git@github.com:penguin-oo/dsh-delegate-router.git
README.md main

dsh-delegate-router

English · 中文

Automatic Flash/Pro routing for DeepSeek Harness subagent calls.

Your main conversation keeps running on a strong model (e.g. V4 Pro). When the
agent delegates a task to a subagent, this plugin classifies the task and sends
light tasks to a cheap model (e.g. V4 Flash) while heavy tasks stay on the
strong model — deterministically, without relying on the model's cooperation.
Every decision is recorded in a ledger you can inspect from the sidebar
(⚡ 分派记录).

  • Auto routing — light tasks run on Flash, heavy tasks stay on Pro.
  • DIY rules — all rules live in ~/.dsh/dsh-delegate-router.json:
    keyword lists, the short-task threshold, the budget cap, and the Beijing
    peak-hour demotion.
  • Manual overrides — per-call provider/model parameters on the
    subagent / subagent_fork tools.
  • /delegate command — per-session mode: /delegate auto | off | flash-all.
  • Decision ledger — the ⚡ 分派记录 panel lists every routed delegation
    (task, route, trigger) for the active session.

Why

DeepSeek V4 Flash costs exactly one third of V4 Pro on every line (official
pricing, effective 2026-08-17; peak/off-peak only scales both). Stock DSH runs
every subagent on the parent's model — a search task delegated from a Pro
session costs the full Pro rate. Measured on real sessions: 4 routed runs,
~156K tokens, ¥0.32 actual (Flash) vs ¥0.96 hypothetical (Pro) — 66.7%
saved
on every routed task. Run node scripts/measure-savings.mjs against
your own ~/.dsh/sessions to reproduce.

Rules (in order)

  1. explicit per-call provider/model → used as-is (manual)
  2. /delegate off → inherit; /delegate flash-all → all Flash
  3. session tokens over budgetCapTokens → Flash (budget)
  4. keyword dominance scoring: heavy wins ties, but a strictly-light task
    beats one incidental heavy word (auto-heavy / auto-light)
  5. task text ≤ shortTaskMaxChars → Flash (auto-short)
  6. unmatched + Beijing peak hours (default 9–12, 14–18) → Flash (peak)
  7. unmatched + unknownToFlash: true (opt-in, aggressive) → Flash (auto-unknown)
  8. otherwise → inherit the parent model

Task text = the subagent call's description + prompt. Matching is precise:
pure-ASCII keywords use word boundaries (list never matches specialist,
design never matches designer); CJK keywords shorter than two characters
are ignored.

Works great with dsh-routing-suite

dsh-routing-suite owns the
thinking-mode / persona layer; this plugin owns the child-model cost
layer
. They stack: light subagent tasks get routed to Flash by this plugin,
then run under the router preset's flash-optimized persona. The routing suite's
own experiments (P11/P24) found the optimal weak persona is flash-specific
and that spec-style personas actively hurt Flash — so Flash + that router
preset is the best-matched combination for cheap delegation, and this plugin
supplies the automatic Flash routing for it.

Honest measurement

Relative prices are guaranteed by the official price table (Flash = 1/3 of Pro
on every line, 2026-08-17 peak/off-peak pricing). Absolute numbers depend on
how much work a run does — LLM runs are nondeterministic, so compare per
token
(or same-task), never raw totals. scripts/measure-savings.mjs prices
your real session logs with the official table; read it before quoting numbers.

Install

dsh plugin --profile web add dsh-delegate-router

Configure

All knobs are optional and live in ~/.dsh/dsh-delegate-router.json:

{
  "flashProvider": "opencode-go",
  "flashModel": "deepseek-v4-flash",
  "proProvider": "opencode-go",
  "proModel": "deepseek-v4-pro",
  "mode": "auto",
  "lightKeywords": ["search", "搜索", "查找", "总结", "summarize", "list", "列出"],
  "heavyKeywords": ["refactor", "重构", "implement", "实现", "debug", "调试"],
  "shortTaskMaxChars": 120,
  "peakDemoteUnknown": true,
  "unknownToFlash": false,
  "peakHours": [[9, 12], [14, 18]],
  "budgetCapTokens": 0
}
  • shortTaskMaxChars: 0 disables the short-task rule; peakDemoteUnknown: false disables peak-hour demotion; unknownToFlash: true sends ANY
    unmatched task to Flash (aggressive — leave false unless you are sure);
    budgetCapTokens: 0 disables the cap.
  • Providers can also come from DSH_DELEGATE_ROUTER_FLASH_PROVIDER /
    DSH_DELEGATE_ROUTER_FLASH_MODEL / DSH_DELEGATE_ROUTER_PRO_PROVIDER /
    DSH_DELEGATE_ROUTER_PRO_MODEL env vars.
  • Restart DSH after editing the file. Switch mode at runtime with
    /delegate <mode>.

Development

npm install
npm run smoke              # manifest sanity
node scripts/test-routing.mjs   # deterministic rule checks
node scripts/measure-savings.mjs # real savings over ~/.dsh/sessions
node scripts/e2e-panel-loop.mjs  # browser E2E against a test instance

License

MIT

Acknowledgements

Built for the DeepSeek Harness plugin ecosystem — thanks to the community on
LINUX DO for feedback and testing.