dsh-rate-limiter
其他 活跃维护

dsh-rate-limiter

Xidong-AI/dsh-rate-limiter

轻量级主动限流插件,可实时拦截超额请求,支持自定义配额规则与动态阈值调整,部署无额外依赖,有效避免接口过载风险,保障服务稳定运行。

1
Stars 标星
0
Forks 分支
1
Watchers 关注
0
Open Issues
JavaScript
主要语言
MIT
开源协议
280 KB
仓库大小
1 个月前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:Xidong-AI/dsh-rate-limiter
git clone https://github.com/Xidong-AI/dsh-rate-limiter.git
git clone git@github.com:Xidong-AI/dsh-rate-limiter.git
README.md master

dsh-rate-limiter

English | 中文

A proactive rate limiter plugin for DeepSeek Harness (dsh): it controls the request rate per provider (token bucket) before model requests are issued, and queues the request with a delay instead of failing when the limit is exceeded — avoiding upstream 429s.

It complements the official dsh-llm-retry (exponential backoff after failure): rate limiting comes first (prevention), backoff comes last (safety net); the two do not interfere with each other.

Features

  • Per-provider token bucket, enforced before the request is sent (proactive prevention)
  • Over-limit requests are queued with a delay instead of rejected (no 429s, no lost requests)
  • Unconfigured providers pass through untouched (zero intrusion)
  • Queued waits honor the abort signal: stopping the user interrupts the wait immediately
  • Hand-written reservation-based token bucket (concurrency-safe), zero third-party rate-limiting dependencies
  • Mounts on agent/request, coexists naturally with dsh-llm-retry

Installation

Install from npm:

dsh plugin --profile web add @xidong-ai/dsh-rate-limiter

npm registry URLs are case-sensitive; use the lowercase package name.

Or install directly from GitHub:

dsh plugin --profile web add github:Xidong-AI/dsh-rate-limiter

For local development, add the checkout directly:

dsh plugin --profile web add .

After installing, dsh --profile web --dump-config should show the plugin entry:

- id: rate-limiter
  name: @xidong-ai/dsh-rate-limiter
  config:
    enabled: true
    providers: {}

Configuration

Configure the token bucket per provider in the profile's cordis.patch.yml (or this plugin's cordis.patch.yml):

- id: rate-limiter
  config:
    enabled: true
    providers:
      nvidia:
        rate: 0.5        # tokens/second (long-term average QPS)
        burst: 1         # bucket capacity (allowed burst requests)
      sensenova:
        rate: 0.02778
        burst: 1
  • rate: refill rate (tokens/second), i.e. the long-term average request rate.
  • burst: bucket capacity, the number of burst requests allowed.
  • Providers not listed are not rate-limited; requests pass through untouched (zero intrusion).
  • enabled: false disables the plugin entirely.

How It Works

The plugin hooks onto the agent/request waterfall: it await next() first to obtain the call config (which carries the provider), then performs a per-provider token bucket check; when tokens are insufficient, it queues the request with a delay (interrupted immediately by the abort signal when the user stops), then returns the config unchanged — it never modifies request content, never changes routing, never swallows errors. It only controls when a request is issued.

The rate-limiting algorithm is a hand-written reservation-based token bucket (concurrency-safe), with zero third-party rate-limiting dependencies.

Relationship with dsh-llm-retry

Plugin Timing Behavior
dsh-rate-limiter Before the request is issued Queue with a delay when over the limit (prevents 429s)
dsh-llm-retry After the request fails Exponential backoff retry (safety net)

They mount at different points (agent/request vs agent/request-error) and coexist naturally.

Uninstall

dsh plugin --profile web remove @xidong-ai/dsh-rate-limiter

Development

npm install
npm run typecheck   # tsc --noEmit
npm run test        # vitest run
npm run build       # esbuild transpiles lib/*.ts → lib/*.js

Acknowledgements

Thanks to the Linux.do community for support.