dsh-session-analyst
Agent 与会话 活跃维护

dsh-session-analyst

dmsobtl/dsh-session-analyst

自动统计Agent会话工具调用成功率、评估token使用效率、检测冗余调用行为,支持跨会话回归对比,可快速定位会话质量波动根因,优化Agent交互表现与资源消耗。

1
Stars 标星
0
Forks 分支
1
Watchers 关注
0
Open Issues
TypeScript
主要语言
None
开源协议
57 KB
仓库大小
1 个月前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:dmsobtl/dsh-session-analyst
git clone https://github.com/dmsobtl/dsh-session-analyst.git
git clone git@github.com:dmsobtl/dsh-session-analyst.git
README.md main

dsh-session-analyst

Session quality analysis plugin for DeepSeek Harness.

Gives the agent (and you) structured insight into session behavior: tool success rates, token efficiency, redundant calls, error patterns, and regression detection.

Install

dsh plugin add dsh-session-analyst

Or add to your cordis.patch.yml:

- id: session-analyst
  plugin: dsh-session-analyst
  config:
    redundantCallThreshold: 3
    excessiveStepThreshold: 10

Tools provided

analyze_session

Parse a session log file (.jsonl or compressed .jsonl.zstd) and return quality metrics.

Agent: I'll analyze the session from the last run.
→ analyze_session({ path: "~/.dsh/sessions/abc123/session.jsonl" })

Returns:
{
  "summary": {
    "totalTurns": 5,
    "totalSteps": 12,
    "totalToolCalls": 8,
    "totalErrors": 1,
    "successRate": 0.875,
    "avgStepsPerTurn": 2.4
  },
  "issues": [
    { "severity": "warning", "code": "REDUNDANT_TOOL_CALL", "message": "..." }
  ],
  "tokenStats": { "efficiency": 0.12, ... },
  "toolStats": { "byName": { "bash": { "count": 5, "errors": 1 }, ... } }
}

compare_sessions

Compare baseline vs current session to detect regressions.

Agent: Compare today's run against yesterday's baseline.
→ compare_sessions({ baseline: "./baseline.jsonl", current: "./today.jsonl" })

Returns:
{
  "verdict": "regressed",
  "regressions": [
    { "dimension": "Tool success rate", "baseline": "100%", "current": "75%", "changePercent": -25 }
  ],
  "delta": { "stepsDelta": +3, "errorsDelta": +2, "tokenDelta": +1500 }
}

Analysis dimensions

Dimension What it detects
Tool success rate Percentage of tool calls that return without error
Redundant calls Same tool + same arguments called multiple times
Token efficiency Ratio of output tokens to total consumed
Excessive steps Turns with >10 steps (possible loop)
Error patterns Tools with >50% error rate
Duration Wall-clock time per turn

Use cases

  • Post-run diagnostics: Agent analyzes its own session after a task to identify inefficiencies
  • Regression detection: Compare sessions before/after a prompt or skill change
  • CI integration: Headless mode runs a task, then analyze_session checks quality gates
  • Skill tuning: Identify which tools are being misused and refine system prompts

Standalone usage (without dsh)

The parser and analyzer are usable as a library:

import { parseSessionFile, analyzeSession, compareSessions } from 'dsh-session-analyst'

const session = await parseSessionFile('./session.jsonl')
const analysis = analyzeSession(session)
console.log(analysis.summary)

Development

npm install
npm test

License

MIT