dsh-vision-plugin
Agent 与会话 活跃维护

dsh-vision-plugin

Zhangbo-cn/dsh-vision-plugin

轻量级视觉能力扩展插件,可给纯文本模型新增图像理解能力,通过路由调度多模态视觉模型完成图像内容解析,无需改动原有模型架构,开箱即用即可让文本模型具备图像识别与分析能力。

0
Stars 标星
0
Forks 分支
0
Watchers 关注
0
Open Issues
TypeScript
主要语言
MIT
开源协议
43 KB
仓库大小
1 个月前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:Zhangbo-cn/dsh-vision-plugin
git clone https://github.com/Zhangbo-cn/dsh-vision-plugin.git
git clone git@github.com:Zhangbo-cn/dsh-vision-plugin.git
README.md main

dsh-vision-plugin

Vision capability for DeepSeek Harness: lets a text-only model "understand" an image by routing to an external OpenAI-compatible multimodal API. Built as a standalone dsh-plugin from the official vision capability seam proposal.

Packages

Package Role
@zhangbo-cn/dsh-vision Service Definition: ctx.vision (registerAdapter, describe, listProviders)
@zhangbo-cn/dsh-vision-openai-compatible Provider: OpenAI-compatible chat-completions adapter
@zhangbo-cn/dsh-tool-vision Consumer: view_image tool

Install

pnpm add @zhangbo-cn/dsh-vision @zhangbo-cn/dsh-vision-openai-compatible @zhangbo-cn/dsh-tool-vision

Mount in your cordis.yml:

- id: vision
  name: '@zhangbo-cn/dsh-vision'

- id: vision-openai-compatible
  name: '@zhangbo-cn/dsh-vision-openai-compatible'
  config:
    baseURL: 'https://api.example.com/v1'   # required at request time
    model: 'gpt-4o'                          # required at request time
    apiKeyEnv: 'OPENAI_API_KEY'              # env var holding the key

- id: tool-vision
  name: '@zhangbo-cn/dsh-tool-vision'

Then ask the model: "use view_image to look at ./screenshot.png" — it reads the file, commits the bytes through the attachment seam, and returns a text description from your configured vision model.

How it works

view_image(file_path, prompt)
  → ctx.fs reads the image bytes
  → attachments.saveImage (durable, content-addressed)
  → ctx.vision.describe({ ref, prompt })
      → vision provider posts a data:image/...;base64 image_url to /chat/completions
      → returns text

Image input reuses the durable ImageAttachmentRef from the attachment seam; output is text (no ImageBlock), so it is independent of whether the harness LLM route itself accepts images.

Requirements

  • DeepSeek Harness with an attachment store (dsh-attachment-local) and filesystem (dsh-fs-local).
  • A configured OpenAI-compatible multimodal endpoint (any OpenAI-chat-completions-compatible vision model).

Development

npm install
npm run build -ws
npx vitest run

Tests include a real Loader composition booting the plugin with a fake in-memory vision provider (52 tests).

License

MIT