dsh-eye-vision
文件与数据 活跃维护

dsh-eye-vision

AlloyPlane/dsh-eye-vision

OpenAI兼容多模态接口,为纯文本模型扩展图像理解、OCR与UI分析能力,自动识别图片中的文字与界面元素,输出结构化结果,无需更换模型即可处理图片输入,集成现有工作流。

1
Stars 标星
0
Forks 分支
1
Watchers 关注
0
Open Issues
JavaScript
主要语言
MIT
开源协议
18 KB
仓库大小
1 个月前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:AlloyPlane/dsh-eye-vision
git clone https://github.com/AlloyPlane/dsh-eye-vision.git
git clone git@github.com:AlloyPlane/dsh-eye-vision.git
README.md main

dsh-eye-vision

Give text-only DeepSeek Harness models eyes — image understanding, OCR, and UI analysis through any OpenAI-compatible multimodal API.

Fork of dsh-free-vision (MIT) v1.0.1, with:

  • custom-provider fixCUSTOM_MODEL_NAME is now forwarded, so arbitrary OpenAI-compatible endpoints (GPT-4o, Qwen-VL, GLM-4V, youtu-vita, vLLM, Ollama…) actually start
  • allowed-directories whitelist — the vision engine can read images from your configured workspace roots, not just the engine CWD and home directory

How it works

You send an image path
  → agent calls the image_understand tool
    → plugin spawns the luma-mcp vision engine (child process)
      → engine calls your multimodal API
        → text description comes back to the text-only model

The main model never needs image input support. Paste the image path, get answers.

Features

  • image_understand tool registered on ctx.tools, visible to every session in the profile
  • Any OpenAI-compatible endpoint via custom provider — bring your own multimodal API
  • Free-tier providers built in: qwen (Qwen3-VL-Flash), volcengine (Doubao), siliconflow (DeepSeek-OCR)
  • Multi-crop for large images (detail preservation)
  • Allowed-directories whitelist (LUMA_ALLOWED_DIRS) — read images from your workspace
  • Proxy vars stripped for direct mainland-China API access
  • Live settings — save via the settings route, no restart needed

Installation

# from a local checkout
dsh plugin --profile web add D:/xd/dsh-eye-vision

# once published
dsh plugin --profile web add dsh-eye-vision

Restart dsh web. The tool appears as image_understand.

Configuration

Settings file: ~/.dsh/free-vision.json (same path as upstream for drop-in compatibility):

{
  "modelProvider": "custom",
  "baseURLs": { "custom": "https://your-api.example.com/v1" },
  "modelName": "your-vision-model",
  "apiKey": "sk-...",
  "allowedDirs": "D:/workspace",
  "toolName": "image_understand"
}

Or use environment variables (fallback chain: settings file > env):

Provider Key env Base URL env
custom CUSTOM_API_KEY CUSTOM_BASE_URL + CUSTOM_MODEL_NAME
qwen DASHSCOPE_API_KEY QWEN_BASE_URL
volcengine VOLCENGINE_API_KEY VOLCENGINE_BASE_URL
siliconflow SILICONFLOW_API_KEY SILICONFLOW_BASE_URL
zhipu ZHIPU_API_KEY ZHIPU_BASE_URL
hunyuan HUNYUAN_API_KEY HUNYUAN_BASE_URL

allowedDirs: semicolon/comma-separated extra roots the engine may read images from (default: engine CWD + home directory).

Usage

看图:D:/path/to/screenshot.png
OCR:D:/path/to/document.png
UI 分析:D:/path/to/design.png (task_type: ui)

Tool arguments: image_source (local path / http(s) URL / data URI), prompt, task_type (auto|general|ocr|ui|debug|describe). PNG/JPG/WebP/GIF up to ~10 MB.

Security

  • API keys live only in the settings file or environment variables — never in this repo
  • Images are sent only to the endpoint you configured
  • Proxy environment variables are deliberately stripped from the engine child process
  • A gitguard-style pre-push scan is recommended before publishing forks

Development

cd dsh-eye-vision
pnpm install   # installs luma-mcp engine + MCP SDK
pnpm test

The engine patches in scripts/patch-luma.mjs re-apply automatically on install (idempotent, pinned to luma-mcp 1.7.1).

License

MIT — see LICENSE. Upstream: dsh-free-vision (MIT) by FuzzySoul.