dsh-pdf-reader
开发工具 活跃维护

dsh-pdf-reader

ralfsqual/dsh-pdf-reader

它是生态功能插件,依托pdfjs-dist实现本地PDF文本提取,无需上传文件即可快速解析PDF内容并输出结构化文本,可直接在模型对话中调用获取PDF文本信息。

0
Stars 标星
0
Forks 分支
0
Watchers 关注
0
Open Issues
JavaScript
主要语言
MIT
开源协议
63 KB
仓库大小
1 个月前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:ralfsqual/dsh-pdf-reader
git clone https://github.com/ralfsqual/dsh-pdf-reader.git
git clone git@github.com:ralfsqual/dsh-pdf-reader.git
README.md main

dsh-pdf-reader

A DeepSeek Harness plugin that adds a read_pdf tool, extracting text from PDF files page by page. Fully local parsing — no upload, no network access.

为 DeepSeek Harness 增加 read_pdf 工具,逐页提取 PDF 文本。纯本地解析,不上传、不联网。

Install / 安装

dsh plugin --profile web add github:ralfsqual/dsh-pdf-reader

Or install from a local checkout / 或本地安装:

dsh plugin --profile web add /path/to/dsh-pdf-reader

Restart dsh web afterwards. The read_pdf tool becomes available to the agent in new sessions.

重启 dsh web 后生效,新会话中 Agent 即可调用 read_pdf。

Usage / 使用

Tell the agent to read a PDF, or combine it with an @ file mention (e.g. via dsh-at-file):

Read @docs/spec.pdf and summarize the requirements.
读取 @docs/report.pdf 并总结要点。

Tool parameters

Parameter Type Description
path string PDF path, relative to the current working directory or an absolute path inside the workspace. / 相对当前工作目录或工作区内绝对路径
maxPages number Max pages to extract, default 500; 0 = no limit. / 最多提取页数
maxChars number Max characters to return, default 120000. / 最多返回字符数

How it works / 工作原理

  • Registered as a model tool via defineTool (@deepseek-ai/dsh-tools).
  • Parses the PDF with pdfjs-dist (2.6.347), extracting text per page with --- page N --- separators.
  • Output is a structured object (pages, extractedPages, truncated, empty, text) with a readable render.

Security / 安全

  • Workspace-confined: path must resolve inside the current working directory; escape attempts (e.g. ../) are rejected. / 路径限定工作区内,越界拒绝。
  • Read-only: the tool only reads the file, never writes, never executes commands. / 只读,绝不写入或执行命令。
  • Local only: parsing happens entirely on the host; no network requests are made. / 纯本地解析,无任何网络请求。
  • Password-protected, corrupted, or non-PDF files produce clear error messages.

Prerequisites / 环境要求

  • DeepSeek Harness with a web (or any agent-capable) profile.
  • Node.js >= 22.19.
  • Works on Windows, macOS, and Linux (host process only; no browser/native code).

Limitations / 已知限制

  • Extracts text layers only. Scanned/image-only PDFs contain no extractable text (the tool reports empty; OCR is out of scope). / 仅提取文本层,扫描件需 OCR。
  • Paths must not contain a leading @ when combined with @-mention plugins. / 与 @ 引用插件合用时路径不能以 @ 开头。
  • Very large PDFs are bounded by maxPages/maxChars; output is truncated with a truncated: true marker.

License

MIT — see LICENSE.