Agent Vision Toolbox
wangyang10/image-vision · skills
image-vision allows models without visual capabilities (such as pure text models such as DeepSeek) to "see images": it automatically calls OpenAI-compatible image recognition model APIs (OpenRouter, SiliconFlow, Zhipu, Kimi, Tongyi Qianwen, local Ollama, etc.) to complete tasks such as image description, question and answer, and OCR.
★ 8GitHub stars
0Forks
2026-08-14Last updated
JavaScriptLanguage
MITLicense
Key features
- When the current model does not have visual capabilities, the image recognition API is automatically called to read images.
- One warehouse, multiple Agent hosts: the same set of skill logic is provided to Codex, Claude Code, DeepSeek Harness (DSH), etc. in the form of standard skills and host plug-ins.
- skills/image-vision/scripts/vision_query.py (Python) and dsh-image-vision/lib/ (JS porting) are two implementations of the same set of logic. When changing the preprocessing/API behavior, you need to synchronize both sides.
- Support local images, http(s) URL, data URL
Install command
dsh plugin --profile web add github:wangyang10/image-vision