Anionex/agent-vision-toolkit
agent-vision-toolkit gives text-only LLM agents vision capabilities without requiring a multimodal model. It provides shell CLI tools (glance for image Q&A, OCR for long screenshots, UI restoration, GUI automation) plus an optional seamless integration that transparently proxies image tool calls so pasted images work directly in DSH. The skill teaches agents when and how to use each vision tool. Also supports Codex, Claude Code, Pi, Oh My Pi, and OpenCode. Note: for native DSH web integration with one-command install and built-in free Gemini vision quota, see the sister project Anionex/dsh-vision-toolkit.
Install
[not a dsh web plugin — shell-based toolkit]agent-vision-toolkit is a shell CLI toolkit, not a DSH web plugin. Easiest install: send your agent this prompt — "Follow the instructions in https://github.com/Anionex/agent-vision-toolkit to install the vision toolkit and skill locally." For Codex users: npx skills add Anionex/agent-vision-toolkit --skill vision-skills -a codex -g --copy -y. For DSH users, see the Seamless Integration section in the README. Note: the DSH-native version is Anionex/dsh-vision-toolkit (separate repo, installs with: dsh plugin --profile web add @anionex/dsh-vision-toolkit).
Compatibility
Works as agent skill (bin/ shell scripts) and drop-in for: Codex, Claude Code, Pi, Oh My Pi, OpenCode, and DSH. Extensions directory available for additional integrations.
Details
- Repo: Anionex/agent-vision-toolkit
- Category: Agent Capabilities
- Stars: 1084
- Last push: 2026-08-27
- First seen: 2026-08-01
Recent updates
See GitHub releases. Companion repo: Anionex/dsh-vision-toolkit for DSH-specific distribution.
FAQ
- How do I add vision to my DSH agent with agent-vision-toolkit?
- Install as a skill: dsh plugin --profile web add github:Anionex/agent-vision-toolkit (verify command in README). The toolkit registers vision tools the agent can call for any image task.
- What is the difference between agent-vision-toolkit and ModLens?
- Both add image understanding to text-only models. ModLens focuses on paste-to-chat image input within DSH web. agent-vision-toolkit is a broader skill covering OCR, GUI automation, and UI restoration across multiple agent frameworks.
Alternatives
liustack/modlens · Anionex/dsh-vision-toolkit · omdsh-dev/dsh-genui