dshplugin.devDeepSeek Harness Plugins
DSH Plugin Vision Toolkit plugin logo
DeepSeek Harness Plugin

DSH Plugin Vision Toolkit

0
Published by YYTbit

Vision toolkit for DeepSeek Harness -- give text-only agents eyes

Developer Toolsdeepseek-harnessdeepseek-vldsh-pluginmultimodal

Get this plugin

Review the source, then continue to the publisher.

dsh plugin add dsh-plugin-vision-toolkit@latest
Get this plugin
Share on X ↗

About this plugin

Source snapshot 8/13/2026

dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.

What it does

Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.

Tools

  • glance -- describe, ask about, or OCR an image
  • ground -- locate a specific element (returns bounding box)
  • detect -- find all instances of an element kind
  • crop -- cut a region from an image

Install

dsh plugin --profile your-profile add dsh-plugin-vision-toolkit

Configuration

Set environment variables:

export VISION_API_KEY=sk-xxx           # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://...     # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2       # Vision model name

Usage examples

# Describe an image
glance screenshot.png

# Ask a question
glance screenshot.png -q "What error is shown?"

# OCR
glance screenshot.png --ocr

# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870

# Find all buttons
detect screenshot.png "buttons"

# Crop a region
crop screenshot.png 450,820,620,870 button.png

How it works

The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:

  1. Reads the image file
  2. Encodes it as base64
  3. Sends it to the vision API with a prompt
  4. Returns the text response

The agent never sees raw pixels -- it gets text descriptions it can reason about.

Supported vision providers

  • DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
  • OpenAI GPT-4V / GPT-4o
  • Any OpenAI-compatible multimodal endpoint

License

MIT -- YYTbit