DeepSee

DeepSee

Give DeepSeek eyes — screenshots to exact text

1 follower

Vision models attend selectively and miss details. DeepSee is the opposite: it decodes any screenshot into exact text — every word with its pixel coordinates, colors, and spacing — so a text-only model can read your UI precisely. It's an open-source MCP server on GitHub that Claude Code, Hermes, and DeepSeek call automatically, or use it right in your browser. Zero VRAM, zero uploads, fully local. Same screenshot, same transcript, every time.