v0.51.57 · MIT
Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in natural language on ANY LLM (Claude, ChatGPT, Gemini, offline Ollama, or any hosted model). 178 tools, 36 AI skills, 55 installer packs. Local, LAN, VPS, or Comfy Cloud.
v0.1.39 · MIT
[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
v0.13.0 · MIT
MCP server for AI image generation
v0.13.0 · MIT
MCP server for AI image generation
v0.8.1 · MIT
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.
v1.2.0 · MIT
Adobe Premiere Pro MCP. 282 tools for AI-driven video editing via MCP, for Codex, Claude, and other MCP clients.
— · MIT
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
v1.13.0 · MIT
MCP server for Adobe Premiere Pro with 288 core AI video editing tools and a capability-aware UXP bridge for supported Premiere workflows.
v1.81.9 · MIT
Kolbo AI MCP Server - Generate images, videos, music, speech, and sound effects from Claude Code
v0.1.0-alpha.4.19 · Apache-2.0
Use openai-codex models and image generation via ChatGPT OAuth for DeepSeek Harness.
v3.1.1 · MIT
DSH plugin: pixel-to-text image reading for text-only models. image_scan/image_ocr/image_sample tools + image-reading skill (34-image trained methodology). Pure local, optional PaddleOCR.
— · MIT
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
— · MIT
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
v1.0.1 · MIT License
Enables generating game-ready 3D models, textures, and audio from natural language by proxying to the remote Gessa API.
v0.4.0 · MIT
Give text-only models eyes: an analyze_image tool for DeepSeek Harness, backed by free Chinese vision APIs (GLM-4V-Flash / Qwen-VL) or any OpenAI-compatible vision endpoint. 给纯文本模型装上眼睛。
— · MIT
Unreal Engine plugin for LLM/GenAI models & MCP UE5 server. OpenAI GPT-5, Deepseek R1, Claude Opus/Sonnet, Gemini 3, Grok 4, Alibaba Qwen, Kimi, ElevenLabs TTS, Inworld, OpenRouter, Groq, GLM, Ollama, Local, Meshy, Tripo, Hunyuan3D, Rodin, fal, Dashscope, Seedream. NPC AI, agentic, chat, 3D gen, TTS, multimodal, image gen. UnrealMCP/UnrealClaude
— · MIT License
Enables agents to generate and edit images, sprites, icons, and animated sprite sheets via Gemini's Nano Banana model. It uses the existing Gemini plan's quota without per-image costs.
— · no license
Atlas Cloud skills for Claude Code, Codex & Gemini CLI — generate images/videos and call 300+ AI models from your coding agent.
— · Apache-2.0
lightweight Python-based MCP (Model Context Protocol) server for local ComfyUI