— · Apache-2.0
Agent Skills for AceDataCloud AI services — music, image, video generation, LLM chat, web search. Compatible with Claude Code, GitHub Copilot, Gemini CLI, OpenAI Codex, and 30+ AI coding agents.
v0.5.2 · MIT
Animated and video live wallpaper for the DeepSeek Harness Web GUI — GIF, animated WebP/APNG and static PNG/JPEG images, plus MP4/WebM video backgrounds, with auto format detection, a dim slider, and UI chrome tone matching.
v3.9.0 · AGPL-3.0
Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Claude Code skill, Reddit-aware, refreshable share links.
— · MIT
Android Full-Stack Device Control Platform: WebRTC/H.264 remote desktop, UI/OCR/image-matching automation, one-click MITM, built-in Frida, proxy/VPN/frp/P2P networking, MCP/Agent, 160+ APIs, designed for multi-device clusters and engineered deployments.
v0.2.4 · no license
向模型暴露 MinerU 文档解析工具,将 PDF/图片/DOCX/PPTX/XLSX 转为结构化 Markdown/JSON | Exposes MinerU document-parsing tools to the model, converting PDF/images/DOCX/PPTX/XLSX into structured Markdown/JSON
— · MIT
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
— · MIT
Precision PPT design skill for OpenCode/Claude Code/Codex, with 40,000+ styles, pixel-perfect build-mode control, and AI image generation
v1.0.0 · no license
AI-powered ECG analysis: submit ECG images and receive diagnosis reports from ARPI's ECG AI.
v3.0.1 · MIT
Agent-first, provider-neutral multimodal OCR CLI for images, PDFs, URLs, JSON schemas, and agentic extraction with Gemini, Kimi, Muse, and OpenRouter.
v0.2.0 · MIT
Enables AI assistants to control Adobe InDesign on Windows via the Model Context Protocol, allowing document inspection, text and style editing, and exports to formats like PDF, images, and EPUB.
v0.17.1 · MIT
AI coding assistant skill (Claude Code, Codex, Gemini CLI, Kimi Code, GitHub Copilot CLI, Aider, OpenCode, OpenClaw) - turn any folder of code, docs, papers, images, or audio/video transcripts into a queryable knowledge graph
— · MIT
Enables AI agents to generate images on a local Stable Diffusion Forge Neo instance, automatically inferring prompt style and sampling parameters from the user's setup and past generations, and providing tools for LoRA search, model management, and module checks.
— · MIT
Unreal Engine plugin for LLM/GenAI models & MCP UE5 server. OpenAI GPT-5, Deepseek R1, Claude Opus/Sonnet, Gemini 3, Grok 4, Alibaba Qwen, Kimi, ElevenLabs TTS, Inworld, OpenRouter, Groq, GLM, Ollama, Local, Meshy, Tripo, Hunyuan3D, Rodin, fal, Dashscope, Seedream. NPC AI, agentic, chat, 3D gen, TTS, multimodal, image gen. UnrealMCP/UnrealClaude
— · MIT License
Enables agents to generate and edit images, sprites, icons, and animated sprite sheets via Gemini's Nano Banana model. It uses the existing Gemini plan's quota without per-image costs.
v1.0.0 · no license
Generate images, videos, voiceovers, and captions from a chat prompt.
v1.0.0 · no license
Generate vector art, vectorize images, and return SVG, PNG, and logo kits to AI agents.
v1.0.1 · no license
Extract text from documents, manipulate PDFs, and perform OCR on images.
v0.1.0 · no license
Filtrix MCP for image/video generation. Portal: https://agent.filtrix.ai/
v0.1.0 · no license
Node-based AI design workflows on the Gemus canvas: generate images, video, 3D, and PPT.
v0.6.0 · no license
Gency AI product image generation - MCP server for Claude, Codex, Cursor, VS Code