v3.0.1 · MIT
Agent-first, provider-neutral multimodal OCR CLI for images, PDFs, URLs, JSON schemas, and agentic extraction with Gemini, Kimi, Muse, and OpenRouter.
v3.9.0 · AGPL-3.0
Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Claude Code skill, Reddit-aware, refreshable share links.
— · MIT
Convert documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection
— · AGPL-3.0
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows
— · Apache-2.0
天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取
— · no license
A list of open-source AI projects you can use to generate income easily.
— · MIT
An MCP server providing tools for image processing operations
— · MIT
Android Full-Stack Device Control Platform: WebRTC/H.264 remote desktop, UI/OCR/image-matching automation, one-click MITM, built-in Frida, proxy/VPN/frp/P2P networking, MCP/Agent, 160+ APIs, designed for multi-device clusters and engineered deployments.
v1.5.9 · MIT
clawdcursor compiles whatever's on screen into one UI map — accessibility tree and OCR fused into stable, addressable elements, with a screenshot only when needed — then drives apps through reusable scripts, verifying every action and routing it through a single safety gate.
v1.0.1 · no license
Extract text from documents, manipulate PDFs, and perform OCR on images.