opendatalab/MinerU
GitHubAn open-source document parsing tool that converts PDFs, images, DOCX, PPTX, and XLSX files into machine-readable Markdown and JSON formats for downstream RAG and processing pipelines.
AI developers preparing high-quality, clean text datasets for pre-training large language models or building technical RAG workflows.
The default high-performance parsing level (effort=medium) does not support image analysis, requiring a switch to a slower mode (effort=high) if images must be parsed.
What matters before you adopt it
The layout destruction, semantic breakage, and formula/table formatting loss that typically occurs when ingesting complex multi-column documents or scanned PDFs into machine-readable training data.
Security evidence without the noise
Trust remains a decision signal; CVEs and scanner evidence explain what is driving the risk.
View trust evidence & security findings▼
Architecture from code7 modules · 2 edges▼
Modules and dependency edges extracted from repository code. This is code evidence, not README inference.
Architecture evidence details
flowchart TD
%% mineru — high-level architecture (DRAFT, refine me)
n0["(root) · 1 file"]
n1["demo · 1 file"]
n2["mineru · 217 files"]
n3["tests · tests · 5 files"]
n1 --> n2
n3 --> n2
class n3 test
classDef test fill:#499894,color:#ffffff,stroke:#397975Evidence, security & integrations▼
Adoption guidance▼
How it works & getting started▼
Nearby repositories worth comparing before adoption.
Structured data gathering from any website using AI-powered scraper, crawler, and browser automation. Scraping and crawling with natural language prompts. Equip your LLM agents with fresh data. AI Studio python SDK for intelligent web data gathering.
holaOS คือ Agentic OS ที่ทำให้ AI Agent สามารถทำงานดิจิทัลได้ทุกอย่างบนคอมพิวเตอร์ของคุณ ไม่ว่าจะเป็นการจัดการไฟล์ ท่องเว็บ หรือรันโปรแกรมต่างๆ ผ่าน Natural Language ด้วย Electron + TypeScript + MCP Protocol ที่ Open Source สำหรับ Developer ที่อยากสร้าง Desktop AI Agent ของตัวเอง
TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed, shared, and equipped across agents and frameworks.