X
Xberg
Rust document intelligence engine for extracting text, tables, metadata, images, and structured data
open-sourcetool-integration
9.4k
Stars
+60
Stars/month
3103
Commits (90d)
10
Releases (6m)
Star Growth
+180 (2.0%)estimated from velocity
Overview
Xberg extracts text, metadata, images, tables, and structured data from 106 formats across 140 file extensions. It includes code intelligence for 371 languages and offers 15 language bindings, CLI, REST API, and an MCP server.
Deep Analysis
Key Differentiator
Single engine handling format detection, reading, OCR, and extraction for 107 formats without requiring pipeline assembly.
⚡ Capabilities
- • Extract text/metadata/images/tables from 107 formats
- • OCR on demand with multiple backends
- • Audio/video transcription
- • Code intelligence for 371 languages
- • Embeddings and search
- • MCP server integration
🔗 Integrations
15 language bindings (Rust, Python, Node.js, Go, etc.)CLI toolREST APIMCP server
✓ Best For
- ✓ Agent pipelines needing structured document extraction
- ✓ RAG systems requiring clean text and table extraction
- ✓ Multi-format data ingestion for AI agents
✗ Not Ideal For
- ✗ End-user chatbot interfaces
- ✗ Generic AI content generation
- ✗ Image generation applications
⚠ Known Limitations
- ⚠ Some capabilities like URL ingestion require specific features
Alternatives
u
unstructured
Open-source ETL for converting documents into structured data for language models
D
Docling
Get your documents ready for gen AI
M
MegaParse
File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.
t
text-extract-api
Local FastAPI for OCR extraction and PII removal from images, PDFs and Office files to Markdown or JSON
Compare Xberg
Maintain Xberg?
Show your live rank in your README, or put Xberg in front of every visitor to AgentoolRank.