70 Best Open-Source Observability & Evaluation in 2026 (Ranked by GitHub Activity)
Monitoring, tracing, and testing infrastructure for running AI agents reliably in production
Top picks right now: OmniRoute, LiteLLM, ragflow. Every tool below is open source and ranked by live GitHub activity (stars, 30-day star growth, commits and releases), refreshed daily.
70 tools
OmniRoute
open-sourceOpenAI-compatible gateway for multi-provider routing, retries, fallbacks, caching, and observability
LiteLLM
freeOpen-source Python SDK and AI gateway for calling 100+ LLMs through a unified OpenAI-compatible interface
ragflow
open-sourceOpen-source RAG engine combining knowledge retrieval and agent capabilities for LLMs
Langfuse
open-sourceOpen-source LLM engineering platform for observability, evaluation, prompt and dataset management
Mastra
freeFrom the team behind Gatsby, Mastra is a framework for building AI-powered applications and agents with a modern TypeScript stack.
MinerU
freeTransforms complex documents like PDFs into LLM-ready markdown/JSON for your Agentic workflows.
Bifrost AI Gateway
open-sourceFastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 Β΅s overhead at 5k RPS.
Promptfoo
open-sourceOpen-source CLI and library for evaluating and red-teaming prompts, agents, RAG systems, and LLM apps
World Monitor
open-sourceAI-powered dashboard for real-time news aggregation, geopolitical monitoring, and infrastructure tracking
SkillSpector
open-sourceScans AI agent skills for prompt injection, data exfiltration, supply-chain risks, and vulnerabilities
Opik
open-sourceDebug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Firecrawl
freeπ₯ The Web Data API for AI - Turn entire websites into LLM-ready markdown or structured data
DeepEval
open-sourceThe LLM Evaluation Framework
phoenix
freeAI Observability & Evaluation
Manifest
open-sourceSmart LLM Routing for OpenClaw. Cut Costs up to 70% π¦π¦
langwatch
freeThe platform for LLM evaluations and AI agent testing
iFixAi
open-sourceAudits AI agents against business KPIs and organizational requirements with AβF scorecards
MLflow
open-sourceOpen-source AI engineering platform for agents, LLMs, and ML models
Weaviate
open-sourceOpen-source cloud-native vector database for semantic search, filtering, RAG, and reranking
Haystack
open-sourceOpen-source AI orchestration framework for modular RAG pipelines and agent workflows
Agenta
freeThe open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.
Netdata
open-sourceThe fastest path to AI-powered full stack observability, even for lean teams.
Prefect
open-sourcePrefect is a workflow orchestration framework for building resilient data pipelines in Python.
voltagent
open-sourceAI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework
DocsGPT
open-sourcePrivate AI platform for agents, assistants and enterprise search. Built-in Agent Builder, Deep research, Document analysis, Multi-model support, and API connectivity for agents.
OpenLIT
open-sourceOpen-source platform for AI agent tracing, evaluations, guardrails, prompts, and GPU monitoring
unstructured
open-sourceOpen-source ETL for converting documents into structured data for language models
txtai
open-sourceπ‘ All-in-one AI framework for semantic search, LLM orchestration and language model workflows
Scrapegraph-ai
open-sourcePython scraper based on AI
OpenLLMetry
open-sourceOpen-source observability for your GenAI or LLM application, based on OpenTelemetry
oumi
open-sourceEasily fine-tune, evaluate and deploy gpt-oss, Qwen3, DeepSeek-R1, or any open source LLM / VLM!
WFGY
freeWFGY is an open-source AI Troubleshooting Atlas for RAG, agents, and real-world AI workflows. Includes the 16-problem map, Global Debug Card, and WFGY 3.0. β Star to help more builders find this repo.
Langroid
open-sourceHarness LLMs with Multi-Agent Programming
garak
open-sourcethe LLM vulnerability scanner
UQLM
open-sourceUQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection
ChainForge
open-sourceAn open-source visual programming environment for battle-testing prompts to LLMs.
helicone
open-sourceπ§ Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 π
STORM
open-sourceAn LLM-powered knowledge curation system that researches a topic and generates a full-length report with citations.
llmware
open-sourceUnified framework for building enterprise RAG pipelines with small, specialized models
Ragas
open-sourceSupercharge Your LLM Application Evaluations π
LLM-eval-survey
freeThe official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".
TensorZero
open-sourceTensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
LangFair
freeLangFair is a Python library for conducting use-case level LLM bias and fairness assessments
Gitingest
open-sourceReplace 'hub' with 'ingest' in any GitHub URL to get a prompt-friendly extract of a codebase
OpenAI Evals
freeEvals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Pezzo
open-sourceπΉοΈ Open-source, developer-first LLMOps platform designed to streamline prompt design, version management, instant delivery, collaboration, troubleshooting, observability and more.
bRAG-langchain
freeEverything you need to know to build your own RAG application
llama-github
open-sourceOpen-source Python library for agentic RAG across public GitHub code, issues, and repository information
Langchain-Chatchat
open-sourceOffline-deployable Chinese knowledge base Q&A with RAG and agents using LangChain and open-source LLMs
Vanna
open-sourceπ€ Chat with your SQL database π. Accurate Text-to-SQL Generation via LLMs using Agentic Retrieval π.
AgentOps
open-sourcePython SDK for monitoring, cost tracking, benchmarking, and debugging AI agents
AgentBench
open-sourceA Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Auto-evaluator
freeEvaluation tool for LLM QA chains
Gorilla
open-sourceGorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
R2R
open-sourceSoTA production-ready AI retrieval system. Agentic Retrieval-Augmented Generation (RAG) with a RESTful API.
Verba
open-sourceRetrieval Augmented Generation (RAG) chatbot powered by Weaviate
text-extract-api
open-sourceLocal FastAPI for OCR extraction and PII removal from images, PDFs and Office files to Markdown or JSON
FastChat
open-sourceAn open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
clip-retrieval
open-sourceEasily compute clip embeddings and build a clip retrieval system with them
VisionAgent
open-sourceThis tool has been deprecated. Use Agentic Document Extraction instead.
TaskingAI
open-sourceThe open source platform for AI-native application development.
UpTrain
open-sourceOpen-source platform to evaluate and improve generative AI applications with 20+ preconfigured evaluations
Pathway
open-sourceReady-to-deploy templates for RAG and enterprise search that sync with live data sources
LangKit
open-sourceOpen-source text metrics toolkit for monitoring language models through input and output signals
LLM Comparator
open-sourceLLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.
Swiss Army Llama
freeA FastAPI service for semantic text search using precomputed embeddings and advanced similarity measures, with built-in support for various file types through textract.
Canopy
open-sourceRetrieval Augmented Generation (RAG) framework and context engine powered by Pinecone
Banana-lyzer
open-sourceOpen source AI Agent evaluation framework for web tasks ππ
GPTDiscord
open-sourceA robust, all-in-one GPT interface for Discord. ChatGPT-style conversations, image generation, AI-moderation, custom indexes/knowledgebase, youtube summarizer, and more!
Repochat
open-sourceChatbot assistant enabling GitHub repository interaction using LLMs with Retrieval Augmented Generation