πŸ“‘

70 Best Open-Source Observability & Evaluation in 2026 (Ranked by GitHub Activity)

Monitoring, tracing, and testing infrastructure for running AI agents reliably in production

Top picks right now: OmniRoute, LiteLLM, ragflow. Every tool below is open source and ranked by live GitHub activity (stars, 30-day star growth, commits and releases), refreshed daily.

70 tools

OmniRoute

open-source

OpenAI-compatible gateway for multi-provider routing, retries, fallbacks, caching, and observability

⭐ 71.9k↑ 11269/moobservability-evaluation

LiteLLM

free

Open-source Python SDK and AI gateway for calling 100+ LLMs through a unified OpenAI-compatible interface

⭐ 60.0k↑ 2999/movoice-agents

ragflow

open-source

Open-source RAG engine combining knowledge retrieval and agent capabilities for LLMs

⭐ 91.6k↑ 2420/moobservability-evaluation

Langfuse

open-source

Open-source LLM engineering platform for observability, evaluation, prompt and dataset management

⭐ 35.3k↑ 1816/moobservability-evaluation

Mastra

free

From the team behind Gatsby, Mastra is a framework for building AI-powered applications and agents with a modern TypeScript stack.

⭐ 28.5k↑ 969/moobservability-evaluation

MinerU

free

Transforms complex documents like PDFs into LLM-ready markdown/JSON for your Agentic workflows.

⭐ 80.9k↑ 3758/moobservability-evaluation

Bifrost AI Gateway

open-source

Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 Β΅s overhead at 5k RPS.

⭐ 8.5k↑ 833/moobservability-evaluation

Promptfoo

open-source

Open-source CLI and library for evaluating and red-teaming prompts, agents, RAG systems, and LLM apps

⭐ 25.6k↑ 1114/moobservability-evaluation

World Monitor

open-source

AI-powered dashboard for real-time news aggregation, geopolitical monitoring, and infrastructure tracking

⭐ 87.6k↑ 6871/moobservability-evaluation
S

SkillSpector

open-source

Scans AI agent skills for prompt injection, data exfiltration, supply-chain risks, and vulnerabilities

⭐ 18.9k↑ 2820/moobservability-evaluation

Opik

open-source

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

⭐ 22.3k↑ 607/moobservability-evaluation

Firecrawl

free

πŸ”₯ The Web Data API for AI - Turn entire websites into LLM-ready markdown or structured data

⭐ 187.4k↑ 14070/mobrowser-web-agents

DeepEval

open-source

The LLM Evaluation Framework

⭐ 18.5k↑ 674/moobservability-evaluation

phoenix

free

AI Observability & Evaluation

⭐ 11.7k↑ 416/moobservability-evaluation

Manifest

open-source

Smart LLM Routing for OpenClaw. Cut Costs up to 70% 🦞🦚

⭐ 7.6k↑ 549/moobservability-evaluation

langwatch

free

The platform for LLM evaluations and AI agent testing

⭐ 4.9k↑ 276/mono-code-agent-builders
i

iFixAi

open-source

Audits AI agents against business KPIs and organizational requirements with A–F scorecards

⭐ 18.0k↑ 20340/moobservability-evaluation
M

MLflow

open-source

Open-source AI engineering platform for agents, LLMs, and ML models

⭐ 28.2k↑ 330/moobservability-evaluation

Weaviate

open-source

Open-source cloud-native vector database for semantic search, filtering, RAG, and reranking

⭐ 16.9k↑ 153/moobservability-evaluation

Haystack

open-source

Open-source AI orchestration framework for modular RAG pipelines and agent workflows

⭐ 26.6k↑ 319/moobservability-evaluation

Agenta

free

The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

⭐ 4.8k↑ 130/moobservability-evaluation
N

Netdata

open-source

The fastest path to AI-powered full stack observability, even for lean teams.

⭐ 80.8k↑ 420/moobservability-evaluation

Prefect

open-source

Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

⭐ 24.0k↑ 317/moobservability-evaluation

voltagent

open-source

AI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework

⭐ 10.7k↑ 584/moobservability-evaluation

DocsGPT

open-source

Private AI platform for agents, assistants and enterprise search. Built-in Agent Builder, Deep research, Document analysis, Multi-model support, and API connectivity for agents.

⭐ 18.3k↑ 80/moobservability-evaluation

OpenLIT

open-source

Open-source platform for AI agent tracing, evaluations, guardrails, prompts, and GPU monitoring

⭐ 2.8k↑ 77/moobservability-evaluation

unstructured

open-source

Open-source ETL for converting documents into structured data for language models

⭐ 15.5k↑ 188/moobservability-evaluation

txtai

open-source

πŸ’‘ All-in-one AI framework for semantic search, LLM orchestration and language model workflows

⭐ 13.0k↑ 102/moobservability-evaluation

Scrapegraph-ai

open-source

Python scraper based on AI

⭐ 23.1k↑ 1928/moobservability-evaluation

OpenLLMetry

open-source

Open-source observability for your GenAI or LLM application, based on OpenTelemetry

⭐ 7.5k↑ 81/moobservability-evaluation

oumi

open-source

Easily fine-tune, evaluate and deploy gpt-oss, Qwen3, DeepSeek-R1, or any open source LLM / VLM!

⭐ 9.4k↑ 76/moobservability-evaluation

WFGY

free

WFGY is an open-source AI Troubleshooting Atlas for RAG, agents, and real-world AI workflows. Includes the 16-problem map, Global Debug Card, and WFGY 3.0. ⭐ Star to help more builders find this repo.

⭐ 1.8k↑ 17/moobservability-evaluation

Langroid

open-source

Harness LLMs with Multi-Agent Programming

⭐ 4.1k↑ 26/movoice-agents
g

garak

open-source

the LLM vulnerability scanner

⭐ 9.4k↑ 60/moobservability-evaluation

UQLM

open-source

UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection

⭐ 1.2k↑ 12/moobservability-evaluation

ChainForge

open-source

An open-source visual programming environment for battle-testing prompts to LLMs.

⭐ 3.0k↑ 11/mono-code-agent-builders

helicone

open-source

🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 πŸ“

⭐ 6.2k↑ 133/moobservability-evaluation

STORM

open-source

An LLM-powered knowledge curation system that researches a topic and generates a full-length report with citations.

⭐ 31.5k↑ 560/moobservability-evaluation

llmware

open-source

Unified framework for building enterprise RAG pipelines with small, specialized models

⭐ 14.8kβ†’ 6/moobservability-evaluation

Ragas

open-source

Supercharge Your LLM Application Evaluations πŸš€

⭐ 15.9k↑ 442/moobservability-evaluation

LLM-eval-survey

free

The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

⭐ 1.6k↑ 3/moobservability-evaluation

TensorZero

open-source

TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.

⭐ 11.7k↑ 89/moobservability-evaluation

LangFair

free

LangFair is a Python library for conducting use-case level LLM bias and fairness assessments

⭐ 262↑ 1/moobservability-evaluation

Gitingest

open-source

Replace 'hub' with 'ingest' in any GitHub URL to get a prompt-friendly extract of a codebase

⭐ 15.8k↑ 248/moobservability-evaluation

OpenAI Evals

free

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

⭐ 19.5k↑ 230/moobservability-evaluation

Pezzo

open-source

πŸ•ΉοΈ Open-source, developer-first LLMOps platform designed to streamline prompt design, version management, instant delivery, collaboration, troubleshooting, observability and more.

⭐ 3.3k↑ 10/moobservability-evaluation

bRAG-langchain

free

Everything you need to know to build your own RAG application

⭐ 4.2k↑ 16/moobservability-evaluation

llama-github

open-source

Open-source Python library for agentic RAG across public GitHub code, issues, and repository information

⭐ 294β†’ 4/moobservability-evaluation

Langchain-Chatchat

open-source

Offline-deployable Chinese knowledge base Q&A with RAG and agents using LangChain and open-source LLMs

⭐ 38.7k↑ 161/moobservability-evaluation

Vanna

open-source

πŸ€– Chat with your SQL database πŸ“Š. Accurate Text-to-SQL Generation via LLMs using Agentic Retrieval πŸ”„.

⭐ 23.8k↑ 109/mono-code-agent-builders

AgentOps

open-source

Python SDK for monitoring, cost tracking, benchmarking, and debugging AI agents

⭐ 5.9k↑ 73/moobservability-evaluation

AgentBench

open-source

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

⭐ 3.8k↑ 78/moobservability-evaluation

Auto-evaluator

free

Evaluation tool for LLM QA chains

⭐ 1.1k↑ 51/moobservability-evaluation

Gorilla

open-source

Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)

⭐ 13.0k↑ 41/movoice-agents

R2R

open-source

SoTA production-ready AI retrieval system. Agentic Retrieval-Augmented Generation (RAG) with a RESTful API.

⭐ 8.0k↑ 42/moobservability-evaluation

Verba

open-source

Retrieval Augmented Generation (RAG) chatbot powered by Weaviate

⭐ 7.7k↑ 13/moobservability-evaluation

text-extract-api

open-source

Local FastAPI for OCR extraction and PII removal from images, PDFs and Office files to Markdown or JSON

⭐ 3.2k↑ 17/moobservability-evaluation

FastChat

open-source

An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.

⭐ 39.6k↑ 16/moobservability-evaluation

clip-retrieval

open-source

Easily compute clip embeddings and build a clip retrieval system with them

⭐ 2.8k↑ 10/moobservability-evaluation

VisionAgent

open-source

This tool has been deprecated. Use Agentic Document Extraction instead.

⭐ 5.3k↑ 4/moobservability-evaluation

TaskingAI

open-source

The open source platform for AI-native application development.

⭐ 5.4k↑ 4/movoice-agents

UpTrain

open-source

Open-source platform to evaluate and improve generative AI applications with 20+ preconfigured evaluations

⭐ 2.4k↑ 4/moobservability-evaluation

Pathway

open-source

Ready-to-deploy templates for RAG and enterprise search that sync with live data sources

⭐ 58.9kβ†’ 84/moobservability-evaluation

LangKit

open-source

Open-source text metrics toolkit for monitoring language models through input and output signals

⭐ 997↑ 3/moobservability-evaluation

LLM Comparator

open-source

LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.

⭐ 526↑ 1/mono-code-agent-builders

Swiss Army Llama

free

A FastAPI service for semantic text search using precomputed embeddings and advanced similarity measures, with built-in support for various file types through textract.

⭐ 1.1k↑ 0/moobservability-evaluation

Canopy

open-source

Retrieval Augmented Generation (RAG) framework and context engine powered by Pinecone

⭐ 1.0k↑ 0/moobservability-evaluation

Banana-lyzer

open-source

Open source AI Agent evaluation framework for web tasks πŸ’πŸŒ

⭐ 330↑ 0/moobservability-evaluation

GPTDiscord

open-source

A robust, all-in-one GPT interface for Discord. ChatGPT-style conversations, image generation, AI-moderation, custom indexes/knowledgebase, youtube summarizer, and more!

⭐ 1.9k↑ 0/moobservability-evaluation

Repochat

open-source

Chatbot assistant enabling GitHub repository interaction using LLMs with Retrieval Augmented Generation

⭐ 318↑ 0/moobservability-evaluation