8 Best FunASR Alternatives in 2026 (Open Source)
FunASR — Open-source speech recognition toolkit for offline, streaming, and edge ASR, VAD, punctuation, and diarization. Provides both OpenAI-compatible API serving and MCP server specifically for AI agent integration alongside comprehensive speech processing pipelines.
Short answer
- Closest match to FunASR: whisperX.
- Most actively developed: Pipecat (2,861 commits in the last 90 days).
- Fastest growing: agents (+1,358 GitHub stars in the last 30 days).
- No commit in 6+ months: WhisperS2T, AudioGPT, Ultravox and RealChar.
These 8 open-source tools do the same job. They are ordered by how closely they match FunASR, with live GitHub data so you can see which projects are actively maintained.
| Tool | GitHub stars | Stars / 30d | Last commit |
|---|---|---|---|
| FunASR(original) | 20.6k | +165 | 2026-10-02 |
| whisperX | 24.3k | +538 | 2026-09-26 |
| WhisperS2T | 580 | +3 | 2024-08-25 |
| Buzz | 21.8k | +535 | 2026-10-02 |
| AudioGPT | 10.2k | -7 | 2023-05-05 |
| Ultravox | 4.6k | +31 | 2025-12-12 |
| agents | 14.4k | +1,358 | 2026-10-01 |
| Pipecat | 16.1k | +833 | 2026-10-02 |
| RealChar | 6.2k | +1 | 2024-02-03 |
1. whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
What sets it apart: Adds word-level timestamps and speaker diarization on top of Whisper — solving the two biggest gaps in OpenAI's original model
Best for: Batch transcription with accurate word-level timestamps; Meeting transcription with speaker identification
2. WhisperS2T
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
What sets it apart: vs WhisperX / HuggingFace Pipeline: 2.3-3X speed improvement through superior pipeline architecture (not just backend optimization) — with multiple inference backend choices and built-in hallucination reduction
Best for: High-volume speech transcription requiring speed optimization; Multilingual audio processing with backend flexibility; Applications needing reduced hallucination output from Whisper
3. Buzz
Buzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper.
What sets it apart: vs Whisper CLI: full GUI with live transcription, speaker ID, and watch folders; vs cloud transcription (AssemblyAI/Deepgram): completely offline with zero data leaving the device
Best for: Offline audio/video transcription with privacy; Live presentation captioning; Batch transcription of media files
4. AudioGPT
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
What sets it apart: It brings speech, singing, general audio, and talking-head capabilities together in one open-source conversational system.
Best for: Researchers and developers experimenting with conversational audio understanding and generation; Projects combining multiple speech, sound, music, and talking-head models
5. Ultravox
A fast multimodal LLM for real-time voice
What sets it apart: vs ASR+LLM pipelines (Whisper+GPT): direct audio-to-embedding projection eliminates ASR latency bottleneck, enabling true real-time voice understanding
Best for: Real-time voice AI agents requiring sub-100ms latency; Custom domain voice applications with proprietary audio data
6. agents
A framework for building realtime voice AI agents 🤖🎙️📹
What sets it apart: The leading open-source framework for realtime voice AI agents with WebRTC infrastructure, semantic turn detection, multi-agent handoff, and native telephony — vs alternatives that bolt voice onto text-first frameworks
Best for: Building production voice AI agents and assistants; Real-time conversational AI with telephony integration; Multi-agent voice workflows with handoffs
7. Pipecat
Open Source framework for voice and multimodal conversational AI
What sets it apart: Only production-grade framework for real-time voice AI with composable pipelines — supports 17+ STT and 20+ TTS providers with ultra-low latency, unlike text-focused agent frameworks
Best for: Building real-time voice AI agents and assistants; Multimodal conversational interfaces with audio, video, and text
8. RealChar
Create and converse with customizable AI characters in real time on web, mobile, and terminal
What sets it apart: vs Character.AI: fully open-source with voice cloning, multi-platform (web+iOS+phone), and pluggable LLM/TTS backends — own your AI characters
Best for: Building interactive AI character experiences with voice; Developers creating multi-platform conversational AI personas
FAQ
- What are the best alternatives to FunASR?
- The closest open-source alternatives to FunASR are whisperX, WhisperS2T and Buzz, followed by AudioGPT, Ultravox and agents. They are ranked by how closely they match what FunASR does.
- Which FunASR alternative is the most popular?
- whisperX has the most GitHub stars among FunASR alternatives, with 24,337 stars.
- Which FunASR alternative is the most actively maintained?
- By recent activity, Pipecat (2,861 commits in the last 90 days) is the most actively developed alternative.