8 Best FunASR Alternatives in 2026 (Open Source)

FunASR — Open-source speech recognition toolkit for offline, streaming, and edge ASR, VAD, punctuation, and diarization. Provides both OpenAI-compatible API serving and MCP server specifically for AI agent integration alongside comprehensive speech processing pipelines.

Short answer

  • Closest match to FunASR: whisperX.
  • Most actively developed: Pipecat (2,861 commits in the last 90 days).
  • Fastest growing: agents (+1,358 GitHub stars in the last 30 days).
  • No commit in 6+ months: WhisperS2T, AudioGPT, Ultravox and RealChar.

These 8 open-source tools do the same job. They are ordered by how closely they match FunASR, with live GitHub data so you can see which projects are actively maintained.

ToolGitHub starsStars / 30dLast commit
FunASR(original)20.6k+1652026-10-02
whisperX24.3k+5382026-09-26
WhisperS2T580+32024-08-25
Buzz21.8k+5352026-10-02
AudioGPT10.2k-72023-05-05
Ultravox4.6k+312025-12-12
agents14.4k+1,3582026-10-01
Pipecat16.1k+8332026-10-02
RealChar6.2k+12024-02-03
  1. 1. whisperX

    WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

    What sets it apart: Adds word-level timestamps and speaker diarization on top of Whisper — solving the two biggest gaps in OpenAI's original model

    Best for: Batch transcription with accurate word-level timestamps; Meeting transcription with speaker identification

  2. 2. WhisperS2T

    An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine

    What sets it apart: vs WhisperX / HuggingFace Pipeline: 2.3-3X speed improvement through superior pipeline architecture (not just backend optimization) — with multiple inference backend choices and built-in hallucination reduction

    Best for: High-volume speech transcription requiring speed optimization; Multilingual audio processing with backend flexibility; Applications needing reduced hallucination output from Whisper

  3. 3. Buzz

    Buzz transcribes and translates audio offline on your personal computer. Powered by OpenAI's Whisper.

    What sets it apart: vs Whisper CLI: full GUI with live transcription, speaker ID, and watch folders; vs cloud transcription (AssemblyAI/Deepgram): completely offline with zero data leaving the device

    Best for: Offline audio/video transcription with privacy; Live presentation captioning; Batch transcription of media files

  4. 4. AudioGPT

    AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

    What sets it apart: It brings speech, singing, general audio, and talking-head capabilities together in one open-source conversational system.

    Best for: Researchers and developers experimenting with conversational audio understanding and generation; Projects combining multiple speech, sound, music, and talking-head models

  5. 5. Ultravox

    A fast multimodal LLM for real-time voice

    What sets it apart: vs ASR+LLM pipelines (Whisper+GPT): direct audio-to-embedding projection eliminates ASR latency bottleneck, enabling true real-time voice understanding

    Best for: Real-time voice AI agents requiring sub-100ms latency; Custom domain voice applications with proprietary audio data

  6. 6. agents

    A framework for building realtime voice AI agents 🤖🎙️📹

    What sets it apart: The leading open-source framework for realtime voice AI agents with WebRTC infrastructure, semantic turn detection, multi-agent handoff, and native telephony — vs alternatives that bolt voice onto text-first frameworks

    Best for: Building production voice AI agents and assistants; Real-time conversational AI with telephony integration; Multi-agent voice workflows with handoffs

  7. 7. Pipecat

    Open Source framework for voice and multimodal conversational AI

    What sets it apart: Only production-grade framework for real-time voice AI with composable pipelines — supports 17+ STT and 20+ TTS providers with ultra-low latency, unlike text-focused agent frameworks

    Best for: Building real-time voice AI agents and assistants; Multimodal conversational interfaces with audio, video, and text

  8. 8. RealChar

    Create and converse with customizable AI characters in real time on web, mobile, and terminal

    What sets it apart: vs Character.AI: fully open-source with voice cloning, multi-platform (web+iOS+phone), and pluggable LLM/TTS backends — own your AI characters

    Best for: Building interactive AI character experiences with voice; Developers creating multi-platform conversational AI personas

FAQ

What are the best alternatives to FunASR?
The closest open-source alternatives to FunASR are whisperX, WhisperS2T and Buzz, followed by AudioGPT, Ultravox and agents. They are ranked by how closely they match what FunASR does.
Which FunASR alternative is the most popular?
whisperX has the most GitHub stars among FunASR alternatives, with 24,337 stars.
Which FunASR alternative is the most actively maintained?
By recent activity, Pipecat (2,861 commits in the last 90 days) is the most actively developed alternative.
8 Best FunASR Alternatives in 2026 (Open Source)