Xiaohongshu Challenges Big Tech AI With Open-Source, On-Premise Voice System

Xiaohongshu Challenges Big Tech AI With Open-Source, On-Premise Voice System

Chinese social media giant Xiaohongshu has released an open-source AI voice interaction system designed for on-premise deployment, a move that directly challenges the dominant API-based models of major technology companies by targeting enterprise needs for data security and cost control.

The company's AI audio team has launched FireRedChat, a full-duplex voice system that allows for natural, human-like conversation with AI. The system is available for private installation, giving companies full control over their data and infrastructure, a key differentiator from services that rely on external cloud APIs.

By making the technology fully open-source and free, Xiaohongshu aims to address critical industry pain points such as high latency, poor control, and sensitivity to background noise. The release provides enterprises with a viable alternative for developing sophisticated voice AI applications, potentially disrupting the market for paid, cloud-based AI services.

According to the company, the goal is to create a voice AI that is not just a command-response tool but an empathetic companion that "knows you, empathizes with you, and can express itself." This vision positions FireRedChat to power a new generation of emotionally intelligent applications, from customer service to mental health support.

A Push for Natural Conversation

Achieving natural, full-duplex voice interaction—where a user can speak and be heard at the same time, much like a human conversation—has long been a significant technical hurdle. Systems must accurately detect conversational turns, handle interruptions gracefully, and distinguish the primary speaker's voice from background noise or other speakers. The reliance on closed-source APIs has also limited customization and raised data privacy concerns for enterprises, obstacles that have hindered the widespread adoption of open-source solutions.

Core Technical Innovations

FireRedChat introduces several key technical features to address these challenges. The system’s design is centered on five core breakthroughs:

  • On-Premise Full-Duplex: It is presented as the industry's first system to combine full-duplex capabilities with a one-click on-premise deployment option, ensuring data security and system scalability.
  • Precision Interruption Handling: It uses a proprietary personalized Voice Activity Detection (pVAD) model to focus on the main speaker and a lightweight End-of-Turn (EoT) detector to identify natural pauses, reducing false interruptions and awkward response delays.
  • Dual-Track Architecture: The system offers two deployment paths: a stable "cascaded" path (ASR → LLM → TTS) for flexible module optimization and a "semi-cascaded" path (AudioLLM → TTS) that directly processes audio for lower latency and more emotionally aware responses.
  • Low End-to-End Latency: Through modular decoupling and stream processing optimization, FireRedChat reportedly achieves end-to-end latency that approaches the performance of industrial-grade, closed-source systems, enabling real-time responsiveness.
  • Emotional Intelligence: By integrating its AudioLLM with the FireRedTTS-2 synthesis engine, the system can detect acoustic cues like tone and pace in a user's voice and generate empathetic, contextually appropriate vocal responses.

An Open and Extensible Framework

FireRedChat's architecture is decoupled into three core modules to balance high performance with maintainability and extensibility. A Turn-taking Controller manages the flow of conversation, while an Interaction Module handles the core processing in either cascaded or semi-cascaded mode. A Dialogue Manager oversees the conversation state and integrates external tools, such as web search or retrieval-augmented generation (RAG).

The entire system—including its core TTS, ASR, pVAD, and EoT models—is open-source and requires no API fees. This modular framework, combined with pre-built integrations for platforms like Dify, is designed to help developers quickly move from a demonstration to a production-ready application.

Performance and Use Cases

Xiaohongshu released benchmark data indicating FireRedChat’s superior performance over other open-source frameworks in key areas. The pVAD model significantly reduces incorrect interruptions caused by noise, while the EoT detector more accurately determines when a user has finished speaking. Its end-to-end response time in local deployments is positioned as a major leap forward for open-source solutions.

The company has identified several typical application scenarios for the technology, including intelligent voice assistants, customer service and call center automation in noisy environments, and applications in education and mental health where empathetic interaction is crucial. With its focus on an open, on-premise model, Xiaohongshu is positioning FireRedChat not just as a technical release but as a strategic asset for the broader developer and enterprise community.

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe