The updated GigaChat Audio model is available through the GigaChat AI assistant and as free open-source code for software developers across the globe
Users of the GigaChat AI assistant can enjoy the updated GigaChat Audio AI, a large language model capable of processing audio files and voice messages without converting speech to text first. The LLM has been trained to understand intonation and had its sound data processing potential enhanced to deliver higher performance.
The AI assistant with more empathy
The updated GigaChat can detect users’ positive or negative emotions based on intonation, voice characteristics, and pronunciation nuances to be more on-point in responses. For example, it will talk more gently to an irritated user or mirror the mood of someone sharing good news.
The model can process recordings up to three hours in length and navigate within them. You can ask it about the exact time when some specific issue was discussed, request a summary of a segment you need, or get a summary of the entire recording with timestamps. The AI is also capable of distinguishing between different speakers in a recording. These features will be useful for reviewing meetings and calls, navigating through recordings, and taking minutes.
The AI assistant can now remember facts that users find important directly in a voice dialogue and rely on them in future sessions. Upon subsequent requests, GigaChat will take into account your previous wishes. For example, it will plan a travel route based on your interests. The user has full control over this function, which means recorded facts can be viewed, edited, or the memory can be disabled at any time in the profile settings.
In-house tests show that GigaChat Audio is as good at understanding and responding to voice queries as the best global solutions. This was confirmed by the Arena-Hard-Audio benchmark, when other neural networks blindly compared the responses of different models to the same voice questions. GigaChat Audio secured a win rate of 70%, almost on par with Gemini 3 Flash preview (77.5%) and higher than Gemini 2.5 Pro (62%). In terms of emotion recognition accuracy, The model achieved 80% accuracy in emotion recognition, outperforming Qwen3-Omni-30B (70%) or Kimi-Audio (62%).
Availability for developers
In addition to the voice model’s integration into the AI assistant, the team has open-sourced GigaChat3.1-Audio-10B, a lightweight version which supports English and multiple other languages. It can be used to create transcription solutions, pronunciation trainers, voice-over quality assessment tools, context-aware voice interpreters, and summarization solutions for long audio recordings.
The team has also open-sourced GigaAM Multilingual, a family of automatic speech recognition models designed for applications including call centers, voice assistants, meeting transcription, interviews, podcasts, voice input, and automatically generated subtitles.
Available on GitVerse and Hugging Face, both the models have been pre-trained on many languages, which allows them to be quickly fine-tuned for additional ones, allowing them to be quickly fine-tuned for additional languages. This would require only a few dozen hours of marked-up audio recordings. Scientific papers about the new models have been accepted at Interspeech 2026, a premier international conference on speech science and technology.
This story was distributed as a release by Jon Stojan under HackerNoon’s Business Blogging Program.