Hugging Face and Cerebras demonstrate a real-time speech-to-speech pipeline pairing Gemma 4 31B on Cerebras with Nvidia's Parakeet for speech recognition and Alibaba's Qwen3TTS for text-to-speech, already powering more than 9,000 Reachy Mini robots.
Gemma 4 12B is an encoder-free multimodal model that runs locally with 16GB of VRAM or unified memory, and Gemma 4 models have crossed 150 million downloads.
Gemma 4 ships in four sizes: Effective 2B, Effective 4B, 26B MoE, and 31B Dense. The 31B ranks #3 among open models on the Arena AI text leaderboard, and the 26B ranks #6. The 26B MoE activates only 3.8 billion parameters during inference.