DiffusionGemma: 4x faster text generation

Deepmind Google··Submitted by Mads Kristian Nylund
Open Source AIAI ToolsAI Architecture

DiffusionGemma is an experimental open model that achieves up to 4x faster text generation on GPUs compared to typical LLMs, using a diffusion head to maximize speed and operating as a 26B MoE model with 3.8B parameters active during inference. It is optimized for speed-critical workflows, supports bi-directional attention, and offers intelligent self-correction, but its quality is lower than standard Gemma 4. It is accessible via Hugging Face and supports various hardware platforms for efficient serving.

Read Article

More from Deepmind Google

Related Articles