
Gemma 4 12B is a mid-sized multimodal model optimized for local deployment on consumer laptops, offering efficient processing of visual and audio data with reduced memory usage. It features a unified architecture that integrates vision and audio into a single language model backbone, enabling direct integration with development tools and frameworks. The model achieves performance comparable to the 26B MoE model but with significantly lower memory requirements, making it suitable for 16GB VRAM systems.
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
