Gemma 4 is presented as a new open-weight, natively multimodal generation spanning dense and Mixture-of-Experts models from 2.3B to 31B parameters. The report describes improved vision and audio encoders across the family, plus a unified encoder-free 12B architecture that consumes raw audio and image patches. It also introduces a thinking mode that generates reasoning traces before answers. According to the abstract, the design targets faster inference, lower memory and compute requirements, and stronger long-context performance, with gains reported across STEM, multimodal, and long-context benchmarks.
No heat snapshots are available in the last 24 hours.