Google Releases Gemma 4 Under Apache 2.0 License

Gemma 4 open AI model graphic

Quick Read

  • Gemma 4 is released under the commercially permissive Apache 2.0 license, enabling unrestricted modification and distribution.
  • The models range from 2B to 31B parameters, with the 31B Dense model ranking third on the global Arena AI open model leaderboard.
  • Edge variants (E2B and E4B) are optimized for offline, low-latency execution on mobile devices, including Android smartphones and Raspberry Pi.

Google has officially released Gemma 4, a new generation of open-source AI models that mark a significant shift in the company’s accessibility strategy. By adopting the commercially permissive Apache 2.0 license, Google is removing previous usage restrictions, allowing developers, enterprises, and researchers to modify and distribute the models with unprecedented freedom. The release comes as developers have already downloaded previous Gemma iterations over 400 million times, fueling a community of more than 100,000 variants.

Expanding the Gemma 4 Ecosystem

The Gemma 4 family is built on the same research and technology powering Google’s proprietary Gemini 3 models, but is optimized for local hardware performance. The release includes four distinct sizes: Effective 2B (E2B) and Effective 4B (E4B) for edge and mobile devices, a 26B Mixture of Experts (MoE) model, and a 31B Dense model. According to internal benchmarks and the Arena AI text leaderboard, the 31B model currently ranks as the third-most capable open model globally, frequently outperforming competitors 20 times its parameter size.

Capabilities in Agentic Workflows and Mobile AI

Engineered for advanced reasoning and autonomous agentic workflows, the models support multi-step planning, function calling, and structured JSON outputs. All four variants are natively multimodal, capable of processing video and images, while the smaller E2B and E4B versions also feature native audio input for real-time speech recognition. To accommodate complex tasks, the larger models support a 256K context window, while the edge-optimized versions offer 128K, enabling the processing of extensive code repositories or lengthy documentation entirely offline.

Deployment Across Diverse Hardware

The strategic focus on efficiency allows these models to operate on varied hardware profiles. The E2B and E4B models are designed for near-zero latency on consumer electronics, including Android smartphones, Raspberry Pi, and NVIDIA Jetson Orin Nano boards. Meanwhile, the 26B and 31B models are optimized for developer workstations, with the 26B MoE variant utilizing only 3.8 billion active parameters during inference to maintain high speed. Google has ensured day-one support across major tools, including Hugging Face, vLLM, llama.cpp, and Ollama, facilitating seamless integration for the global developer community.

The shift to an Apache 2.0 license represents a critical pivot in Google’s AI strategy, effectively lowering the barrier to entry for developers seeking to build sovereign, private, and offline-capable AI applications that compete directly with proprietary cloud-based services.

|
Contributor:Azat TV Editorial
|
Publisher:Azat TV

LATEST NEWS