Gemma 4 is a family of open-weight, multimodal models by Google DeepMind, available in pre-trained and instruction-tuned variants across five sizes (E2B, E4B, 12B, 26B A4B, and 31B) in both Dense and Mixture-of-Experts (MoE) architectures. Built to run on everything from mobile devices to high-end servers, Gemma 4 is optimized for reasoning, coding, agentic workflows, and text generation.
Key Features
Multimodal Inputs: Supports text and variable-resolution images across all sizes, plus native video and audio on E2B, E4B, and 12B. Output is text-only.
Expanded Context & Multilingual: Up to 256K context window (128K on smaller models) with support for over 140 languages.
Reasoning & Agents: Features configurable thinking modes, native function calling for autonomous agents, and native system prompt support.
Hybrid Attention: Interleaves local sliding window and full global attention with Proportional RoPE (p-RoPE) and unified Keys/Values to optimize memory and speed for long-context tasks.
Source: https://huggingface.co/google/gemma-4-E4B-it
No comments:
Post a Comment