
@googlegemma
The official home of Google's Gemma. Lightweight, state-of-the-art open models by Google DeepMind, built on Gemini tech. What will you build? 🚀💻
Explore a new community-built playground allowing you to try Gemma models right in your browser, no installation needed! You no longer need to jump between repositories. Explore an interactive timeline of our models, then switch over to the playground to test them securely on-device! Powered by WebGPU & Transformers.js, test 10+ models and explore the interactive "Gemma Journey" timeline.
Try it out: web-gemma.vercel.app GitHub repository: github.com/NSTiwari/WebGe… Built by @NSTiwari21
This week, Gemma surpassed 900 million downloads! 🎉 All the way from Gemma 1 and ShieldGemma to MedGemma and Gemma 4, we'll keep supporting open source. More to come!
What happens when you put Gemma 4 into the hands of developers across 20+ countries? So much innovation! Like this app that turns an image prompt into a working web app in seconds, running on-device! Want to get involved? Join or host an event!
Everything you need to get started: Find a hackathon in your area: luma.com/gemma-naek Apply for sponsorship: goo.gle/build-with-gem… Original hackathon project post: x.com/theShawwwn/sta…
Gemma 4 just crossed 300 million downloads. Thank you to the developers, researchers, and open-source community building with us. Your work and feedback drive this project forward. Let's keep building!
Racing into the future at Goodwood Festival of Speed!🏎️⚡️ We teamed up with Formula E to showcase a next-gen onboard AI assistant, powered by Gemma open models! Here is a look at the tech stack built by the Formula E team:
- @Raspberry_Pi streams telemetry directly to the Pixel 10 Pro for local database storage - Pixel 10 Pro processes telemetry via AICore (ML Kit) for handling real-time driver voice queries - Gemma 4 E2B/E4B optimized to run on <1.45GB VRAM - Gemma operates as a local ADK-powered agent, syncing via A2A with a @googlecloud agent for post-race analytics
Extremely fast multimodal inference! Damage Scout uses Gemma 4 on @cerebras running at an impressive 2,300+ toks/s! It analyzes rental car walkaround videos and generates annotated damage reports with box coordinates in <6 seconds.
Built on the blazing fast infrastructure over at @Cerebras! Watch the full demo and breakdown here: x.com/cerebras/statu…
Voice AI without the wait! ⏱️ Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the brain for voice AI at ultra-fast inference speeds. Add it to a fully open-source, cascaded speech-to-speech stack that can be used to power existing voice apps! 🗣️
Try the demo by @andimarafioti and team: huggingface.co/spaces/smolage… Read the blog: huggingface.co/blog/cerebras-…
Build with Gemma community hackathons are going on now in over 20 countries. Join a hackathon near you to create awesome AI applications with Gemma 4, or apply to host your own!
Find an event near you: luma.com/gemma-naek Apply for sponsorship: goo.gle/build-with-gem…
Ready to customize Gemma, but not sure how? Have your agents assist you by using the gemma-trainer skill! It helps agents set your training configs, manage training runs and evaluate results. Let’s make fine-tuning Gemma 4 possible for everyone!
You can find more information at the links before: - Gemma Skills repo: github.com/google-gemma/g… - Blog post: dev.to/googleai/maste…
Every millisecond counts for voice agents. Introducing Gemma 4 31B on LiveKit Inference! Optimized for real-time voice agents: - 354ms time to first audio - 192ms time to first token - Beats GPT-4.1 in agentic tool use (31B achieves 76.9% on tau2bench) Fast, capable, and efficient!
Learn more and test the integration right here: livekit.com/products/infer…
We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is a breakdown of what’s being fixed and updated in this release: 🧵👇
🏎️ Flash Attention: We've enabled uniform Flash Attention 4 (FA4) support on NVIDIA Hopper GPUs! Expect prefill throughput to jump by 25-70% and time-to-first-token (TTFT) to drop by up to 31%. ⚡ (2/5)
An offline AI tutor, a medieval bard, and a real-world adventure game. What do they have in common? They were all built using Gemma 4! With MTP, our 12B Unified model, and QAT checkpoints, you get ultimate flexibility. Deploy anywhere from edge to workstation under Apache 2.0.
Read the full blog: blog.google/innovation-and…
Gemma 4 is live on @cerebras, the fastest multimodal inference ever! Running on Gemma 4 31B open-weight model at a blistering 1,500+ tokens/sec. That's a 15x speedup, unlocking real-time visual and agentic loops without the GPU lag.
Read the full article by @cerebras: cerebras.ai/blog/gemma-4-o…
Hugging Face Gemma Challenge results are in! 📈 Over 6 days, more than 100 AI agents and humans collaborated to make Gemma 4 inference 5x faster on a single NVIDIA A10G GPU. - Fastest result: 491.8 TPS (fastest overall, but resulted in a drop in model quality in other areas) - Fastest lossless: 315 TPS A great example of what humans and agents can achieve when they work together.
Beyond the speedups, the agents' coordination stood out. They naturally pooled resources, compute-poor agents focused on debugging while others ran heavy experiments, and even self-policed to prevent "reward hacking." Full breakdown of the results: huggingface.co/spaces/agent-c…
Decoding centuries-old text can be a complex challenge. This project shows how to fine-tune Gemma 4 E2B on a single GPU. The model translates Classical Korean to modern text. The accuracy improved from 4% to 80%!
Read the full walkthrough here: dev.to/googleai/turni…
Gemma 4 now works on-device using React Native! You can now run Gemma 4 fully offline in your cross-platform apps with local hardware acceleration: ⚡ Vulkan delegate on Android ⚡ MLX delegate on Apple Silicon See Gemma 4's vision and tool-use capabilities in action, instantly reading a flyer and scheduling a calendar event, 100% on-device.
Shout out to @swmansion for support! Check out the opensource demo app: github.com/software-mansi… LLM docs: docs.swmansion.com/react-native-e…
Just dropped: The Gemma 4 Technical Report! Dive into the decisions behind our latest open-weight multimodal models, Gemma 4 (E2B–31B). Discover details on how the Gemma team approaches architecture, efficiency and responsibility. The report also covers the recent encoder-free 12B variant, quantization-aware training and multi-token prediction (MTP) drafters.
📃Read the full technical report here: arxiv.org/abs/2607.02770
Ready to build the next generation of AI agents? We’re partnering with @AIatAMD, @FireworksAI_HQ & @lablabai for the AMD Developer Hackathon: ACT II! Build with Gemma open models on AMD GPUs & Fireworks AI for a chance to win from our $6,000 Gemma prize pool.
Sign up with AMD below to secure your spot by 8pm CET July 6. The submission deadline is July 11. lablab.ai/ai-hackathons/…
“Agentic kernel optimization is the future of on-device inference” @xenovacom used Fable 5 to write kernels that pushed Gemma 4 to a massive 255 tok/s on WebGPU with M4. He shared the demo, so you can try in your browser!!
Original post by @xenovacom x.com/xenovacom/stat…
Gemma 4 31B at over 1,800 tokens per second! Gemma 4 is now in Public Preview on Cerebras.
Give it a try: chat.cerebras.ai Photo credit: x.com/RayFernando133…
Gemma 4 is the first multimodal model on Cerebras! ️ What can you build with Gemma 4 31B running at 1500 tokens per second? Join the Cerebras x Gemma 4 24-hour virtual hackathon this Sunday to compete for $5,000 in prizes. Participants get early access to Gemma 4 on Cerebras.
Register here: luma.com/cerebras-piwl
Gemma 4 just hit 200M downloads in only 2.5 months! For context, total downloads across the entire Gemma family of models were at 100M when we launched Gemma 3. The community's acceleration is incredible. Thank you to everyone building with Gemma. Watch how developers are driving real-world impact:
Want to host your own Gemma hackathon? We’re sponsoring 1-day hackathons on Kaggle to help developers dive into open models! From building lightweight tools to tackling your community's unique challenges, this is your chance to lead the charge with Gemma 4.👇
Apply now to become a host: goo.gle/build-with-gem…
16 parallel runs of Gemma 4 26B A4B on a single NVIDIA DGX Spark! Pushing 18 tok/s per instance and a 300 tok/s aggregate. It can even hit 32 parallel runs. This level of concurrency highlights how efficient the architecture is.
Model link: huggingface.co/nvidia/Gemma-4… Original post by @onusoz , x.com/onusoz/status/…
AI has entered orbit! NASA and Loft Orbital are now analyzing images directly in orbit. We're proud to see Gemma helping power this huge space tech milestone from behind the scenes.
Read the full story here: techcrunch.com/2026/06/15/a-s…
Teamwork makes the dream work. Now running locally. Watch Gemma 4 26B orchestrate 10 parallel sub-agents to code an SVG art gallery in seconds. Hitting 100+ tokens/sec, imagine how you can scale this for complex tasks or local chatbots for entire teams!!
Start experimenting with parallel workflows and build your own multi-agent systems today! github.com/google-gemma/c…
Gemma 4 E2B goes super fast on Intel AI PCs thanks to LiteRT NPU support on OpenVINO! ⚡1.3x faster prefill performance over GPU 📈2.8x improvement in performance-per-watt 🔋Runs background LLM tasks with zero thermal throttling or heavy battery drain
Read the full technical breakdown and documentation here: - Intel Blog: intel.com/content/www/us… - LiteRT Docs: developers.google.com/edge/litert/ne…
Want to teach Gemma to master chess? Check out this awesome community project showing how to fine-tune Gemma 4 12B on your own data, 100% locally! Running text, images, and audio on just 8GB VRAM makes custom models more accessible than ever.
Check out the full breakdown by @akshay_pachaar here: x.com/akshay_pachaar…
Real-time social robotics, from the cloud to your local device. Watch Ian from our DevX team use Gemini Live for a seamless voice chat with Reachy Mini. Then, stick around until the end to see the robot running locally on Gemma 4!
Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
⚡ Blazing Fast: By shifting the decode bottleneck from memory-bandwidth to compute, DiffusionGemma delivers up to a 4x speedup on standard accelerators. (1000+ tokens per second on a single NVIDIA H100, 700+ tokens per second on NVIDIA GeForce RTX 5090!)
💻 Accessible Hardware: A 26B Mixture of Experts (MoE) model that activates only 3.8B parameters during inference. Fits comfortably within 18GB VRAM limits of high-end dedicated consumer GPUs when quantized.
🔄 Bi-directional Attention: Generating 256 tokens in parallel allows every token to attend to all others. Unlocks significant advantages for non-linear domains like in-line editing, code infilling, and mathematical graphs.
🧠 Intelligent Self-Correction: Similar to AI image generators, the model iteratively refines its own output. It evaluates the entire text block at once to seamlessly close formatting and fix mistakes in real-time.
🛠️ Broad Ecosystem Support: Download the weights today from Hugging Face. Begin building and serving efficiently with MLX, vLLM, Unsloth, Hugging Face Transformers, RedHat, NVIDIA NeMo, NIM and Gemini Enterprise Agent Platform Model Garden or NVIDIA NIM. We can't wait to see
Read more in our blog: blog.google/innovation-and…
Introducing the Fast Gemma Challenge with Hugging Face Over the next few days, dozens of agents will collaborate to make Gemma 4 E4B even faster!
Join the challenge and submit your agents! huggingface.co/spaces/gemma-c…