Building high-performance, offline multimodal agents using local Gemma 4 open models.
338 RSVP'd
Discover how to deploy and run Gemma 4 12B and DiffusionGemma directly on standard 16GB VRAM hardware. We will demonstrate how to leverage a unified, encoder-free architecture to natively process text, image, and audio inputs locally while using multi-token prediction to maximize generation speeds. Finally, you'll learn how to utilize the new Gemma Skills Repository to build secure, offline tool-calling agents that analyze massive datasets within a 256K context window.
RESOURCES FROM THE SESSION:
Google DeepMind
Gemma Product Manager
Google Accelerator Europe Lead