GDG AI for Science - Australia
Join the GDG AI for Science community to discuss Gemini Robotics ER 2.0, Google DeepMind’s flagship foundation model for...
108 RSVP'd
Join the GDG AI for Science community to discuss Gemini Robotics ER 2.0, Google DeepMind’s flagship foundation model for physical AI and embodied reasoning. Built on Gemini 3.5 Flash, Gemini Robotics ER 2.0 acts as a high-level brain for robots that can interpret continuous visual and audio feeds, conduct spatial and temporal reasoning, and orchestrate complex multi-step tasks before handing off motor execution to Vision-Language-Action (VLA) models or low-level robotic APIs. From real-time 2D spatial pointing to continuous video progress classification and multi-robot collaboration, ER 2.0 enables machines to perceive, plan, self-correct, and operate safely in the physical world.
In this collaborative session, we will unpack the model, code, and documentation to understand:
What Gemini Robotics ER 2.0 is: A multimodal vision-language model (VLM) specialized in physical, spatial, and temporal reasoning for autonomous robots and physical agents.
How it works: We will look into its underlying Gemini 3.5 Flash foundation, 128k context window, dynamic tool/API orchestration capabilities, and sub-second bidirectional streaming endpoint (gemini-robotics-er-2-streaming-preview) via the Live API.
What Gemini Robotics ER 2.0 is capable of: Core performance breakthroughs across spatial understanding, precision moment-finding, continuous video progress classification (57.4% accuracy), generalized instrument reading (dials, digital screens, scales), multi-robot coordination (e.g., Apptronik Apollo 2 & Franka F3 Duo), and safety instruction following.
How to use it: During the session, we will work directly through the provided official Google Colab notebook. Together, we will execute hands-on Python code using the google-genai SDK to run object pointing, bounding box detection, agentic code execution, and task progress tracking in real time.
What it CANNOT do (Limitations & Safety): We will review safety guardrails, operational limitations (such as potential visual hallucinations, computational latency trade-offs, and API key restrictions), as well as prohibited high-risk physical domains (e.g., safety-critical healthcare or hazardous operations).
Whatever else? This is a virtual, facilitated session so interaction and active participation are strongly encouraged! Discussion is aimed to go wherever the group finds most interesting.
This event is for anyone interested in the intersection of artificial intelligence, robotics, computer vision, and physical AI - including researchers, software developers, robotics engineers, scientific coders exploring physical automation, students, and AI enthusiasts.
Review the Gemini Robotics ER 2 Model Card and Gemini API Robotics Overview.
Have a Google account to open and run the hands-on exercise in Google Colab.
Obtain a Gemini API key from Google AI Studio (we can also walk through key setup during the session).
This is an exciting opportunity to directly talk to and learn from researchers using cutting edge Robotics models. As our speakers are donating their time to walk us through this pivotal technology we want to ensure an engaged event. Please reserve your spot only if you are committed to joining us for this session.
Limited availability to join in-person (lunch to follow).
Or join virtually to engage with discussion.
Friday, September 25, 2026
2:00 AM – 3:00 AM (UTC)
University of Sydney
PhD Candidate