Google Introduces Speculative Decoding for Faster Gemma 4 AI Models
Google’s Gemma 4 open AI models now use an experimental technique called speculative decoding to accelerate text generation by up to three times. This approach enables faster local inference, making advanced large language models more accessible to users on their own hardware.