Ollama Accelerates Gemma 4 on Apple Silicon with Multi-Token Prediction

Ollama 0.31 now offers significantly faster Gemma 4 performance on Apple Silicon, achieving up to 90% speed improvement through multi-token prediction (MTP) powered by MLX. This enhancement is particularly beneficial for coding agents, as shown by the Aider polyglot benchmark.

Source: Ollama Blog