Ollama Accelerates Gemma 4 Performance on Apple Silicon with Multi-Token Prediction

Ollama 0.31 now offers significantly faster Gemma 4 performance on Apple Silicon, achieving up to a 90% speed increase for coding agents. This improvement is enabled by multi-token prediction (MTP) powered by MLX. Benchmarks using the Aider polyglot confirm these substantial performance gains.

Source: Ollama