Ollama 0.31 now offers significantly faster Gemma 4 performance on Apple Silicon, achieving up to a 90% speed increase for coding agents. This improvement is enabled by multi-token prediction (MTP) powered by MLX. Benchmarks using the Aider polyglot confirm these substantial performance gains.
Source: Ollama