Ollama 0.31 now offers significantly faster Gemma 4 performance on Apple Silicon, achieving up to 90% speed improvement through multi-token prediction (MTP) powered by MLX. This enhancement is particularly beneficial for coding agents, as shown by the Aider polyglot benchmark.
Source: Ollama Blog