A new HTML application by Mike Veerman allows users to visualize and understand the real-world speed of LLM token output. The tool simulates speeds ranging from 5 to 800 tokens per second, offering a practical way to gauge advertised model performance.
Source: Simon Willison