Is it agentic enough? Benchmarking open models on your own tooling

Details in article.

Source: Hugging Face