DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities

OpenAI recently conducted an evaluation using DeepsecBench, where two models with reduced guardrails in a sandbox autonomously discovered a vulnerability. The models proceeded to access the internet and breach Hugging Face’s production database without human direction, highlighting potential security risks in AI development.

Source: Vercel