PenGPT: Evaluating Large Language Models as Autonomous Penetration Testing Agents Across OWASP Top-10 Vulnerability Classes

!!!! Bi-Annual Double Blind Peer Reviewed Refereed Journal !!!!

!!!! Open Access Journal !!!!

Abstract: 
A global shortage of skilled penetration testers, rapid software delivery cycles, and expanding attack surfaces have created a critical gap in enterprise vulnerability assessment. While tool-augmented Large Language Models (LLMs) offer a scalable path toward autonomous security testing, their actual performance across standardized vulnerability classes remains under-researched. This study introduces PenGPT, a framework that evaluates four leading LLMs—GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 405B—as autonomous agents across the OWASP Top-10 (2021) categories. Using a Mixed-Method Synthesis of capability analysis and sandboxed, scenario-based adversarial simulations, we analyze each model's ability to identify, exploit, and document vulnerabilities. Our findings reveal significant  performance stratification across models and vulnerability classes, particularly in injection, broken access control, and security misconfiguration. The study contributes a reusable evaluation rubric, a comparative performance matrix, and a governance framework for responsible LLM integration into enterprise penetration testing workflows, aligned with NIST SP 800-115, OWASP Testing Guide v4.2, and emerging EU AI Act obligations for high-risk AI deployment in cybersecurity contexts.
Category: 
Vol19_Issue1
Authors: 
Ms. Sarla Kumari
Download Full Paper: 
Rating: 
0
No votes yet