Abstract:Generative AI agents, software systems powered by Large Language Models (LLMs), are emerging as a promising approach to automate cybersecurity tasks. Among the others, penetration testing is a challenging field due to the task complexity and the diverse strategies to simulate cyber-attacks. Despite growing interest and initial studies in automating penetration testing with generative agents, there remains a significant gap in the form of a comprehensive and standard framework for their evaluation and development. This paper introduces AutoPenBench, an open benchmark for evaluating generative agents in automated penetration testing. We present a comprehensive framework that includes 33 tasks, each representing a vulnerable system that the agent has to attack. Tasks are of increasing difficulty levels, including in-vitro and real-world scenarios. We assess the agent performance with generic and specific milestones that allow us to compare results in a standardised manner and understand the limits of the agent under test. We show the benefits of AutoPenBench by testing two agent architectures: a fully autonomous and a semi-autonomous supporting human interaction. We compare their performance and limitations. For example, the fully autonomous agent performs unsatisfactorily achieving a 21% Success Rate (SR) across the benchmark, solving 27% of the simple tasks and only one real-world task. In contrast, the assisted agent demonstrates substantial improvements, with 64% of SR. AutoPenBench allows us also to observe how different LLMs like GPT-4o or OpenAI o1 impact the ability of the agents to complete the tasks. We believe that our benchmark fills the gap with a standard and flexible framework to compare penetration testing agents on a common ground. We hope to extend AutoPenBench along with the research community by making it available under <a class="link-external link-https" href="https://github.com/lucagioacchini/auto-pen-bench" rel="external noopener nofollow">this https URL</a>.

Generative AI for pentesting: the good, the bad, the ugly

Getting pwn'd by AI: Penetration Testing with Large Language Models

AutoPenBench: Benchmarking Generative Agents for Penetration Testing

AI-Augmented Ethical Hacking: A Practical Examination of Manual Exploitation and Privilege Escalation in Linux Environments

AI-Enhanced Ethical Hacking: A Linux-Focused Experiment

Review of Generative AI Methods in Cybersecurity

Generative AI for Cyber Security: Analyzing the Potential of ChatGPT, DALL-E, and Other Models for Enhancing the Security Space

Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements

Is Generative AI the Next Tactical Cyber Weapon For Threat Actors? Unforeseen Implications of AI Generated Cyber Attacks

PentestGPT: An LLM-empowered Automatic Penetration Testing Tool

Generative AI in Cybersecurity

From ChatGPT to ThreatGPT: Impact of Generative AI in Cybersecurity and Privacy

Hacking, The Lazy Way: LLM Augmented Pentesting

Generative Adversarial Network (GAN)-Based Autonomous Penetration Testing for Web Applications

Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security

Identifying and Mitigating the Security Risks of Generative AI

Impacts and Risk of Generative AI Technology on Cyber Defense

Generative Artificial Intelligence and the Future of Software Testing

Accelerating Software Quality: Unleashing the Power of Generative AI for Automated Test-Case Generation and Bug Identification