Abstract:Large Language Models (LLMs) have increasingly become pivotal in content generation with notable societal impact. These models hold the potential to generate content that could be deemed harmful.Efforts to mitigate this risk include implementing safeguards to ensure LLMs adhere to social ethics.However, despite such measures, the phenomenon of "jailbreaking" -- where carefully crafted prompts elicit harmful responses from models -- persists as a significant challenge. Recognizing the continuous threat posed by jailbreaking tactics and their repercussions for the trustworthy use of LLMs, a rigorous assessment of the models' robustness against such attacks is essential. This study introduces an comprehensive evaluation framework and conducts an large-scale empirical experiment to address this need. We concentrate on 10 cutting-edge jailbreak strategies across three categories, 1525 questions from 61 specific harmful categories, and 13 popular LLMs. We adopt multi-dimensional metrics such as Attack Success Rate (ASR), Toxicity Score, Fluency, Token Length, and Grammatical Errors to thoroughly assess the LLMs' outputs under jailbreak. By normalizing and aggregating these metrics, we present a detailed reliability score for different LLMs, coupled with strategic recommendations to reduce their susceptibility to such vulnerabilities. Additionally, we explore the relationships among the models, attack strategies, and types of harmful content, as well as the correlations between the evaluation metrics, which proves the validity of our multifaceted evaluation framework. Our extensive experimental results demonstrate a lack of resilience among all tested LLMs against certain strategies, and highlight the need to concentrate on the reliability facets of LLMs. We believe our study can provide valuable insights into enhancing the security evaluation of LLMs against jailbreak within the domain.

A Voter-Based Stochastic Rejection-Method Framework for Asymptotically Safe Language Model Outputs

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Black-box Uncertainty Quantification Method for LLM-as-a-Judge

Pareto Optimal Learning for Estimating Large Language Model Errors

Risk Aware Benchmarking of Large Language Models

Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation

A Framework for Real-time Safeguarding the Text Generation of Large Language Model

Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations

Output Scouting: Auditing Large Language Models for Catastrophic Responses

Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward

VulnLLMEval: A Framework for Evaluating Large Language Models in Software Vulnerability Detection and Patching

A Simple and Provable Scaling Law for the Test-Time Compute of Large Language Models

Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability

A Formalism and Approach for Improving Robustness of Large Language Models Using Risk-Adjusted Confidence Scores

Reranking Laws for Language Generation: A Communication-Theoretic Perspective

LLM4VV: Exploring LLM-as-a-Judge for Validation and Verification Testsuites

Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

A Watermark for Low-entropy and Unbiased Generation in Large Language Models

Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks

A Probabilistic Perspective on Unlearning and Alignment for Large Language Models