Abstract:The rise of Large Language Models (LLMs) has significantly advanced many applications on software engineering tasks, particularly in code generation. Despite the promising performance, LLMs are prone to generate hallucinations, which means LLMs might produce outputs that deviate from users' intent, exhibit internal inconsistencies, or misalign with the factual knowledge, making the deployment of LLMs potentially risky in a wide range of applications. Existing work mainly focuses on investing the hallucination in the domain of natural language generation (NLG), leaving a gap in understanding the types and extent of hallucinations in the context of code generation. To bridge the gap, we conducted a thematic analysis of the LLM-generated code to summarize and categorize the hallucinations present in it. Our study established a comprehensive taxonomy of hallucinations in LLM-generated code, encompassing 5 primary categories of hallucinations depending on the conflicting objectives and varying degrees of deviation observed in code generation. Furthermore, we systematically analyzed the distribution of hallucinations, exploring variations among different LLMs and their correlation with code correctness. Based on the results, we proposed HalluCode, a benchmark for evaluating the performance of code LLMs in recognizing hallucinations. Hallucination recognition and mitigation experiments with HalluCode and HumanEval show existing LLMs face great challenges in recognizing hallucinations, particularly in identifying their types, and are hardly able to mitigate hallucinations. We believe our findings will shed light on future research about hallucination evaluation, detection, and mitigation, ultimately paving the way for building more effective and reliable code LLMs in the future.

De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding

Exploring and Evaluating Hallucinations in LLM-Powered Code Generation

LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation

Code Hallucination

On Mitigating Code LLM Hallucinations with API Documentation

Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Embedding and Gradient Say Wrong: A White-Box Method for Hallucination Detection

CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification

Alleviating Hallucinations of Large Language Models through Induced Hallucinations

CodeMirage: Hallucinations in Code Generated by Large Language Models

Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach

A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation

Towards Mitigating Hallucination in Large Language Models via Self-Reflection

Distinguishing Ignorance from Error in LLM Hallucinations

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Mitigating Entity-Level Hallucination in Large Language Models

Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs

Unravelling the Mysteries of Hallucination in Large Language Models: Strategies for Precision in Artificial Intelligence Language Generation

Developing a Reliable, General-Purpose Hallucination Detection and Mitigation Service: Insights and Lessons Learned