Abstract:Logs are a first-hand source of information for software maintenance and failure diagnosis. Log parsing, which converts semi-structured log messages into structured templates, is a prerequisite for automated log analysis tasks such as anomaly detection, troubleshooting, and root cause analysis. However, existing log parsers fail in real-world systems for three main reasons. First, traditional heuristics-based parsers require handcrafted features and domain knowledge, which are difficult to generalize at scale. Second, existing large language model-based parsers rely on periodic offline processing, limiting their effectiveness in real-time use cases. Third, existing online parsing algorithms are susceptible to log drift, where slight log changes create false positives that drown out real anomalies. To address these challenges, we propose HELP, a Hierarchical Embeddings-based Log Parser. HELP is the first online semantic-based parser to leverage LLMs for performant and cost-effective log parsing. We achieve this through a novel hierarchical embeddings module, which fine-tunes a text embedding model to cluster logs before parsing, reducing querying costs by multiple orders of magnitude. To combat log drift, we also develop an iterative rebalancing module, which periodically updates existing log groupings. We evaluate HELP extensively on 14 public large-scale datasets, showing that HELP achieves significantly higher F1-weighted grouping and parsing accuracy than current state-of-the-art online log parsers. We also implement HELP into Iudex's production observability platform, confirming HELP's practicality in a production environment. Our results show that HELP is effective and efficient for high-throughput real-world log parsing.

Experience Mining Google's Production Console Logs.

Mining Console Logs for Large-Scale System Problem Detection.

Detecting Large-Scale System Problems by Mining Console Logs

Online System Problem Detection by Mining Patterns of Console Logs

System Problem Detection by Mining Console Logs - Escholarship

Experience Report: Log Mining Using Natural Language Processing and Application to Anomaly Detection

A Distributed Data Mining System Framework for Mobile Internet Access Log Based on Hadoop.

Towards Automated Log Parsing for Large-Scale Log Data Analysis

The Unified Logging Infrastructure for Data Analytics at Twitter

A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?

Tools and Benchmarks for Automated Log Parsing

LogSed: Anomaly Diagnosis Through Mining Time-Weighted Control Flow Graph in Logs

Contextual analysis of program logs for understanding system behaviors

Finding Needles in the Haystack: Harnessing Syslogs for Data Center Management

GLog: A high level graph analysis system using MapReduce

Data Preprocessing Method and API for Mining Processes from Cloud-Based Application Event Logs

HELP: Hierarchical Embeddings-based Log Parsing

Visual Analysis of E-Commerce User Behavior Based on Log Mining

A Flexible Architecture for Statistical Learning and Data Mining from System Log Streams

Experience Report: Deep Learning-based System Log Analysis for Anomaly Detection