Abstract:With the ever-increasing quantity and variety of data worldwide, the Web has become a rich repository of mathematical formulae. This necessitates the creation of robust and scalable systems for Mathematical Information Retrieval, where users search for mathematical information using individual formulae (query-by-expression) or a combination of keywords and formulae. Often, the pages that best satisfy users' information needs contain expressions that only approximately match the query formulae. For users trying to locate or re-find a specific expression, browse for similar formulae, or who are mathematical non-experts, the similarity of formulae depends more on the relative positions of symbols than on deep mathematical semantics. We propose the Maximum Subtree Similarity (MSS) metric for query-by-expression that produces intuitive rankings of formulae based on their appearance, as represented by the types and relative positions of symbols. Because it is too expensive to apply the metric against all formulae in large collections, we first retrieve expressions using an inverted index over tuples that encode relationships between pairs of symbols, ranking hits using the Dice coefficient. The top-k formulae are then re-ranked using MSS. Our approach obtains state-of-the-art performance on the NTCIR-11 Wikipedia formula retrieval benchmark and is efficient in terms of both index space and overall retrieval time. Retrieval systems for other graphical forms, including chemical diagrams, flowcharts, figures, and tables, may also benefit from adopting our approach.

The Tangent Search Engine: Improved Similarity Metrics and Scalability for Math Formula Search

Formula Citation Graph Based Mathematical Information Retrieval

A Mathematics Retrieval System for Formulae in Layout Presentations

Wikimirs: A Mathematical Information Retrieval System For Wikipedia

Performance Evaluation and Optimization of Math-Similarity Search

Formula Ranking Within an Article.

A Symbol Dominance Based Formulae Recognition Approach For Pdf Documents

ICST Math Retrieval System for NTCIR-11 Math-2 Task.

MathIRs: Retrieval System for Scientific Documents

The Effectiveness of Graph Contrastive Learning on Mathematical Information Retrieval

Mathematical Information Retrieval Trends and Techniques

Wikimirs 3.0: A Hybrid Mir System Based On The Context, Structure And Importance Of Formulae In A Document

Discovering Mathematical Objects of Interest -- A Study of Mathematical Notations

A Semantic Search Engine for Mathlib4

Combining Text and Formula Queries in Math Information Retrieval: Evaluation of Query Results Merging Strategies

Discovery and Recognition of Formula Concepts using Machine Learning

Making Math Searchable in Wikipedia

Which one is better: presentation-based or content-based math search?

Methods and Tools to Advance the Retrieval of Mathematical Knowledge from Digital Libraries for Search-, Recommendation-, and Assistance-Systems

Semantic Preserving Bijective Mappings of Mathematical Formulae between Document Preparation Systems and Computer Algebra Systems

MSQL: Efficient Similarity Search in Metric Spaces Using SQL.