Abstract:In 2006, Geoffrey Hinton proposed the concept of training “Deep Neural Networks (DNNs)” and an improved model training method to break the bottleneck of neural network development. More recently, the introduction of AlphaGo in 2016 demonstrated the powerful learning ability of deep learning and its enormous potential. Deep learning has been increasingly used to develop state-of-the-art software engineering (SE) research tools due to its ability to boost performance for various SE tasks. There are many factors, e.g., deep learning model selection, internal structure differences, and model optimization techniques, that may have an impact on the performance of DNNs applied in SE. Few works to date focus on summarizing, classifying, and analyzing the application of deep learning techniques in SE. To fill this gap, we performed a survey to analyze the relevant studies published since 2006. We first provide an example to illustrate how deep learning techniques are used in SE. We then conduct a background analysis (BA) of primary studies and present four research questions to describe the trend of DNNs used in SE (BA), summarize and classify different deep learning techniques (RQ1), analyze the data processing including data collection, data classification, data pre-processing, and data representation (RQ2). In RQ3, we depicted a range of key research topics using DNNs and investigated the relationships between DL-based model adoption and multiple factors (i.e., DL architectures, task types, problem types, and data types). We also summarized commonly used datasets for different SE tasks. In RQ4, we summarized the widely used optimization algorithms and provided important evaluation metrics for different problem types, including regression, classification, recommendation, and generation. Based on our findings, we present a set of current challenges remaining to be investigated and outline a proposed research road map highlighting key opportunities for future work.

Deep Learning for Approximate Nearest Neighbour Search: A Survey and Future Directions

A comprehensive survey and experimental comparison of graph-based approximate nearest neighbor search

Approximate Nearest Neighbor Search on High Dimensional Data — Experiments, Analyses, and Improvement

Learning to Hash for Indexing Big Data - A Survey

ParlayANN: Scalable and Deterministic Parallel Graph-Based Approximate Nearest Neighbor Search Algorithms

Optimizing Graph-based Approximate Nearest Neighbor Search: Stronger and Smarter

Results of the Big ANN: NeurIPS'23 competition

Fast Approximate Nearest Neighbor Search with the Navigating Spreading-out Graph.

Subspace Collision: An Efficient and Accurate Framework for High-dimensional Approximate Nearest Neighbor Search

Distance Comparison Operators for Approximate Nearest Neighbor Search: Exploration and Benchmark

Approximate Nearest Neighbour Search on Dynamic Datasets: An Investigation

A Survey on Deep Hashing Methods

Complementary Hashing for Approximate Nearest Neighbor Search

Deep Learning for Content-Based Image Retrieval: A Comprehensive Study

A Revisit of Hashing Algorithms for Approximate Nearest Neighbor Search

Local Deep Learning Quantization for Approximate Nearest Neighbor Search

A Survey on Deep Learning for Software Engineering

High Dimensional Similarity Search with Satellite System Graph: Efficiency, Scalability, and Unindexed Query Compatibility

Deep Learning for Matching in Search and Recommendation.

A Comprehensive Survey of Neural Architecture Search: Challenges and Solutions

FLEX: A Fast and Light-weight Learned Index for kNN Search in High-Dimensional Space