Abstract:Nonparametric classification models, such as K-Nearest Neighbor (KNN), have become particularly powerful tools in machine learning and data mining, due to their simplicity and flexibility. However, the testing time of the KNN classifier becomes unacceptable and the KNN's performance deteriorates significantly when applied to data sets with millions of dimensions. We observe that state-of-the-art approximate nearest neighbor (ANN) methods aim to either reduce the number of distance comparisons based on tree structure or decrease the cost of distance computation by dimension reduction methods. In this paper, we propose a doubly approximate nearest neighbor classification strategy, which marries the two branches which compress the dimensions for decreasing distance computation cost as well as reduce the number of distance comparison instead of full scan. Under this strategy, we build a compressed dimensional tree (CD-Tree) to avoid unnecessary distance calculations. In each decision node, we propose a novel feature selection paradigm by optimizing the feature selection vector as well as the separator (indicator variables for splitting instances) with the maximum margin. An efficient algorithm is then developed to find the globally optimal solution with convergence guarantee. Furthermore, we also provide a data-dependent generalization error bound for our model, which reveals a new insight for the design of ANN classification algorithms. Our empirical studies show that our algorithm consistently obtains competitive or better classification results on all data sets, yet we can also achieve three orders of magnitude faster than state-of-the-art libraries on very high dimensions.

Combining the Strength of Pattern Frequency and Distance for Classification

Combining Multiple K-Nearest Neighbor Classifiers Using Different Distance Functions

Multiple Classifiers Combination Based on Specialists' FIelds

Combining KNN algorithm and other classifiers.

Improving Performance of the k-Nearest Neighbor Classifier by Combining Feature Selection with Feature Weighting.

Decision by Maximum of Posterior Probability Average with Weights: a Method of Multiple Classifiers Combination

Improved KNN Algorithm based on Probability and Adaptive K Value.

Combining Feature Selection with Feature Weighting for k-NN Classifier

Doubly Approximate Nearest Neighbor Classification

Combining Classifiers in Software Quality Prediction: A Neural Network Approach

Distance-based Features in Pattern Classification

Feature Combination and the kNN Framework in Object Classification

Optimal Combination of Feature Weight Learning and Classification Based on Local Approximation

Optimal Combination of Feature Selection and Classification Via Local Hyperplane Based Learning Strategy

Hybrid Dynamic K-Nearest-neighbour and Distance and Attribute Weighted Method for Classification

K-Nearest Neighbour and Support Vector Machine Hybrid Classification

A New Multiple Classifiers Combination Algorithm

Decision Combination Based on the Characterisation of Predictive Accuracy.

Discriminative feature combination selection for enhancing multiclass classification

Evaluating Classifier Combination in Object Classification

A New Supervised Feature Selection Method for Pattern Classification.