Abstract:We study the problem of constructing ε-coresets for the (k, z)-clustering problem in a doubling metric M(X, d). An ε-coreset is a weighted subset S ⊆ X with weight function w : S → ℝ ≥0 , such that for any k-subset C ∈ [X] k , it holds that Σ x∈S w(x) · d z (x, C) ∈ (1 ± ε) · Σ x∈X d z (x, C). We present an efficient algorithm that constructs an ε-coreset for the (k, z)-clustering problem in M(X, d), where the size of the coreset only depends on the parameters k, z, ε and the doubling dimension ddim(M). To the best of our knowledge, this is the first efficient c-coreset construction of size independent of |X| for general clustering problems in doubling metrics. To this end, we establish the first relation between the doubling dimension of M(X, d) and the shattering dimension (or VC-dimension) of the range space induced by the distance d. Such a relation is not known before, since one can easily construct instances in which neither one can be bounded by (some function of) the other. Surprisingly, we show that if we allow a small (1 ± ε)-distortion of the distance function d (the distorted distance is called the smoothed distance function), the shattering dimension can be upper bounded by O(ε -O(ddim(M)) ). For the purpose of coreset construction, the above bound does not suffice as it only works for unweighted spaces. Therefore, we introduce the notion of τ-error probabilistic shattering dimension, and prove a (drastically better) upper bound of O(ddim(M)·log(1/ε)+log log 1/τ) for the probabilistic shattering dimension for weighted doubling metrics. As it turns out, an upper bound for the probabilistic shattering dimension is enough for constructing a small coreset. We believe the new relation between doubling and shattering dimensions is of independent interest and may find other applications. Furthermore, we study robust coresets for (k, z)-clustering with outliers in a doubling metric. We show an improved connection between α-approximation and robust coresets. This also leads to improvement upon the previous best known bound of the size of robust coreset for Euclidean space [Feldman and Langberg, STOC 11]. The new bound entails a few new results in clustering and property testing. As another application, we show constant-sized (ε, k, z)centroid sets in doubling metrics can be constructed by extending our coreset construction. Prior to our result, constantsized centroid sets for general clustering problems were only known for Euclidean spaces. We can apply our centroid set to accelerate the local search algorithm (studied in [Friggstad et al., FOCS 2016]) for the (k, z)-clustering problem in doubling metrics.

Coresets for Clustering in Euclidean Spaces: Importance Sampling is Nearly Optimal

On Coresets for Clustering in Small Dimensional Euclidean Spaces

On Optimal Coreset Construction for Euclidean $(k,z)$-Clustering

Coresets for Clustering with Missing Values

Near-optimal Coresets for Robust Clustering

Epsilon-Coresets for Clustering (with Outliers) in Doubling Metrics.

Coresets for Clustering in Geometric Intersection Graphs

Improved Coresets for Clustering with Capacity and Fairness Constraints

Coresets for Constrained Clustering: General Assignment Constraints and Improved Size Bounds

Sensitivity Sampling for $k$-Means: Worst Case and Stability Optimal Coreset Bounds

Coresets for Clustering with Fairness Constraints.

Near-Optimal Quantum Coreset Construction Algorithms for Clustering

Coresets for Clustering in Graphs of Bounded Treewidth

D S ] 2 0 Ju n 20 19 Coresets for Clustering with Fairness Constraints

Coresets for Kernel Clustering

Coresets for Time Series Clustering

Deterministic Clustering in High Dimensional Spaces: Sketches and Approximation

The Johnson-Lindenstrauss Lemma for Clustering and Subspace Approximation: From Coresets to Dimension Reduction

Provable Imbalanced Point Clustering