Abstract:We study the clustering problem for mixtures of bounded covariance distributions, under a fine-grained separation assumption. Specifically, given samples from a $k$-component mixture distribution $D = \sum_{i =1}^k w_i P_i$, where each $w_i \ge \alpha$ for some known parameter $\alpha$, and each $P_i$ has unknown covariance $\Sigma_i \preceq \sigma^2_i \cdot I_d$ for some unknown $\sigma_i$, the goal is to cluster the samples assuming a pairwise mean separation in the order of $(\sigma_i+\sigma_j)/\sqrt{\alpha}$ between every pair of components $P_i$ and $P_j$. Our contributions are as follows: For the special case of nearly uniform mixtures, we give the first poly-time algorithm for this clustering task. Prior work either required separation scaling with the maximum cluster standard deviation (i.e. $\max_i \sigma_i$) [DKK+22b] or required both additional structural assumptions and mean separation scaling as a large degree polynomial in $1/\alpha$ [BKK22]. For general-weight mixtures, we point out that accurate clustering is information-theoretically impossible under our fine-grained mean separation assumptions. We introduce the notion of a clustering refinement -- a list of not-too-small subsets satisfying a similar separation, and which can be merged into a clustering approximating the ground truth -- and show that it is possible to efficiently compute an accurate clustering refinement of the samples. Furthermore, under a variant of the "no large sub-cluster'' condition from in prior work [BKK22], we show that our algorithm outputs an accurate clustering, not just a refinement, even for general-weight mixtures. As a corollary, we obtain efficient clustering algorithms for mixtures of well-conditioned high-dimensional log-concave distributions. Moreover, our algorithm is robust to $\Omega(\alpha)$-fraction of adversarial outliers.

Learning general Gaussian mixtures with efficient score matching

Learning Mixtures of Gaussians Using Diffusion Models

Learning Mixtures of Gaussians Using the DDPM Objective

Fit Like You Sample: Sample-Efficient Generalized Score Matching from Fast Mixing Diffusions

Bayesian estimation and prediction for certain mixtures

Sample-Efficient Private Learning of Mixtures of Gaussians

SQ Lower Bounds for Learning Bounded Covariance GMMs

SQ Lower Bounds for Learning Mixtures of Linear Classifiers

Learning Mixtures of Arbitrary Distributions over Large Discrete Domains.

Learning Mixtures of Gaussians with Censored Data

Entropic characterization of optimal rates for learning Gaussian mixtures

On the best approximation by finite Gaussian mixtures

Probabilistic Inference from Arbitrary Uncertainty using Mixtures of Factorized Generalized Gaussians

Clustering Mixtures with Almost Optimal Separation in Polynomial Time

Implicit High-Order Moment Tensor Estimation and Learning Latent Variable Models

Privately Learning Mixtures of Axis-Aligned Gaussians

Fast deep mixtures of Gaussian process experts

Clustering Mixtures of Bounded Covariance Distributions Under Optimal Separation

The Informativeness of K -Means for Learning Mixture Models

Pseudo Independent Conditional Approximation for Training the Mixtures of Gaussian Processes