Dataset Distillation by Tabular Alignment via Moment Embeddings
Eduard Barnoviciu, Corneliu FloreaDataset distillation has achieved strong results in computer vision, but is largely underexplored in the tabular domain. We introduce a tabular dataset distillation method that projects the data through many random embedders (views) to achieve invariance to transformation and to focus on the consistency between real and synthetic sets. The synthetic set is determined through a formulation of distribution matching between the many-view projection of the original and distilled dataset. Building on this approach, our proposal achieves three goals: (1) we formulate the Tabular Alignment via Moment Embeddings (TAME) method and, by extensive empirical evaluation, we show its efficiency; (2) we evaluate TAME on a benchmark of 18 tabular datasets, with strong baselines and evaluation metrics; and (3) we present a structured set of studies analyzing the impact of losses, dataset geometry, embedder architecture, instances per class (IPC) budget and downstream classifiers. We show that the proposed TAME method consistently surpasses baselines on neural classifiers, while remaining competitive with strong coreset baselines on tree-based classifiers (RF, XGBoost). Performance is further increased, especially for tree-based classifiers with a lightweight validation method. Extensive evaluation, including on additional large-scale sets and ablation experiments, allow a better understanding of the method.