TexSortML: A Dataset and Baseline Evaluation for Garment Classification in Automated Textile Sorting
Julian Schanz, Martin Kohnle, Fabian Kopf, Sebastian Geldhäuser, Mesut Cetin, Alexandra TeynorThe growing volume of textile waste, driven by fast fashion and shortened garment lifespans, underscores the urgent need for scalable, automated sorting solutions. Progress is hindered by the absence of realistic, public datasets reflecting the variability of industrial textile sorting. We address this gap with TexSortML, a dataset of 751 post-consumer garments across 11 categories, captured under three controlled difficulty stages that simulate increasing handling disorder on a conveyor belt. On this dataset, we compare Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), and CLIP-based zero-shot classifiers under five-fold cross-validation. A ViT fine-tuned directly on our data is the most accurate and most stable model at every stage, losing only 9.3 percentage points between the most and least controlled condition, against 16.6–16.9 points for the CNN baselines. The variant additionally pre-trained on external second-hand clothing data matches it only under controlled presentation and degrades sharply once garments are rotated or crumpled. CLIP-based zero-shot classifiers remain far behind and degrade the fastest. Rather than proposing a new architecture, this work contributes the dataset and a reproducible benchmark, and shows that extra training data can hurt when it is cleaner than the deployment condition. The dataset is publicly available on Zenodo.