MM‐CAD: A Multi‐Modal CAD Dataset and Benchmark for Cross‐Modal Geometric Learning
Anush Bharathi, Ananthakrishnan A, Ramanathan MuthuganapathyAbstract
Computer‐Aided Design (CAD) boosts modern manufacturing, yet design reuse remains constrained by the absence of large, openly available CAD repositories with rich multi‐modal annotations suitable for search/retrieval. Recent large‐scale efforts to annotate public datasets rely on hash‐based redundancy removal that leaves no semantic structure, and on captioning by Vision‐Language Models (VLMs) using rendered images alone, which struggles to capture geometric and procedural information. We introduce MM‐CAD, a multi‐modal CAD dataset designed to level‐up retrieval and retrieval‐augmented generation models for engineering geometry, comprising two complementary parts. MM‐CAD:A brings 33,816 unique CAD models from eleven widely used benchmark datasets under a common identifier scheme, with isometric renderings, point clouds, and humanly‐curated multi‐level text captions, and 4,376 real hand‐drawn user sketches among others. MM‐CAD:B curates 192,626 models from the 1M‐model ABC corpus through a seven‐stage pipeline centered on Manifold‐Aware Adaptive Sampling (MAAS), which organizes models into semantically coherent neighborhoods rather than merely removing duplicates, directly supplying the hard negatives that contrastive retrieval training requires. Every retained model is annotated through a metadata‐grounded pipeline that conditions caption generation on parsed construction sequences rather than rendered views alone, producing three‐level text descriptions, multi‐level contour sketches, a hierarchical application taxonomy, and photorealistic in‐context images that largely preserve source CAD geometry, a modality not previously available at this scale on CAD data. We further introduce a joint retrieval architecture that aligns sketch, text, image, B‐Rep, and point cloud encoders in a single latent space through Matryoshka‐nested contrastive objectives, establishing the first unified cross‐modal retrieval benchmark for large‐scale CAD.