DOI: 10.1145/3837124 ISSN: 2836-6573
SemDistill: Bootstrapping a Low-Latency, Low-Cost Semantic Table Annotator from Noisy LLM
Yuhang Ge, Zhangyan Ye, Yuren Mao, Yunjun Gao
Relational tables widely exist in enterprise data lakes and repositories, yet they often lack semantic context such as column types and relationships, creating a bottleneck for automated data management.
Semantic Table Annotation,
comprising tasks like
Column Type Annotation
(CTA) and
Column Property Annotation
(CPA), is essential to restoring missing semantic context. While Large Language Models (LLMs) excel at these tasks, their scalable deployment is hindered by prohibitive costs, high latency, and privacy information leakage. Distilling LLM capabilities into compact local models offers a viable solution; however, these models inevitably overfit to the structured noise in the teacher's predictions. In this paper, we formalize the setting of label-efficient distillation from noisy LLMs. Our empirical analysis reveals two error patterns: (1)
Systematic Ontological Drift,
a global class-conditional bias where LLMs tend to over-generalize specific concepts; and (2)
High-Confidence Hallucinations,
stubborn high-confidence errors on ambiguous data. To address these, we propose e SemDistill, a framework that bootstraps high-performance local annotators from noisy LLM outputs using only scarce trusted anchors. e SemDistill uses reliability-based partitioning and tier-specific optimization: it models global drift with anchor-guided correction, profiles residual hallucinations through learning trajectories, and trains trusted, semi-trusted, and noisy tiers with different objectives. Extensive experiments demonstrate that e SemDistill effectively addresses the cost-quality trade-off. It outperforms the LLM teacher (GPT-5-mini) and the distillation baseline while reducing inference costs by
3800x
and latency by
2600x
.