A Multi-Source Heterogeneous Data Empowerment Method for AI Product Profiling that Integrates Distant Supervision and Large Language Models
Yuan Wang, Yurong Li, Haotan Liu, Wei Ma, Wenhui Hu, Yong HuangAbstract
Purpose
This paper aims to support AI-product evaluation and management by automatically extracting AI products and their multi-dimensional attributes from multi-source heterogeneous data, supporting personalized product-design optimization and enterprise market insight.
Design/methodology/approach
An end-to-end AI-product profiling framework is built over a constructed multi-source heterogeneous dataset. It integrates a Document Image Transformer (DiT) and large language models (e.g., Qwen, ChatGPT) for webpage data, report layout analysis and target-chart recognition; chart processing filters relevance at the title level, with attributes taken from captions and surrounding text rather than pixel content. Product names are extracted by a domain-adapted distantly supervised pipeline (AI-product lexicon, dependency-based context filtering and LLM-assisted boundary disambiguation with BERT/RoBERTa-CRF), and fine-grained attributes by a LoRA-fine-tuned Qwen-32B model.
Findings
On the constructed dataset, the framework attains F1 above 90 % on all three core tasks: target-chart recognition (BERT, 96.5 %), product-name extraction (RoBERTa + CRF, 91.7 %) and attribute extraction (Qwen-32B, 92.4 %). The distantly supervised pipeline is competitive with larger generative models while relying mainly on weakly labelled data and targeted human verification.
Research limitations
Performance depends on lexicon coverage and on the quality and completeness of source data; some attributes (e.g., reproducibility or license of open-source products) are sparsely documented and remain hard to extract, occasionally preventing structured output. Broader cross-domain validation is left to future work.
Practical implications
The framework provides a reusable pipeline for constructing and updating AI-product profiles from heterogeneous sources, supporting market insight, technology decision-making and personalized product-design optimization; its potential to reduce manual effort is a design implication rather than a quantified outcome.
Originality/value
The study formulates AI-product profiling as a distinct task and contributes a framework that adapts – rather than merely combines – distant supervision, document-layout analysis and LLM-based attribute extraction to three defining characteristics of AI products: data heterogeneity, attribute specialization and dynamic iteration.