DOI: 10.1145/3837112 ISSN: 2836-6573
FlowPipe: LLM-Enhanced Conditional Generative Flow Networks for Data Preparation Pipeline Construction
Kunyu Ni, Lei Cao, Jie He, Xiaotong Zhang, Jianfeng Jin, Junyu Dong, Yanwei Yu
Data preparation pipelines are a primary mechanism for improving data quality in machine learning workflows, transforming raw, error-prone tables into learning-ready data via sequential cleaning and feature transformation operators. However, automated pipeline construction remains computationally prohibitive due to the combinatorial complexity of operator sequences and the high cost of end-to-end evaluation. While Reinforcement Learning provides a principled discrete search paradigm, state-of-the-art (SOTA) Multi-DQN architectures suffer from three fundamental limitations:
structural dissonance,
where decoupled value estimators hinder long-horizon credit assignment;
semantic detachment,
where dataset context is treated as a superficial additive bias rather than strictly conditioning the agent's reasoning; and
exploration inefficiency
in a vast, sparse optimization landscape with many invalid pipeline states. To address these challenges, we propose
FlowPipe
, a unified framework that reformulates pipeline synthesis as conditional probabilistic flow generation over a directed acyclic graph. First, FlowPipe employs a Conditional Generative Flow Networks (C-GFlowNets) optimized via a Trajectory Balance objective, establishing a direct gradient path from terminal validation rewards to early actions to ensure holistic credit assignment. Second, to resolve semantic detachment, we introduce
Deep Semantic Modulation
via Feature-wise Linear Modulation (FiLM), which allows LLM-derived logical priors to multiplicatively modulate the policy's internal activation maps, thereby structurally adapting the decision logic to the dataset context. Finally, to mitigate exploration inefficiency, we incorporate failure awareness into the flow objective to prune semantically invalid states early and concentrate search mass on high-potential regions. Extensive experiments on two benchmark suites comprising 74 real-world datasets show that FlowPipe significantly outperforms SOTA baselines, improving accuracy by an average of
11.96%
while achieving a
12.5×
speedup in training convergence. The source code is available at https://github.com/KunyuNi/FlowPipe.