PipelineEDU: Creating Synthetic Data for Educational Research and Applications
Austin Nicolas, Shahnewaz Karim SakibCurrent educational AI research often relies on large empirical datasets, whose usage raises ethical concerns due to the risk of exposing sensitive student information, including personal identifiers. To address the gap between the need for data-driven educational AI research and ethical data sourcing, this work introduces a purely synthetic dataset generated through a replicable pipeline. Because the dataset is generated from aggregate distributions and specific mappings rather than individual student records, it reduces risks associated with exposing real student data. However, limited similarity to real-world data can reduce downstream utility. The proposed pipeline generates a synthetic educational dataset, with no real-world counterpart, that includes student demographics, educational history, career aspirations, and more. Data is generated by aggregate empirical distributions and weighted mappings between explicitly related variables. Some mappings follow correlations reported in prior studies, while others aim to enable exploratory analysis.