DOI: 10.12688/f1000research.190597.1 ISSN: 2046-1402

A Confound-Annotated Course-Catalog Dataset with Parsed Conjunctive-Normal-Form Prerequisites for Twenty-Two Universities of the United Arab Emirates

Sherzod Turaev, Saja Al-Dabet, Mary John, Mamoun Awad, Nazar Zaki, Khaled Shuaib
Background The prerequisite structure of a curriculum governs the order in which knowledge is built and the time a student needs to graduate, yet the datasets on which its quantitative study depends record prerequisites as flat lists of required courses, are usually confined to a single institution at a single point in time, and discard the alternative requirements that catalogs state with the word or. No machine-readable, multi-institution curriculum dataset of the national higher-education system of the United Arab Emirates previously existed, and the source catalogs are published as unstructured or semi-structured documents ranging from rendered web pages to image-based study plans. Methods From the published catalogs of twenty-two universities we extracted, normalized, reconciled, and audited a uniform course-level record set of 52,802 courses across 56 catalog editions, together with 33,945 parsed prerequisite relations and a layer of 756-degree programs. Prerequisites were parsed into conjunctive-normal form, so that alternatives are preserved as boolean structure rather than flattened into mandatory codes; 7,030 of the relations, or 20.7 percent, are alternatives. One institution is covered by an eleven-edition panel spanning the decade from 2015/16 to 2025/26, and every record is annotated with the measurement confounds, namely notation drift, selective disclosure, subject-code renumbering, and prerequisite-operator ambiguity, that document-derived curriculum data inherit from their sources. Conclusions A correctness audit, coded blind against the source under a protocol fixed before coding began, places course-code agreement at 100 percent, credit and prerequisite agreement in the mid-to-high nineties, and substantive title accuracy near 99 percent. The dataset is released with the complete harvesting, parsing, and verification code and the full audit bundle, and it supports the cross-institution and longitudinal study of prerequisite structure for a national system for which such analysis was not previously feasible.