SEM-PDPL: Semantic Exposure Graphs for Privacy-Law-Informed Risk Assessment of Public Social-Media Data
Heba IsmailPublic social-media content often contains self-disclosed personal attributes that appear low-risk in isolation but become privacy-relevant when linked across posts, platform accounts, or user-level traces. Existing research has advanced privacy-sensitive content detection, de-anonymization analysis, social-media research ethics, and privacy-compliance workflows; however, limited work operationalizes how personal-data disclosures combine structurally and how these structures can be translated into auditable governance actions. This paper proposes SEM-PDPL, a computational, privacy-law-informed risk-assessment framework for modeling public social-media exposure as semantic exposure graphs and mapping graph patterns to controls aligned with the United Arab Emirates Personal Data Protection Law (PDPL) and compatible with GDPR principles. SEM-PDPL combines governance scoping; a PDPL-informed disclosure taxonomy; hybrid extraction using rule-based methods; named-entity recognition; fine-tuned BERT; and schema-constrained large language model annotation, followed by graph construction at post, platform, corpus, and persona levels. The framework is evaluated on a synthetic multi-platform corpus of 1095 posts generated for 150 personas across 290 platform accounts. Results show that, within this controlled synthetic corpus, fine-tuned BERT provides the strongest extraction performance among six evaluated methods, achieving a macro-F1 of 0.975. Graph analysis shows that exposure density increases with aggregation, rising from 0.275 at post level to 1.000 at corpus level, and from 0.859 at platform level to 0.967 at persona level. Across all graph resolutions, quasi-identifiers emerge as the dominant weighted-degree and betweenness node, indicating that ordinary location, employer, school, and demographic cues often function as bridges connecting sensitive categories such as health and biometric data to identifying information. These findings indicate that, within this controlled corpus, privacy risk in public social-media data is not only attribute-based but also structurally graph-shaped. SEM-PDPL contributes an explainable and reproducible framework for identifying exposure hubs, sensitive bridges, and aggregation risks before applying masking, minimization, exclusion, retention, or review controls. The framework does not automate legal compliance; rather, it provides evidence-based decision support for privacy-aware social-media analytics.