A Source-Record Audit of DeepSeek-R1 Responses to Romanized Sindhi Prompts
Irum Naz Sodhar, Dil Nawaz Hakro, Abdul Hafeez Buller, Umair Ramzan Sheikh, Suad Mohammed Al Qassabi, Osama Al Rahbi, Akhtar Hussain, Mohammed Izaan KariSindhi language models struggle with Romanized Sindhi as there is variation in the spellings of everyday words, and short prompts do not explicitly state the task or target language. In this study, the remaining source record of a small DeepSeek-R1 probe, which was based on seven canonical prompts and fourteen alleged A/B records, was audited. The screenshot captions were compared with the visible inputs and outputs, and duplication of images was checked. It was only relabeled when there was a single unambiguous canonical prompt that matched the visible input to a record. Task choice and exact target-string occurrence were then coded separately. Six records were auditable, three of which were recovered by relabeling; eight records were not auditable due to being duplicated, contaminated with a supplied answer, associated with a different phrase, or missing the user input or lost after label repair. Four of the six audited records adhered to the experimenter’s task, one provided a reasonably good English translation when the experimenter did not include the target language, and one resulted in an incompatible cross-linguistic reading. Five had the target (English/Sindhi) string and one did not. The screenshots do not demonstrate clear “standard” and “chain of thought” conditions: all audited B-labeled inputs are in the form of a simple prompt and there are no paired wrapper texts remaining. These are descriptive counts, not accuracy estimates and do not provide proof of causal benefit of prompt explicitness. The audit demonstrates the necessity of full prompt–response documentation, independently verified linguistic references, reruns, information-matched prompts, and preregistered evaluation plans in studies with low resources.