DOI: 10.1177/10519815261491758 ISSN: 1051-9815

Reliability, quality, and readability levels of ChatGPT-5.2 responses regarding office ergonomics: Analysis of content based on Google Trends

Ayşe Ünal, Osman Celbiş

Background

Office ergonomics plays a critical role in preventing work-related musculoskeletal disorders and promoting employee well-being. As large language models such as ChatGPT are increasingly used to obtain health-related information, evaluating the quality, reliability, readability, and practical applicability of artificial intelligence (AI)-generated ergonomics content has become increasingly important.

Objective

To examine the reliability, quality, accuracy, readability, and clinical applicability levels of the responses given by ChatGPT-5.2 to office ergonomics search terms identified by Google Trends, through expert evaluation.

Methods

Nineteen keywords were selected from the terms obtained by searching “office ergonomics” on Google Trends on January 7, 2026. Responses were generated individually in incognito mode using ChatGPT-5.2 (December 2025 version) and assessed by two experts (physiotherapist, forensic medicine specialist) using Journal of the American Medical Association Benchmarking Criteria (JAMA), Global Quality Score (GQS), Modified DISCERN Score (MDS), accuracy scale, Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES) and Patient Education Material Assessment Tool for printed materials (PEMAT-P); inter-rater agreement was analyzed using Intraclass Correlation Coefficients (ICC) and Cronbach's alpha.

Results

The mean GQS score was 3.00 ± 0.74; the MDS score was 2.89 ± 0.31; and the accuracy score was 3.58 ± 0.51. The FKGL level was 7.66 ± 2.76 and the FRES score was 54.17 ± 14.91, with 47.4% of the content requiring university-level readability. The PEMAT-P comprehensibility level was 82.53 ± 11.03, and applicability was 49.21 ± 20.76. Inter-rater agreement was good-to-very high (ICC = 0.779–0.974).

Conclusions

Although ChatGPT-5.2 generally provided acceptable levels of accuracy and content quality in this exploratory analysis, the findings should be interpreted in light of the study's methodological limitations, including the restricted keyword set, single-day assessment, and evaluation of a single model version.