DOI: 10.1093/bjs/znag087.010 ISSN: 0007-1323

SP 1.04 Can ChatGPT Be “Taught” to Accurately Triage A GP Referral Letter for Hernia Repair? A Comparison with Consultant Decisions

Paul Ingley, Shivam Bhanderi, Amelia Mclaughlin, Abdulrahman Jama, Chandra Cheruvu

Abstract

Aims

This study assessed the capability of ChatGPT to emulate a Consultant General Surgeon in triaging a GP hernia referral and its ability to learn to do so when presented hernia signs and symptoms to assess its potential for autonomy.

Methods

60 patients, referred to a UK tertiary centre for hernia repair were analysed. Cardinal features of a hernia (i.e. lump, pain, reducibility, cough impulse and location) were extracted along with patient comorbidities and age. ChatGPT 5 was asked to review these letters initially in a raw format and subsequently the processed form alongside rules for triage as provided by the consultant general surgeon. ChatGPT was then asked to assign each case an outcome from clinic review, discharge to GP or list for surgery. Concordance rates and statistical agreement was assessed with Cohens k and McNemar’s test.

Results

ChatGPT agreed with consultant in 76% of cases when provided with letter alone (k = 0.2, McNemar 0.04) and in 87% of cases with processed data was (k = 0.4, McNemar 0.03) agreement with specialist. ChatGPT tended to favour clinic review rather than discharge or adding to waiting list.

Conclusion

Whilst ChatGPT achieved 76% concordance with consultant decision with raw data and 87% with processed, it accurately identified the hernia type in all cases and always offered a safe triage outcome. AI is not yet capable of replacing specialist decision making, it may serve as a valuable decision support tool.

More from our Archive