Vision-language large learning model, GPT4V, accurately classifies the Boston Bowel Preparation Scale score

5Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

Introduction Large learning models (LLMs) such as GPT are advanced artificial intelligence (AI) models. Originally developed for natural language processing, they have been adapted for multi-modal tasks with vision-language input. One clinically relevant task is scoring the Boston Bowel Preparation Scale (BBPS). While traditional AI techniques use large amounts of data for training, we hypothesise that vision-language LLM can perform this task with fewer examples. Methods We used the GPT4V vision-language LLM developed by OpenAI, via the OpenAI application programming interface. A standardised prompt instructed the model to grade BBPS with contextual references extracted from the original paper describing the BBPS by Lai et al (GIE 2009). Performance was tested on the HyperKvasir dataset, an open dataset for automated BBPS grading. Results Of 1794 images, GPT4V returned valid results for 1772 (98%). It had an accuracy of 0.84 for two-class classification (BBPS 0-1 vs 2-3) and 0.74 for four-class classification (BBPS 0, 1, 2, 3). Macro-averaged F1 scores were 0.81 and 0.63, respectively. Qualitatively, most errors arose from misclassification of BBPS 1 as 2. These results compare favourably with current methods using large amounts of training data, which achieve an accuracy in the range of 0.8-0.9. Conclusion This study provides proof-of-concept that a vision-language LLM is able to perform BBPS classification accurately, without large training datasets. This represents a paradigm shift in AI classification methods in medicine, where many diseases lack sufficient data to train traditional AI models. An LLM with appropriate examples may be used in such cases.

Cite

CITATION STYLE

APA

Lim, D. Y. Z., Tan, Y. B., Ho, J. R. Y., Carkarine, S., Chew, T. W. V., Ke, Y., … Tan, D. (2025). Vision-language large learning model, GPT4V, accurately classifies the Boston Bowel Preparation Scale score. BMJ Open Gastroenterology, 12(1). https://doi.org/10.1136/bmjgast-2024-001496

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free