Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study

7Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Background Multimodal large-language models (LLMs) have recently gained the ability to interpret images. However, their accuracy on anatomy tasks remains unclear. Methods A cross-sectional, atlas-based benchmark study was conducted in which six publicly accessible chat endpoints, including paired “deep-reasoning” and “low-latency” modes from OpenAI, Microsoft Copilot, and Google Gemini, identified 260 numbered landmarks on 26 high-resolution plates from a classical anatomic atlas. Each image was processed twice per model. Two blinded anatomy lecturers scored responses, including accuracy, run-to-run consistency, and per-label latency, which were compared with χ2 and Kruskal–Wallis tests. Results Overall accuracy differed significantly among models (χ2=73.2, P<0.001). OpenAI o3 achieved the highest correctness (53.1%), outperforming its sibling GPT-4o and both Copilot variants, but required the longest inference time. Musculoskeletal structures were recognised more accurately than neurovascular targets, reflecting the greater visual complexity of fine vessels and nerves. Consistency ranged from 43.5% (Gemini Flash) to 65.0% (GPT-4o); deeper modes improved stability for Copilot and Gemini but not accuracy. Median per-label latency spanned three orders of magnitude, from 0.5 s for Gemini Flash to 33 s for o3. Conclusions Currently, publicly available multimodal LLMs can only moderately identify oral and maxillofacial landmarks, and no endpoint is sufficiently reliable to serve as a stand-alone answer key. Higher accuracy was achievable with a trade-off in latency, highlighting the need for domain-specific tuning and human oversight. This atlas benchmark study introduced here provides a reproducible yardstick for future model refinement and educational integration.

Cite

CITATION STYLE

APA

Nguyen, V. A., Vuong, T. Q. T., & Nguyen, V. H. (2025). Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study. PLOS ONE, 20(10 October). https://doi.org/10.1371/journal.pone.0335775

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free