Abstract
Background Multimodal large-language models (LLMs) have recently gained the ability to interpret images. However, their accuracy on anatomy tasks remains unclear. Methods A cross-sectional, atlas-based benchmark study was conducted in which six publicly accessible chat endpoints, including paired “deep-reasoning” and “low-latency” modes from OpenAI, Microsoft Copilot, and Google Gemini, identified 260 numbered landmarks on 26 high-resolution plates from a classical anatomic atlas. Each image was processed twice per model. Two blinded anatomy lecturers scored responses, including accuracy, run-to-run consistency, and per-label latency, which were compared with χ2 and Kruskal–Wallis tests. Results Overall accuracy differed significantly among models (χ2=73.2, P<0.001). OpenAI o3 achieved the highest correctness (53.1%), outperforming its sibling GPT-4o and both Copilot variants, but required the longest inference time. Musculoskeletal structures were recognised more accurately than neurovascular targets, reflecting the greater visual complexity of fine vessels and nerves. Consistency ranged from 43.5% (Gemini Flash) to 65.0% (GPT-4o); deeper modes improved stability for Copilot and Gemini but not accuracy. Median per-label latency spanned three orders of magnitude, from 0.5 s for Gemini Flash to 33 s for o3. Conclusions Currently, publicly available multimodal LLMs can only moderately identify oral and maxillofacial landmarks, and no endpoint is sufficiently reliable to serve as a stand-alone answer key. Higher accuracy was achievable with a trade-off in latency, highlighting the need for domain-specific tuning and human oversight. This atlas benchmark study introduced here provides a reproducible yardstick for future model refinement and educational integration.
Cite
CITATION STYLE
Nguyen, V. A., Vuong, T. Q. T., & Nguyen, V. H. (2025). Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study. PLOS ONE, 20(10 October). https://doi.org/10.1371/journal.pone.0335775
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.