Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Benchmarking Large Language Models for Dental Clinical Decision Support: A BERT Score Analysis of Claude Opus 4.5

Author 1: Achmad Zam Zam Aghasy Author 2: Muhammad Lutfan Lazuardi Author 3: Hari Kusnanto Josef
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 17, No. 1 · Published 2026

DOI: https://doi.org/10.14569/IJACSA.2026.0170158

Abstract

The integration of Large Language Models (LLMs) into clinical decision support systems represents a significant advancement in healthcare informatics. This study presents a comprehensive evaluation framework for benchmarking LLM-generated dental treatment recommendations using BERT Score as the primary semantic similarity metric. We evaluated Claude Opus 4.5 as a Clinical Decision Support System (CDSS) across 116 dental case reports extracted from the Case Reports in Dentistry journal (2024-2025), spanning nine dental specialties. The BERT Score was calculated using the RoBERTa-large model to measure semantic alignment between AI-generated treatment plans and gold-standard published treatments. Results demonstrated strong semantic alignment with a mean BERT Score F1 of 0.8199 with a standard deviation of 0.0144 (95 per cent confidence interval: 0.8172-0.8225), significantly exceeding the 0.80 threshold (t = 14.90, p < 0.001, d = 1.38). Cross-specialty analysis revealed consistent performance across all nine dental domains (Kruskal-Wallis H = 3.07, p = 0.879), indicating robust generalizability. A significant negative correlation was observed between BERT Score and response time (ρ = -0.371, p < 0.001), suggesting a speed-accuracy trade-off in LLM reasoning. This study contributes a reproducible benchmarking methodology for evaluating LLM performance in specialized clinical domains and demonstrates the potential of BERT Score as a scalable evaluation metric for AI-generated clinical text.

Keywords

How to Cite this Article

Aghasy, A. Z. Z., Lazuardi, M. L., & Josef, H. K. (2026). Benchmarking Large Language Models for Dental Clinical Decision Support: A BERT Score Analysis of Claude Opus 4.5. International Journal of Advanced Computer Science and Applications, 17(1). https://doi.org/10.14569/IJACSA.2026.0170158

Aghasy, Achmad Zam Zam, et al.. "Benchmarking Large Language Models for Dental Clinical Decision Support: A BERT Score Analysis of Claude Opus 4.5." International Journal of Advanced Computer Science and Applications, vol. 17, no. 1, 2026, https://doi.org/10.14569/IJACSA.2026.0170158.

@article{Aghasy2026,
  title     = {Benchmarking Large Language Models for Dental Clinical Decision Support: A BERT Score Analysis of Claude Opus 4.5},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {17},
  number    = {1},
  year      = {2026},
  publisher = {The Science and Information Organization},
  author    = {Achmad Zam Zam Aghasy and Muhammad Lutfan Lazuardi and Hari Kusnanto Josef},
  doi       = {10.14569/IJACSA.2026.0170158},
  url       = {https://doi.org/10.14569/IJACSA.2026.0170158}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.