March 2026 in “ArXiv.org” This review presents a comprehensive evaluation of medical reasoning using large language models, highlighting a significant gap between exam-level performance and true clinical decision-making accuracy.
September 2024 in “arXiv (Cornell University)” This study evaluated various NLP models for detecting bias in medical curricula, finding that fine-tuned BERT models perform well, whereas LLMs, despite being state-of-the-art in many tasks, are unsuitable for this application.
1 citations
,
September 2025 in “PLOS Digital Health” This study found that all four evaluated large language models generated a significant proportion of inappropriate responses when prompted with or without LGBTQIA+ identity terms, with more severe bias observed in prompts mentioning LGBTQIA+ identities.
1 citations
,
October 2023 In this study, the authors found that syntax-based neural networks performed comparably to pre-trained Transformers on tasks involving definitely unseen sentences, suggesting they are a more transparent and parameter-efficient alternative for certain Natural Language Processing applications.
January 2026 in “Diagnostics” This study reports that publicly available large language models are currently less accurate than human experts in diagnosing trichoscopic images, suggesting the need for further development and specialized training for these AI tools in trichology.