6 citations
,
July 2024 in “The Journal of the American Board of Family Medicine” This study found that while GPT-4 shows high accuracy and efficiency in clinical decision making, physicians' critical thinking and lifelong learning skills remain essential, particularly in addressing and interpreting AI errors in medical settings.
September 2024 in “arXiv (Cornell University)” This study evaluated various NLP models for detecting bias in medical curricula, finding that fine-tuned BERT models perform well, whereas LLMs, despite being state-of-the-art in many tasks, are unsuitable for this application.
June 2025 in “British Journal of Dermatology” This study detailed the implementation of an autonomous AI device in an NHS skin cancer pathway, showing that it achieved a sensitivity of 97.3% for diagnosing skin cancers and exceeded sensitivity targets compared to specialists with a negative predictive value over 99.7%.
March 2026 in “Journal of Evidence-Based Medicine” This study found that AI chatbots provided highly accurate responses to plastic surgery exam questions, indicating their reliability as learning tools for medical students in this field.
February 2024 in “arXiv (Cornell University)” In this study, researchers found that differences in skin condition distribution are the main source of errors when AI algorithms classify dermatological conditions from new, previously unseen sources, and proposed steps to improve their generalizability based on available information.