4 citations
,
May 2023 in “arXiv (Cornell University)” This study found that current automatic text generation metrics do not predict human preference judgments for long-form question answering, and suggests adopting a multi-faceted evaluation approach focusing on aspects like factuality and completeness.
4 citations
,
November 2023 in “ArXiv.org” This study demonstrates that a proposed multi-stage framework improves the accuracy and faithfulness of drug-related responses generated by language models, compared to traditional methods.
March 2026 in “ArXiv.org” This review presents a comprehensive evaluation of medical reasoning using large language models, highlighting a significant gap between exam-level performance and true clinical decision-making accuracy.
March 2026 in “Journal of Evidence-Based Medicine” This study found that AI chatbots provided highly accurate responses to plastic surgery exam questions, indicating their reliability as learning tools for medical students in this field.
1 citations
,
September 2025 in “PLOS Digital Health” This study found that all four evaluated large language models generated a significant proportion of inappropriate responses when prompted with or without LGBTQIA+ identity terms, with more severe bias observed in prompts mentioning LGBTQIA+ identities.