A Critical Evaluation of Evaluations for Long-Form Question Answering
May 2023
in “
arXiv (Cornell University)
”
Studysummary This study found that current automatic text generation metrics do not predict human preference judgments for long-form question answering, and suggests adopting a multi-faceted evaluation approach focusing on aspects like factuality and completeness.
Our plain-language summary. Not medical advice or a treatment recommendation. Consult a qualified healthcare professional before changing treatment. Full disclaimer
The study critically evaluates the methods used for assessing long-form question answering (LFQA), involving domain experts from seven areas to provide preference judgments and justifications for answer pairs. The analysis highlights the importance of comprehensiveness in evaluations and reveals that current automatic metrics do not predict human preferences, though some correlate with specific answer aspects like coherence. The authors advocate for a multi-faceted evaluation approach focusing on aspects such as factuality and completeness, rather than a single overall score. They have made their annotations and code publicly available to encourage further research in LFQA evaluation.