A Critical Evaluation of Evaluations for Long-Form Question Answering

    January 2023
    Fu-Liu Xu, Yining Song, Mohit Iyyer, Eunsol Choi
    Image
    Studysummary This study highlights challenges in evaluating long-form question answering, suggesting that current automated metrics do not align with human judgments, and it advocates for a multi-faceted evaluation approach.
    Automatically generated from the study's abstract, not written by a person, and not a review of the full paper. Not medical advice or a treatment recommendation. Read the original study, and consult a qualified healthcare professional before changing treatment. Full disclaimer
    Discuss this study in the Community →