LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization

arXiv cs.CL / 4/29/2026

📰 NewsDeveloper Stack & InfrastructureIdeas & Deep AnalysisModels & Research

共有:

Key Points

Long document summarization research is held back by evaluation metrics that weakly match human judgments and fail to provide actionable, deficiency-focused guidance.
LongSumEval proposes a unified approach that treats summary quality as answerability and factual alignment using structured question-answer pairs, producing interpretable scores and targeted feedback.
The QA-based framework aims to close the gap between evaluation and generation objectives by generating feedback that directly indicates coverage gaps and factual inconsistencies.
Meta-evaluation across seven benchmarks shows stronger agreement with human judgments than existing metrics, and the feedback enables meaningful self-refinement without retraining.
The authors plan to release code and datasets on GitHub to support reproducibility and further research on verifiable, controllable text generation quality control.

Abstract

Evaluating long document summaries remains the primary bottleneck in summarization research. Existing metrics correlate weakly with human judgments and produce aggregate scores without explaining deficiencies or guiding improvement, preventing effective refinement in applications requiring verifiable accuracy. We introduce LongSumEval, a unified framework bridging evaluation and generation through structured question-answering feedback. The framework operationalizes summary quality as answerability and factual alignment of question-answer pairs, generating interpretable scores and actionable feedback that identifies coverage gaps and factual inconsistencies. This resolves the misalignment where evaluation operates independently of generation objectives. Meta-evaluation of our QA-based evaluation module across seven benchmarks demonstrates substantially stronger agreement with human judgments compared to established metrics. Structured feedback enables significant quality improvements through self-refinement without retraining. By demonstrating that evaluation feedback can serve as executable instructions for generation, this work establishes a generalizable paradigm for aligning assessment with improvement, with direct implications for controllable text generation requiring verifiable accuracy and transparent quality control. All code and datasets will be released in GitHub for reproducibility.