Localizing and Mitigating Errors in Long-form Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Sachdeva, Rachneet, Song, Yixiao, Iyyer, Mohit, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
by: Sachdeva, Rachneet, et al.
Published: (2025)
by: Sachdeva, Rachneet, et al.
Published: (2025)
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
by: Sachdeva, Rachneet, et al.
Published: (2023)
by: Sachdeva, Rachneet, et al.
Published: (2023)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024)
by: Song, Yixiao, et al.
Published: (2024)
Are Emergent Abilities in Large Language Models just In-Context Learning?
by: Lu, Sheng, et al.
Published: (2023)
by: Lu, Sheng, et al.
Published: (2023)
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs
by: Fang, Haishuo, et al.
Published: (2024)
by: Fang, Haishuo, et al.
Published: (2024)
Suri: Multi-constraint Instruction Following for Long-form Text Generation
by: Pham, Chau Minh, et al.
Published: (2024)
by: Pham, Chau Minh, et al.
Published: (2024)
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
by: Baumgärtner, Tim, et al.
Published: (2025)
by: Baumgärtner, Tim, et al.
Published: (2025)
Literary Evidence Retrieval via Long-Context Language Models
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
Citation Failure: Definition, Analysis and Efficient Mitigation
by: Buchmann, Jan, et al.
Published: (2025)
by: Buchmann, Jan, et al.
Published: (2025)
M2QA: Multi-domain Multilingual Question Answering
by: Engländer, Leon, et al.
Published: (2024)
by: Engländer, Leon, et al.
Published: (2024)
NeoQA: Evidence-based Question Answering with Generated News Events
by: Glockner, Max, et al.
Published: (2025)
by: Glockner, Max, et al.
Published: (2025)
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025)
by: Mekala, Anmol, et al.
Published: (2025)
Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation
by: Jafari, Nazanin, et al.
Published: (2026)
by: Jafari, Nazanin, et al.
Published: (2026)
Token Weighting for Long-Range Language Modeling
by: Helm, Falko, et al.
Published: (2025)
by: Helm, Falko, et al.
Published: (2025)
Attribute or Abstain: Large Language Models as Long Document Assistants
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
Frankentext: Stitching random text fragments into long-form narratives
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
VeriFastScore: Speeding up long-form factuality evaluation
by: Rajendhran, Rishanth, et al.
Published: (2025)
by: Rajendhran, Rishanth, et al.
Published: (2025)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
by: Dycke, Nils, et al.
Published: (2025)
by: Dycke, Nils, et al.
Published: (2025)
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification
by: Bates, Luke, et al.
Published: (2023)
by: Bates, Luke, et al.
Published: (2023)
Enhancing Depression Detection via Question-wise Modality Fusion
by: Mandal, Aishik, et al.
Published: (2025)
by: Mandal, Aishik, et al.
Published: (2025)
Towards Automated Error Discovery: A Study in Conversational AI
by: Petrak, Dominic, et al.
Published: (2025)
by: Petrak, Dominic, et al.
Published: (2025)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
by: Samuel, Vinay, et al.
Published: (2026)
by: Samuel, Vinay, et al.
Published: (2026)
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026)
by: Kim, Yekyung, et al.
Published: (2026)
BEARCUBS: A benchmark for computer-using web agents
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
by: Baumgärtner, Tim, et al.
Published: (2026)
by: Baumgärtner, Tim, et al.
Published: (2026)
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
by: Kaur, Rachneet, et al.
Published: (2025)
by: Kaur, Rachneet, et al.
Published: (2025)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
OLAPH: Improving Factuality in Biomedical Long-form Question Answering
by: Jeong, Minbyul, et al.
Published: (2024)
by: Jeong, Minbyul, et al.
Published: (2024)
Improving Attributed Long-form Question Answering with Intent Awareness
by: Zhao, Xinran, et al.
Published: (2026)
by: Zhao, Xinran, et al.
Published: (2026)
Agentic LLMs for Question Answering over Tabular Data
by: Tyagi, Rishit, et al.
Published: (2025)
by: Tyagi, Rishit, et al.
Published: (2025)
Commitment Checklist: Auditing Author Commitments in Peer Review
by: Chen, Chung-Chi, et al.
Published: (2026)
by: Chen, Chung-Chi, et al.
Published: (2026)
FoRAG: Factuality-optimized Retrieval Augmented Generation for Web-enhanced Long-form Question Answering
by: Cai, Tianchi, et al.
Published: (2024)
by: Cai, Tianchi, et al.
Published: (2024)
Document Structure in Long Document Transformers
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Overview of PerpectiveArg2024: The First Shared Task on Perspective Argument Retrieval
by: Falk, Neele, et al.
Published: (2024)
by: Falk, Neele, et al.
Published: (2024)
LLM Roleplay: Simulating Human-Chatbot Interaction
by: Tamoyan, Hovhannes, et al.
Published: (2024)
by: Tamoyan, Hovhannes, et al.
Published: (2024)
Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
Similar Items
-
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
by: Sachdeva, Rachneet, et al.
Published: (2025) -
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
by: Sachdeva, Rachneet, et al.
Published: (2023) -
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024) -
Are Emergent Abilities in Large Language Models just In-Context Learning?
by: Lu, Sheng, et al.
Published: (2023) -
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs
by: Fang, Haishuo, et al.
Published: (2024)