Beyond Output Critique: Self-Correction via Task Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Rahmani, Hossein A., Wan, Mengting, Zhou, Pei, Yang, Longqi, Craswell, Nick, Yilmaz, Emine, Jauhar, Sujay Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Correcting Large Language Models: Generation vs. Multiple Choice
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
Towards Understanding Bias in Synthetic Data for Evaluation
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
Synthetic Test Collections for Retrieval Evaluation
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Overview of the TREC 2023 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Overview of the TREC 2021 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
by: Chen, Yinzhu, et al.
Published: (2026)
by: Chen, Yinzhu, et al.
Published: (2026)
Group Preference Alignment: Customized LLM Response Generation from In-Situ Conversations
by: Mondal, Ishani, et al.
Published: (2025)
by: Mondal, Ishani, et al.
Published: (2025)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Teaching Language Models To Gather Information Proactively
by: Huang, Tenghao, et al.
Published: (2025)
by: Huang, Tenghao, et al.
Published: (2025)
Understanding the Role of User Profile in the Personalization of Large Language Models
by: Wu, Bin, et al.
Published: (2024)
by: Wu, Bin, et al.
Published: (2024)
Overview of the TREC 2022 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
TnT-LLM: Text Mining at Scale with Large Language Models
by: Wan, Mengting, et al.
Published: (2024)
by: Wan, Mengting, et al.
Published: (2024)
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
by: Baek, Jinheon, et al.
Published: (2024)
by: Baek, Jinheon, et al.
Published: (2024)
Conversational User-AI Intervention: A Study on Prompt Rewriting for Improved LLM Response Generation
by: Sarkar, Rupak, et al.
Published: (2025)
by: Sarkar, Rupak, et al.
Published: (2025)
Beyond Internal Data: Constructing Complete Datasets for Fairness Testing
by: Ramineni, Varsha, et al.
Published: (2025)
by: Ramineni, Varsha, et al.
Published: (2025)
Knowledge-Augmented Large Language Models for Personalized Contextual Query Suggestion
by: Baek, Jinheon, et al.
Published: (2023)
by: Baek, Jinheon, et al.
Published: (2023)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
by: Kamoi, Ryo, et al.
Published: (2026)
by: Kamoi, Ryo, et al.
Published: (2026)
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
by: Yao, Zonghai, et al.
Published: (2024)
by: Yao, Zonghai, et al.
Published: (2024)
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
by: Shi, Taiwei, et al.
Published: (2024)
by: Shi, Taiwei, et al.
Published: (2024)
Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
by: Rahmani, Zahra, et al.
Published: (2025)
by: Rahmani, Zahra, et al.
Published: (2025)
Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
by: Finkelshtein, Ben, et al.
Published: (2025)
by: Finkelshtein, Ben, et al.
Published: (2025)
Large language models can accurately predict searcher preferences
by: Thomas, Paul, et al.
Published: (2023)
by: Thomas, Paul, et al.
Published: (2023)
MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation
by: Rahmani, Mahdi, et al.
Published: (2025)
by: Rahmani, Mahdi, et al.
Published: (2025)
Neon: News Entity-Interaction Extraction for Enhanced Question Answering
by: Singhania, Sneha, et al.
Published: (2024)
by: Singhania, Sneha, et al.
Published: (2024)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
by: Wen, Xueru, et al.
Published: (2025)
by: Wen, Xueru, et al.
Published: (2025)
The Critique of Critique
by: Sun, Shichao, et al.
Published: (2024)
by: Sun, Shichao, et al.
Published: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
Idea2Plan: Exploring AI-Powered Research Planning
by: Huang, Jin, et al.
Published: (2025)
by: Huang, Jin, et al.
Published: (2025)
Self-Correction Distillation for Structured Data Question Answering
by: Zhu, Yushan, et al.
Published: (2025)
by: Zhu, Yushan, et al.
Published: (2025)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
by: Khan, Omer Jauhar
Published: (2025)
by: Khan, Omer Jauhar
Published: (2025)
Merging Improves Self-Critique Against Jailbreak Attacks
by: Gallego, Victor
Published: (2024)
by: Gallego, Victor
Published: (2024)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
by: Ke, Pei, et al.
Published: (2023)
by: Ke, Pei, et al.
Published: (2023)
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints
by: Ferraz, Thomas Palmeira, et al.
Published: (2024)
by: Ferraz, Thomas Palmeira, et al.
Published: (2024)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
by: Wu, Xueqing, et al.
Published: (2024)
by: Wu, Xueqing, et al.
Published: (2024)
Similar Items
-
Self-Correcting Large Language Models: Generation vs. Multiple Choice
by: Rahmani, Hossein A., et al.
Published: (2025) -
Towards Understanding Bias in Synthetic Data for Evaluation
by: Rahmani, Hossein A., et al.
Published: (2025) -
Synthetic Test Collections for Retrieval Evaluation
by: Rahmani, Hossein A., et al.
Published: (2024) -
Overview of the TREC 2023 deep learning track
by: Craswell, Nick, et al.
Published: (2025) -
Overview of the TREC 2021 deep learning track
by: Craswell, Nick, et al.
Published: (2025)