Beyond Output Critique: Self-Correction via Task Distillation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Rahmani, Hossein A., Wan, Mengting, Zhou, Pei, Yang, Longqi, Craswell, Nick, Yilmaz, Emine, Jauhar, Sujay Kumar |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Self-Correcting Large Language Models: Generation vs. Multiple Choice
par: Rahmani, Hossein A., et autres
Publié: (2025)
par: Rahmani, Hossein A., et autres
Publié: (2025)
Towards Understanding Bias in Synthetic Data for Evaluation
par: Rahmani, Hossein A., et autres
Publié: (2025)
par: Rahmani, Hossein A., et autres
Publié: (2025)
Synthetic Test Collections for Retrieval Evaluation
par: Rahmani, Hossein A., et autres
Publié: (2024)
par: Rahmani, Hossein A., et autres
Publié: (2024)
Overview of the TREC 2023 deep learning track
par: Craswell, Nick, et autres
Publié: (2025)
par: Craswell, Nick, et autres
Publié: (2025)
Overview of the TREC 2021 deep learning track
par: Craswell, Nick, et autres
Publié: (2025)
par: Craswell, Nick, et autres
Publié: (2025)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
par: Chen, Yinzhu, et autres
Publié: (2026)
par: Chen, Yinzhu, et autres
Publié: (2026)
Group Preference Alignment: Customized LLM Response Generation from In-Situ Conversations
par: Mondal, Ishani, et autres
Publié: (2025)
par: Mondal, Ishani, et autres
Publié: (2025)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
par: Rahmani, Hossein A., et autres
Publié: (2024)
par: Rahmani, Hossein A., et autres
Publié: (2024)
Teaching Language Models To Gather Information Proactively
par: Huang, Tenghao, et autres
Publié: (2025)
par: Huang, Tenghao, et autres
Publié: (2025)
Understanding the Role of User Profile in the Personalization of Large Language Models
par: Wu, Bin, et autres
Publié: (2024)
par: Wu, Bin, et autres
Publié: (2024)
Overview of the TREC 2022 deep learning track
par: Craswell, Nick, et autres
Publié: (2025)
par: Craswell, Nick, et autres
Publié: (2025)
TnT-LLM: Text Mining at Scale with Large Language Models
par: Wan, Mengting, et autres
Publié: (2024)
par: Wan, Mengting, et autres
Publié: (2024)
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
par: Baek, Jinheon, et autres
Publié: (2024)
par: Baek, Jinheon, et autres
Publié: (2024)
Conversational User-AI Intervention: A Study on Prompt Rewriting for Improved LLM Response Generation
par: Sarkar, Rupak, et autres
Publié: (2025)
par: Sarkar, Rupak, et autres
Publié: (2025)
Beyond Internal Data: Constructing Complete Datasets for Fairness Testing
par: Ramineni, Varsha, et autres
Publié: (2025)
par: Ramineni, Varsha, et autres
Publié: (2025)
Knowledge-Augmented Large Language Models for Personalized Contextual Query Suggestion
par: Baek, Jinheon, et autres
Publié: (2023)
par: Baek, Jinheon, et autres
Publié: (2023)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
par: Fu, Xiao, et autres
Publié: (2025)
par: Fu, Xiao, et autres
Publié: (2025)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
par: Kamoi, Ryo, et autres
Publié: (2026)
par: Kamoi, Ryo, et autres
Publié: (2026)
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
par: Rahmani, Hossein A., et autres
Publié: (2024)
par: Rahmani, Hossein A., et autres
Publié: (2024)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
par: Gou, Zhibin, et autres
Publié: (2023)
par: Gou, Zhibin, et autres
Publié: (2023)
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
par: Yao, Zonghai, et autres
Publié: (2024)
par: Yao, Zonghai, et autres
Publié: (2024)
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
par: Shi, Taiwei, et autres
Publié: (2024)
par: Shi, Taiwei, et autres
Publié: (2024)
Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
par: Rahmani, Zahra, et autres
Publié: (2025)
par: Rahmani, Zahra, et autres
Publié: (2025)
Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
par: Finkelshtein, Ben, et autres
Publié: (2025)
par: Finkelshtein, Ben, et autres
Publié: (2025)
Large language models can accurately predict searcher preferences
par: Thomas, Paul, et autres
Publié: (2023)
par: Thomas, Paul, et autres
Publié: (2023)
MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation
par: Rahmani, Mahdi, et autres
Publié: (2025)
par: Rahmani, Mahdi, et autres
Publié: (2025)
Neon: News Entity-Interaction Extraction for Enhanced Question Answering
par: Singhania, Sneha, et autres
Publié: (2024)
par: Singhania, Sneha, et autres
Publié: (2024)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
par: Wen, Xueru, et autres
Publié: (2025)
par: Wen, Xueru, et autres
Publié: (2025)
The Critique of Critique
par: Sun, Shichao, et autres
Publié: (2024)
par: Sun, Shichao, et autres
Publié: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
par: Lin, Zicheng, et autres
Publié: (2024)
par: Lin, Zicheng, et autres
Publié: (2024)
Idea2Plan: Exploring AI-Powered Research Planning
par: Huang, Jin, et autres
Publié: (2025)
par: Huang, Jin, et autres
Publié: (2025)
Self-Correction Distillation for Structured Data Question Answering
par: Zhu, Yushan, et autres
Publié: (2025)
par: Zhu, Yushan, et autres
Publié: (2025)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
par: Wang, Yiming, et autres
Publié: (2024)
par: Wang, Yiming, et autres
Publié: (2024)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
par: Thakur, Nandan, et autres
Publié: (2025)
par: Thakur, Nandan, et autres
Publié: (2025)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
par: Zhang, Yifei, et autres
Publié: (2026)
par: Zhang, Yifei, et autres
Publié: (2026)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
par: Khan, Omer Jauhar
Publié: (2025)
par: Khan, Omer Jauhar
Publié: (2025)
Merging Improves Self-Critique Against Jailbreak Attacks
par: Gallego, Victor
Publié: (2024)
par: Gallego, Victor
Publié: (2024)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
par: Ke, Pei, et autres
Publié: (2023)
par: Ke, Pei, et autres
Publié: (2023)
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints
par: Ferraz, Thomas Palmeira, et autres
Publié: (2024)
par: Ferraz, Thomas Palmeira, et autres
Publié: (2024)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
par: Wu, Xueqing, et autres
Publié: (2024)
par: Wu, Xueqing, et autres
Publié: (2024)
Documents similaires
-
Self-Correcting Large Language Models: Generation vs. Multiple Choice
par: Rahmani, Hossein A., et autres
Publié: (2025) -
Towards Understanding Bias in Synthetic Data for Evaluation
par: Rahmani, Hossein A., et autres
Publié: (2025) -
Synthetic Test Collections for Retrieval Evaluation
par: Rahmani, Hossein A., et autres
Publié: (2024) -
Overview of the TREC 2023 deep learning track
par: Craswell, Nick, et autres
Publié: (2025) -
Overview of the TREC 2021 deep learning track
par: Craswell, Nick, et autres
Publié: (2025)