Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Dongwei, Zhang, Alvin, Wang, Andrew, Andrews, Nicholas, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
by: Uzunoglu, Arda, et al.
Published: (2026)
by: Uzunoglu, Arda, et al.
Published: (2026)
Benchmarking Language Model Creativity: A Case Study on Code Generation
by: Lu, Yining, et al.
Published: (2024)
by: Lu, Yining, et al.
Published: (2024)
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
by: Shi, Taiwei, et al.
Published: (2024)
by: Shi, Taiwei, et al.
Published: (2024)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
Test-time Recursive Thinking: Self-Improvement without External Feedback
by: Zhuang, Yufan, et al.
Published: (2026)
by: Zhuang, Yufan, et al.
Published: (2026)
MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
by: Byerly, Adam, et al.
Published: (2025)
by: Byerly, Adam, et al.
Published: (2025)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
by: Li, Tianjian, et al.
Published: (2025)
by: Li, Tianjian, et al.
Published: (2025)
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
by: Byerly, Adam, et al.
Published: (2024)
by: Byerly, Adam, et al.
Published: (2024)
Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
by: Yasir, Tahreem, et al.
Published: (2026)
by: Yasir, Tahreem, et al.
Published: (2026)
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
by: Rufail, Andrew, et al.
Published: (2025)
by: Rufail, Andrew, et al.
Published: (2025)
LLMs are Superior Feedback Providers: Bootstrapping Reasoning for Lie Detection with Self-Generated Feedback
by: Banerjee, Tanushree, et al.
Published: (2024)
by: Banerjee, Tanushree, et al.
Published: (2024)
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
by: Ormerod, Christopher
Published: (2025)
by: Ormerod, Christopher
Published: (2025)
Teaching LLMs to Abstain across Languages via Multilingual Feedback
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
by: Ye, Xiao, et al.
Published: (2024)
by: Ye, Xiao, et al.
Published: (2024)
Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins
by: Chan, Amanda, et al.
Published: (2025)
by: Chan, Amanda, et al.
Published: (2025)
LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
LLMs Struggle with Abstract Meaning Comprehension More Than Expected
by: Alhazmi, Hamoud, et al.
Published: (2026)
by: Alhazmi, Hamoud, et al.
Published: (2026)
Time-Reversal Provides Unsupervised Feedback to LLMs
by: Varun, Yerram, et al.
Published: (2024)
by: Varun, Yerram, et al.
Published: (2024)
Self-Refinement of Language Models from External Proxy Metrics Feedback
by: Ramji, Keshav, et al.
Published: (2024)
by: Ramji, Keshav, et al.
Published: (2024)
AudienceView: AI-Assisted Interpretation of Audience Feedback in Journalism
by: Brannon, William, et al.
Published: (2024)
by: Brannon, William, et al.
Published: (2024)
A Systematic Study of Pseudo-Relevance Feedback with LLMs
by: Jedidi, Nour, et al.
Published: (2026)
by: Jedidi, Nour, et al.
Published: (2026)
Evaluating the Evaluators: Are readability metrics good measures of readability?
by: Cachola, Isabel, et al.
Published: (2025)
by: Cachola, Isabel, et al.
Published: (2025)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
by: Wang, Hexuan, et al.
Published: (2026)
by: Wang, Hexuan, et al.
Published: (2026)
Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework
by: Ma, Xilai, et al.
Published: (2026)
by: Ma, Xilai, et al.
Published: (2026)
Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation
by: Xiao, Meiman, et al.
Published: (2026)
by: Xiao, Meiman, et al.
Published: (2026)
Code Aesthetics with Agentic Reward Feedback
by: Xiao, Bang, et al.
Published: (2025)
by: Xiao, Bang, et al.
Published: (2025)
Generating Planning Feedback for Open-Ended Programming Exercises with LLMs
by: Demirtaş, Mehmet Arif, et al.
Published: (2025)
by: Demirtaş, Mehmet Arif, et al.
Published: (2025)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
by: Gehring, Jonas, et al.
Published: (2024)
by: Gehring, Jonas, et al.
Published: (2024)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
by: Moon, Seungjun, et al.
Published: (2023)
by: Moon, Seungjun, et al.
Published: (2023)
Long-context LLMs Struggle with Long In-context Learning
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
Improving the Validity of Automatically Generated Feedback via Reinforcement Learning
by: Scarlatos, Alexander, et al.
Published: (2024)
by: Scarlatos, Alexander, et al.
Published: (2024)
Generating Feedback-Ladders for Logical Errors in Programming using Large Language Models
by: Heickal, Hasnain, et al.
Published: (2024)
by: Heickal, Hasnain, et al.
Published: (2024)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
by: Wang, Xingyao, et al.
Published: (2023)
by: Wang, Xingyao, et al.
Published: (2023)
RLSF: Fine-tuning LLMs via Symbolic Feedback
by: Jha, Piyush, et al.
Published: (2024)
by: Jha, Piyush, et al.
Published: (2024)
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Similar Items
-
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
by: Jiang, Dongwei, et al.
Published: (2024) -
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025) -
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
by: Uzunoglu, Arda, et al.
Published: (2026) -
Benchmarking Language Model Creativity: A Case Study on Code Generation
by: Lu, Yining, et al.
Published: (2024) -
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
by: Shi, Taiwei, et al.
Published: (2024)