Provably Learning from Language Feedback
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Wanqiao, Nie, Allen, Zheng, Ruijie, Modi, Aditya, Swaminathan, Adith, Cheng, Ching-An |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
par: Cheng, Ching-An, et autres
Publié: (2024)
par: Cheng, Ching-An, et autres
Publié: (2024)
The Importance of Directional Feedback for LLM-based Optimizers
par: Nie, Allen, et autres
Publié: (2024)
par: Nie, Allen, et autres
Publié: (2024)
How to Solve Contextual Goal-Oriented Problems with Offline Datasets?
par: Fan, Ying, et autres
Publié: (2024)
par: Fan, Ying, et autres
Publié: (2024)
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
par: Chang, Trenton, et autres
Publié: (2025)
par: Chang, Trenton, et autres
Publié: (2025)
On Overcoming Miscalibrated Conversational Priors in LLM-based Chatbots
par: Herlihy, Christine, et autres
Publié: (2024)
par: Herlihy, Christine, et autres
Publié: (2024)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
par: Chowdhury, Sayak Ray, et autres
Publié: (2024)
par: Chowdhury, Sayak Ray, et autres
Publié: (2024)
Provable Interactive Learning with Hindsight Instruction Feedback
par: Misra, Dipendra, et autres
Publié: (2024)
par: Misra, Dipendra, et autres
Publié: (2024)
Reasoning Elicitation in Language Models via Counterfactual Feedback
par: Hüyük, Alihan, et autres
Publié: (2024)
par: Hüyük, Alihan, et autres
Publié: (2024)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
par: Schnabel, Tobias, et autres
Publié: (2025)
par: Schnabel, Tobias, et autres
Publié: (2025)
Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs
par: Tomlinson, Kiran, et autres
Publié: (2026)
par: Tomlinson, Kiran, et autres
Publié: (2026)
RLHF and IIA: Perverse Incentives
par: Xu, Wanqiao, et autres
Publié: (2023)
par: Xu, Wanqiao, et autres
Publié: (2023)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
par: Zheng, Qinqing, et autres
Publié: (2024)
par: Zheng, Qinqing, et autres
Publié: (2024)
CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback
par: Zhang, Wenbo, et autres
Publié: (2024)
par: Zhang, Wenbo, et autres
Publié: (2024)
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
par: Zhao, Siyan, et autres
Publié: (2025)
par: Zhao, Siyan, et autres
Publié: (2025)
Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models
par: Shukla, Divyaksh, et autres
Publié: (2026)
par: Shukla, Divyaksh, et autres
Publié: (2026)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
par: Bansal, Hritik, et autres
Publié: (2023)
par: Bansal, Hritik, et autres
Publié: (2023)
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials
par: Mandal, Shreyasi, et autres
Publié: (2024)
par: Mandal, Shreyasi, et autres
Publié: (2024)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
par: Xu, Ruichen, et autres
Publié: (2025)
par: Xu, Ruichen, et autres
Publié: (2025)
Geometry of Decision Making in Language Models
par: Joshi, Abhinav, et autres
Publié: (2025)
par: Joshi, Abhinav, et autres
Publié: (2025)
Saten: Sparse Augmented Tensor Networks for Post-Training Compression of Large Language Models
par: Solgi, Ryan, et autres
Publié: (2025)
par: Solgi, Ryan, et autres
Publié: (2025)
Offline Learning and Forgetting for Reasoning with Large Language Models
par: Ni, Tianwei, et autres
Publié: (2025)
par: Ni, Tianwei, et autres
Publié: (2025)
On Provable Length and Compositional Generalization
par: Ahuja, Kartik, et autres
Publié: (2024)
par: Ahuja, Kartik, et autres
Publié: (2024)
Learning Personalized Agents from Human Feedback
par: Liang, Kaiqu, et autres
Publié: (2026)
par: Liang, Kaiqu, et autres
Publié: (2026)
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Reinforcement Learning from Denoising Feedback
par: He, Qi, et autres
Publié: (2026)
par: He, Qi, et autres
Publié: (2026)
Benchmarking Benchmark Leakage in Large Language Models
par: Xu, Ruijie, et autres
Publié: (2024)
par: Xu, Ruijie, et autres
Publié: (2024)
When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal
par: Phalod, Aditya Ajay
Publié: (2026)
par: Phalod, Aditya Ajay
Publié: (2026)
Reinforcement Learning without Human Feedback for Last Mile Fine-Tuning of Large Language Models
par: Solway, Alec
Publié: (2024)
par: Solway, Alec
Publié: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
par: Wu, Diyuan, et autres
Publié: (2025)
par: Wu, Diyuan, et autres
Publié: (2025)
Task-Aware Calibration: Provably Optimal Decoding in LLMs
par: Tomov, Tim, et autres
Publié: (2026)
par: Tomov, Tim, et autres
Publié: (2026)
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
par: Solgi, Ryan, et autres
Publié: (2025)
par: Solgi, Ryan, et autres
Publié: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
par: Chen, Yanxi, et autres
Publié: (2024)
par: Chen, Yanxi, et autres
Publié: (2024)
Probing the Decision Boundaries of In-context Learning in Large Language Models
par: Zhao, Siyan, et autres
Publié: (2024)
par: Zhao, Siyan, et autres
Publié: (2024)
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
par: Allen-Zhu, Zeyuan, et autres
Publié: (2023)
par: Allen-Zhu, Zeyuan, et autres
Publié: (2023)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
par: Luo, Renjie, et autres
Publié: (2025)
par: Luo, Renjie, et autres
Publié: (2025)
SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes
par: Yang, Yifan, et autres
Publié: (2025)
par: Yang, Yifan, et autres
Publié: (2025)
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions
par: Nair, Inderjeet, et autres
Publié: (2024)
par: Nair, Inderjeet, et autres
Publié: (2024)
IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts
par: Basak, Udvas, et autres
Publié: (2024)
par: Basak, Udvas, et autres
Publié: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
par: Jung, Jaehun, et autres
Publié: (2024)
par: Jung, Jaehun, et autres
Publié: (2024)
SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning
par: Hu, Chenzhi, et autres
Publié: (2026)
par: Hu, Chenzhi, et autres
Publié: (2026)
Documents similaires
-
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
par: Cheng, Ching-An, et autres
Publié: (2024) -
The Importance of Directional Feedback for LLM-based Optimizers
par: Nie, Allen, et autres
Publié: (2024) -
How to Solve Contextual Goal-Oriented Problems with Offline Datasets?
par: Fan, Ying, et autres
Publié: (2024) -
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
par: Chang, Trenton, et autres
Publié: (2025) -
On Overcoming Miscalibrated Conversational Priors in LLM-based Chatbots
par: Herlihy, Christine, et autres
Publié: (2024)