Clarify: Improving Model Robustness With Natural Language Corrections
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Yoonho, Lam, Michelle S., Vasconcelos, Helena, Bernstein, Michael S., Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Calibrating Language Models with Adaptive Temperature Scaling
von: Xie, Johnathan, et al.
Veröffentlicht: (2024)
von: Xie, Johnathan, et al.
Veröffentlicht: (2024)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
von: Xie, Johnathan, et al.
Veröffentlicht: (2024)
von: Xie, Johnathan, et al.
Veröffentlicht: (2024)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
von: Rafailov, Rafael, et al.
Veröffentlicht: (2023)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2023)
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
von: Liu, Yuejiang, et al.
Veröffentlicht: (2024)
von: Liu, Yuejiang, et al.
Veröffentlicht: (2024)
Conservative Prediction via Data-Driven Confidence Minimization
von: Choi, Caroline, et al.
Veröffentlicht: (2023)
von: Choi, Caroline, et al.
Veröffentlicht: (2023)
Resolving Intent Ambiguities by Retrieving Discriminative Clarifying Questions
von: Dhole, Kaustubh D.
Veröffentlicht: (2020)
von: Dhole, Kaustubh D.
Veröffentlicht: (2020)
Harnessing Large Language Models: Fine-tuned BERT for Detecting Charismatic Leadership Tactics in Natural Language
von: Saeid, Yasser, et al.
Veröffentlicht: (2024)
von: Saeid, Yasser, et al.
Veröffentlicht: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
von: Baan, Joris, et al.
Veröffentlicht: (2026)
von: Baan, Joris, et al.
Veröffentlicht: (2026)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
von: Galeone, Cosimo, et al.
Veröffentlicht: (2026)
On the Robustness of Reward Models for Language Model Alignment
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
Planning In Natural Language Improves LLM Search For Code Generation
von: Wang, Evan, et al.
Veröffentlicht: (2024)
von: Wang, Evan, et al.
Veröffentlicht: (2024)
RLVF: Learning from Verbal Feedback without Overgeneralization
von: Stephan, Moritz, et al.
Veröffentlicht: (2024)
von: Stephan, Moritz, et al.
Veröffentlicht: (2024)
Protein Language Models Diverge from Natural Language: Comparative Analysis and Improved Inference
von: Hart, Anna, et al.
Veröffentlicht: (2026)
von: Hart, Anna, et al.
Veröffentlicht: (2026)
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries
von: Yang, Blair, et al.
Veröffentlicht: (2024)
von: Yang, Blair, et al.
Veröffentlicht: (2024)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
Improving Large Language Models with Concept-Aware Fine-Tuning
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
Structured Prompts Improve Evaluation of Language Models
von: Aali, Asad, et al.
Veröffentlicht: (2025)
von: Aali, Asad, et al.
Veröffentlicht: (2025)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
von: Li, Zhuoyun, et al.
Veröffentlicht: (2026)
von: Li, Zhuoyun, et al.
Veröffentlicht: (2026)
Fewer Truncations Improve Language Modeling
von: Ding, Hantian, et al.
Veröffentlicht: (2024)
von: Ding, Hantian, et al.
Veröffentlicht: (2024)
Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation
von: Lee, JoonHo, et al.
Veröffentlicht: (2024)
von: Lee, JoonHo, et al.
Veröffentlicht: (2024)
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
von: Tsui, Ken
Veröffentlicht: (2025)
von: Tsui, Ken
Veröffentlicht: (2025)
Improving Context-Aware Preference Modeling for Language Models
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
Correcting Large Language Model Behavior via Influence Function
von: Zhang, Han, et al.
Veröffentlicht: (2024)
von: Zhang, Han, et al.
Veröffentlicht: (2024)
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
LaFFi: Leveraging Hybrid Natural Language Feedback for Fine-tuning Language Models
von: Li, Qianxi, et al.
Veröffentlicht: (2023)
von: Li, Qianxi, et al.
Veröffentlicht: (2023)
Enhancing Antibiotic Stewardship using a Natural Language Approach for Better Feature Representation
von: Lee, Simon A., et al.
Veröffentlicht: (2024)
von: Lee, Simon A., et al.
Veröffentlicht: (2024)
ProgCo: Program Helps Self-Correction of Large Language Models
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
Uncovering Customer Issues through Topological Natural Language Analysis
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2024)
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2024)
Assessing Large Language Models on Climate Information
von: Bulian, Jannis, et al.
Veröffentlicht: (2023)
von: Bulian, Jannis, et al.
Veröffentlicht: (2023)
Improving Code Generation by Training with Natural Language Feedback
von: Chen, Angelica, et al.
Veröffentlicht: (2023)
von: Chen, Angelica, et al.
Veröffentlicht: (2023)
ROPO: Robust Preference Optimization for Large Language Models
von: Liang, Xize, et al.
Veröffentlicht: (2024)
von: Liang, Xize, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Calibrating Language Models with Adaptive Temperature Scaling
von: Xie, Johnathan, et al.
Veröffentlicht: (2024) -
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025) -
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
von: Xie, Johnathan, et al.
Veröffentlicht: (2024) -
A Critical Evaluation of AI Feedback for Aligning Large Language Models
von: Sharma, Archit, et al.
Veröffentlicht: (2024) -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
von: Rafailov, Rafael, et al.
Veröffentlicht: (2023)