The Art of Saying No: Contextual Noncompliance in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Brahman, Faeze, Kumar, Sachin, Balachandran, Vidhisha, Dasigi, Pradeep, Pyatkin, Valentina, Ravichander, Abhilasha, Wiegreffe, Sarah, Dziri, Nouha, Chandu, Khyathi, Hessel, Jack, Tsvetkov, Yulia, Smith, Noah A., Choi, Yejin, Hajishirzi, Hannaneh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023)
von: Yin, Da, et al.
Veröffentlicht: (2023)
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
von: Miranda, Lester James V., et al.
Veröffentlicht: (2024)
von: Miranda, Lester James V., et al.
Veröffentlicht: (2024)
RewardBench: Evaluating Reward Models for Language Modeling
von: Lambert, Nathan, et al.
Veröffentlicht: (2024)
von: Lambert, Nathan, et al.
Veröffentlicht: (2024)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
von: Graf, Victoria, et al.
Veröffentlicht: (2026)
von: Graf, Victoria, et al.
Veröffentlicht: (2026)
Small Reward Models via Backward Inference
von: Wang, Yike, et al.
Veröffentlicht: (2026)
von: Wang, Yike, et al.
Veröffentlicht: (2026)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
von: Rao, Kavel, et al.
Veröffentlicht: (2023)
von: Rao, Kavel, et al.
Veröffentlicht: (2023)
Fine-grained Hallucination Detection and Editing for Language Models
von: Mishra, Abhika, et al.
Veröffentlicht: (2024)
von: Mishra, Abhika, et al.
Veröffentlicht: (2024)
Reasoning Up the Instruction Ladder for Controllable Language Models
von: Zheng, Zishuo, et al.
Veröffentlicht: (2025)
von: Zheng, Zishuo, et al.
Veröffentlicht: (2025)
Meta-Reinforcement Learning with Self-Reflection for Agentic Search
von: Xiao, Teng, et al.
Veröffentlicht: (2026)
von: Xiao, Teng, et al.
Veröffentlicht: (2026)
ComPO: Community Preferences for Language Model Personalization
von: Kumar, Sachin, et al.
Veröffentlicht: (2024)
von: Kumar, Sachin, et al.
Veröffentlicht: (2024)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
von: Lyu, Xinxi, et al.
Veröffentlicht: (2024)
von: Lyu, Xinxi, et al.
Veröffentlicht: (2024)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
von: Chen, Tong, et al.
Veröffentlicht: (2025)
von: Chen, Tong, et al.
Veröffentlicht: (2025)
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
von: Wang, Yike, et al.
Veröffentlicht: (2025)
von: Wang, Yike, et al.
Veröffentlicht: (2025)
Generalizing Verifiable Instruction Following
von: Pyatkin, Valentina, et al.
Veröffentlicht: (2025)
von: Pyatkin, Valentina, et al.
Veröffentlicht: (2025)
Large-Scale Data Selection for Instruction Tuning
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
Merge to Learn: Efficiently Adding Skills to Language Models with Model Merging
von: Morrison, Jacob, et al.
Veröffentlicht: (2024)
von: Morrison, Jacob, et al.
Veröffentlicht: (2024)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning
von: Wang, Haojin, et al.
Veröffentlicht: (2026)
von: Wang, Haojin, et al.
Veröffentlicht: (2026)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
von: Baheti, Ashutosh, et al.
Veröffentlicht: (2024)
von: Baheti, Ashutosh, et al.
Veröffentlicht: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
P^3SUM: Preserving Author's Perspective in News Summarization with Diffusion Language Models
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
von: Liu, Yuhan, et al.
Veröffentlicht: (2023)
MacGyver: Are Large Language Models Creative Problem Solvers?
von: Tian, Yufei, et al.
Veröffentlicht: (2023)
von: Tian, Yufei, et al.
Veröffentlicht: (2023)
Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
von: Ahuja, Kabir, et al.
Veröffentlicht: (2024)
von: Ahuja, Kabir, et al.
Veröffentlicht: (2024)
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
von: Lambert, Nathan, et al.
Veröffentlicht: (2024)
von: Lambert, Nathan, et al.
Veröffentlicht: (2024)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language Models
von: Feng, Shangbin, et al.
Veröffentlicht: (2023)
von: Feng, Shangbin, et al.
Veröffentlicht: (2023)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
FACTS&EVIDENCE: An Interactive Tool for Transparent Fine-Grained Factual Verification of Machine-Generated Text
von: Boonsanong, Varich, et al.
Veröffentlicht: (2025)
von: Boonsanong, Varich, et al.
Veröffentlicht: (2025)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
von: Jiang, Liwei, et al.
Veröffentlicht: (2025)
von: Jiang, Liwei, et al.
Veröffentlicht: (2025)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024) -
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024) -
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023) -
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
von: Miranda, Lester James V., et al.
Veröffentlicht: (2024) -
RewardBench: Evaluating Reward Models for Language Modeling
von: Lambert, Nathan, et al.
Veröffentlicht: (2024)