Saved in:
| Main Authors: | Liu, Grace, Christian, Brian, Dumbalska, Tsvetomira, Bakker, Michiel A., Dubey, Rachit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.04721 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reward Model Interpretability via Optimal and Pessimal Tokens
by: Christian, Brian, et al.
Published: (2025)
by: Christian, Brian, et al.
Published: (2025)
Reward Models Inherit Value Biases from Pretraining
by: Christian, Brian, et al.
Published: (2026)
by: Christian, Brian, et al.
Published: (2026)
What do we mean when we talk about socioeconomic status? Implications for measurement, mechanisms and interventions from a critical review on adolescent mental health
by: Mirela Zaneva, et al.
Published: (2024)
by: Mirela Zaneva, et al.
Published: (2024)
Adding New Capability in Existing Scientific Application with LLM Assistance
by: Dubey, Anshu, et al.
Published: (2025)
by: Dubey, Anshu, et al.
Published: (2025)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)
by: Vaccaro, Michelle, et al.
Published: (2026)
From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM
by: Fulay, Suyash, et al.
Published: (2025)
by: Fulay, Suyash, et al.
Published: (2025)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
by: Du, Yufeng, et al.
Published: (2025)
by: Du, Yufeng, et al.
Published: (2025)
AI and Collective Decisions: Strengthening Legitimacy and Losers' Consent
by: Fulay, Suyash, et al.
Published: (2026)
by: Fulay, Suyash, et al.
Published: (2026)
Digital Epidemiology: Leveraging Social Media for Insight into Epilepsy and Mental Health
by: Dahiya, Liza, et al.
Published: (2024)
by: Dahiya, Liza, et al.
Published: (2024)
Benchmarking Overton Pluralism in LLMs
by: Poole-Dayan, Elinor, et al.
Published: (2025)
by: Poole-Dayan, Elinor, et al.
Published: (2025)
Belief Engine: Configurable and Inspectable Stance Dynamics in Multi-Agent LLM Deliberation
by: Yang, Joshua C., et al.
Published: (2026)
by: Yang, Joshua C., et al.
Published: (2026)
Agora: Teaching the Skill of Consensus-Finding with AI Personas Grounded in Human Voice
by: Ravi, Prerna, et al.
Published: (2026)
by: Ravi, Prerna, et al.
Published: (2026)
AI for Service: Proactive Assistance with AI Glasses
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024)
by: Bayat, Reza, et al.
Published: (2024)
Tell Me Why: Incentivizing Explanations
by: Srinivasan, Siddarth, et al.
Published: (2025)
by: Srinivasan, Siddarth, et al.
Published: (2025)
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
by: Liu, Ruikang, et al.
Published: (2025)
by: Liu, Ruikang, et al.
Published: (2025)
AI Readiness in Healthcare through Storytelling XAI
by: Dubey, Akshat, et al.
Published: (2024)
by: Dubey, Akshat, et al.
Published: (2024)
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
by: Cao, Xiaoyang, et al.
Published: (2025)
by: Cao, Xiaoyang, et al.
Published: (2025)
Data Shifts Hurt CoT: A Theoretical Study
by: Yin, Lang, et al.
Published: (2025)
by: Yin, Lang, et al.
Published: (2025)
Detecting AI Assistance in Abstract Complex Tasks
by: King, Tyler, et al.
Published: (2025)
by: King, Tyler, et al.
Published: (2025)
Independence Is Not an Issue in Neurosymbolic AI
by: Faronius, Håkan Karlsson, et al.
Published: (2025)
by: Faronius, Håkan Karlsson, et al.
Published: (2025)
Zero-shot counting with a dual-stream neural network model
by: Thompson, Jessica A. F., et al.
Published: (2024)
by: Thompson, Jessica A. F., et al.
Published: (2024)
Can AI mediation improve democratic deliberation?
by: Tessler, Michael Henry, et al.
Published: (2026)
by: Tessler, Michael Henry, et al.
Published: (2026)
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
by: Zhou, Shu, et al.
Published: (2026)
by: Zhou, Shu, et al.
Published: (2026)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
by: Armandpour, Mohammadreza, et al.
Published: (2026)
by: Armandpour, Mohammadreza, et al.
Published: (2026)
PHAX: A Structured Argumentation Framework for User-Centered Explainable AI in Public Health and Biomedical Sciences
by: İlgen, Bahar, et al.
Published: (2025)
by: İlgen, Bahar, et al.
Published: (2025)
AssistanceZero: Scalably Solving Assistance Games
by: Laidlaw, Cassidy, et al.
Published: (2025)
by: Laidlaw, Cassidy, et al.
Published: (2025)
Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
by: Hong, Shen Zhou, et al.
Published: (2026)
by: Hong, Shen Zhou, et al.
Published: (2026)
Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs
by: Dubey, Shivam
Published: (2025)
by: Dubey, Shivam
Published: (2025)
When Verification Hurts: Asymmetric Effects of Multi-Agent Feedback in Logic Proof Tutoring
by: Yasir, Tahreem, et al.
Published: (2026)
by: Yasir, Tahreem, et al.
Published: (2026)
Invoice Information Extraction: Methods and Performance Evaluation
by: Yashwant, Sai, et al.
Published: (2025)
by: Yashwant, Sai, et al.
Published: (2025)
Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem Solving
by: Zenkri, Oussama, et al.
Published: (2026)
by: Zenkri, Oussama, et al.
Published: (2026)
When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning
by: Qiu, Chenghao, et al.
Published: (2026)
by: Qiu, Chenghao, et al.
Published: (2026)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
by: Zhang, Yi-Ge, et al.
Published: (2025)
by: Zhang, Yi-Ge, et al.
Published: (2025)
AI-Enhanced Operator Assistance for UNICOS Applications
by: Tam, Bernard, et al.
Published: (2025)
by: Tam, Bernard, et al.
Published: (2025)
ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
by: Sun, Liwen, et al.
Published: (2024)
by: Sun, Liwen, et al.
Published: (2024)
User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance
by: Verghese, Mrinal, et al.
Published: (2024)
by: Verghese, Mrinal, et al.
Published: (2024)
AstraAI: LLMs, Retrieval, and AST-Guided Assistance for HPC Codebases
by: Natarajan, Mahesh, et al.
Published: (2026)
by: Natarajan, Mahesh, et al.
Published: (2026)
Value of Assistance for Grasping
by: Masarwy, Mohammad, et al.
Published: (2023)
by: Masarwy, Mohammad, et al.
Published: (2023)
Causal Predictive Optimization and Generation for Business AI
by: Zhao, Liyang, et al.
Published: (2025)
by: Zhao, Liyang, et al.
Published: (2025)
Similar Items
-
Reward Model Interpretability via Optimal and Pessimal Tokens
by: Christian, Brian, et al.
Published: (2025) -
Reward Models Inherit Value Biases from Pretraining
by: Christian, Brian, et al.
Published: (2026) -
What do we mean when we talk about socioeconomic status? Implications for measurement, mechanisms and interventions from a critical review on adolescent mental health
by: Mirela Zaneva, et al.
Published: (2024) -
Adding New Capability in Existing Scientific Application with LLM Assistance
by: Dubey, Anshu, et al.
Published: (2025) -
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)