When Do LLM Preferences Predict Downstream Behavior?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Slama, Katarina, Souly, Alexandra, Bansal, Dishank, Davidson, Henry, Summerfield, Christopher, Luettgau, Lennart |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ask don't tell: Reducing sycophancy in large language models
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
von: Summerfield, Christopher, et al.
Veröffentlicht: (2025)
von: Summerfield, Christopher, et al.
Veröffentlicht: (2025)
One-shot emergency psychiatric triage across 15 frontier AI chatbots
von: Weilnhammer, Veith, et al.
Veröffentlicht: (2026)
von: Weilnhammer, Veith, et al.
Veröffentlicht: (2026)
TaskMet: Task-Driven Metric Learning for Model Learning
von: Bansal, Dishank, et al.
Veröffentlicht: (2023)
von: Bansal, Dishank, et al.
Veröffentlicht: (2023)
Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
von: Dohnány, Sebastian, et al.
Veröffentlicht: (2025)
von: Dohnány, Sebastian, et al.
Veröffentlicht: (2025)
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
Evaluating whether AI models would sabotage AI safety research
von: Kirk, Robert, et al.
Veröffentlicht: (2026)
von: Kirk, Robert, et al.
Veröffentlicht: (2026)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
UK AISI Alignment Evaluation Case-Study
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
von: Aggarwal, Dishank, et al.
Veröffentlicht: (2024)
von: Aggarwal, Dishank, et al.
Veröffentlicht: (2024)
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
von: Malik, Sameer, et al.
Veröffentlicht: (2025)
von: Malik, Sameer, et al.
Veröffentlicht: (2025)
Subjective Behaviors and Preferences in LLM: Language of Browsing
von: Sundaresan, Sai, et al.
Veröffentlicht: (2025)
von: Sundaresan, Sai, et al.
Veröffentlicht: (2025)
Conversational AI increases political knowledge as effectively as self-directed internet search
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
von: Szeider, Stefan
Veröffentlicht: (2025)
von: Szeider, Stefan
Veröffentlicht: (2025)
Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study
von: Bansal, Kaushal
Veröffentlicht: (2026)
von: Bansal, Kaushal
Veröffentlicht: (2026)
When Agents Disagree With Themselves: Measuring Behavioral Consistency in LLM-Based Agents
von: Mehta, Aman
Veröffentlicht: (2026)
von: Mehta, Aman
Veröffentlicht: (2026)
Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores
von: Panda, Shevya, et al.
Veröffentlicht: (2026)
von: Panda, Shevya, et al.
Veröffentlicht: (2026)
Scaling Laws for Predicting Downstream Performance in LLMs
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
Seven simple steps for log analysis in AI systems
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
Large Language Models and Algorithm Execution: Application to an Arithmetic Function
von: Slama, Farah Ben, et al.
Veröffentlicht: (2026)
von: Slama, Farah Ben, et al.
Veröffentlicht: (2026)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
von: Ghate, Kshitish, et al.
Veröffentlicht: (2025)
von: Ghate, Kshitish, et al.
Veröffentlicht: (2025)
When Does Predictive Inverse Dynamics Outperform Behavior Cloning?
von: Schäfer, Lukas, et al.
Veröffentlicht: (2026)
von: Schäfer, Lukas, et al.
Veröffentlicht: (2026)
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
von: Chen, Haoxian, et al.
Veröffentlicht: (2024)
von: Chen, Haoxian, et al.
Veröffentlicht: (2024)
Decoupled Behavioral Cloning for Scalable Inductive Generalization in RL from Specifications
von: Subramanian, Vignesh, et al.
Veröffentlicht: (2026)
von: Subramanian, Vignesh, et al.
Veröffentlicht: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
von: Karnik, Sathwik, et al.
Veröffentlicht: (2025)
von: Karnik, Sathwik, et al.
Veröffentlicht: (2025)
Do LLM Agents Exhibit Social Behavior?
von: Leng, Yan, et al.
Veröffentlicht: (2023)
von: Leng, Yan, et al.
Veröffentlicht: (2023)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
von: Ensign, Danielle, et al.
Veröffentlicht: (2025)
von: Ensign, Danielle, et al.
Veröffentlicht: (2025)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
Do Large Language Models Mentalize When They Teach?
von: Harootonian, Sevan K., et al.
Veröffentlicht: (2026)
von: Harootonian, Sevan K., et al.
Veröffentlicht: (2026)
Preference Learning Algorithms Do Not Learn Preference Rankings
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
von: Li, Miaomiao, et al.
Veröffentlicht: (2025)
von: Li, Miaomiao, et al.
Veröffentlicht: (2025)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
von: Bansal, Hritik, et al.
Veröffentlicht: (2023)
Online hand gesture recognition using Continual Graph Transformers
von: Slama, Rim, et al.
Veröffentlicht: (2025)
von: Slama, Rim, et al.
Veröffentlicht: (2025)
Synergistic Weak-Strong Collaboration by Aligning Preferences
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ask don't tell: Reducing sycophancy in large language models
von: Dubois, Magda, et al.
Veröffentlicht: (2026) -
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025) -
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
von: Summerfield, Christopher, et al.
Veröffentlicht: (2025) -
One-shot emergency psychiatric triage across 15 frontier AI chatbots
von: Weilnhammer, Veith, et al.
Veröffentlicht: (2026) -
TaskMet: Task-Driven Metric Learning for Model Learning
von: Bansal, Dishank, et al.
Veröffentlicht: (2023)