Model Agreement via Anchoring
Fuente:
arXiv
Saved in:
| Main Authors: | Eaton, Eric, Goel, Surbhi, Hussing, Marcel, Kearns, Michael, Roth, Aaron, Sengupta, Sikata Bela, Sorrell, Jessica |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Oracle-Efficient Reinforcement Learning for Max Value Ensembles
by: Hussing, Marcel, et al.
Published: (2024)
by: Hussing, Marcel, et al.
Published: (2024)
Replicable Reinforcement Learning with Linear Function Approximation
by: Eaton, Eric, et al.
Published: (2025)
by: Eaton, Eric, et al.
Published: (2025)
Intersectional Fairness in Reinforcement Learning with Large State and Constraint Spaces
by: Eaton, Eric, et al.
Published: (2025)
by: Eaton, Eric, et al.
Published: (2025)
Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
by: Hussing, Marcel, et al.
Published: (2024)
by: Hussing, Marcel, et al.
Published: (2024)
Robotic Manipulation Datasets for Offline Compositional Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2023)
by: Hussing, Marcel, et al.
Published: (2023)
Behavior-Consistent Deep Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2026)
by: Hussing, Marcel, et al.
Published: (2026)
Model Ensembling for Constrained Optimization
by: Globus-Harris, Ira, et al.
Published: (2024)
by: Globus-Harris, Ira, et al.
Published: (2024)
Tractable Agreement Protocols
by: Collina, Natalie, et al.
Published: (2024)
by: Collina, Natalie, et al.
Published: (2024)
Probabilistic Stability Guarantees for Feature Attributions
by: Jin, Helen, et al.
Published: (2025)
by: Jin, Helen, et al.
Published: (2025)
Multi-Objective Reinforcement Learning for Large-Scale Tote Allocation in Human-Robot Collaborative Fulfillment Centers
by: Sengupta, Sikata, et al.
Published: (2026)
by: Sengupta, Sikata, et al.
Published: (2026)
Distributed Continual Learning
by: Le, Long, et al.
Published: (2024)
by: Le, Long, et al.
Published: (2024)
Can we hop in general? A discussion of benchmark selection and design using the Hopper environment
by: Voelcker, Claas A, et al.
Published: (2024)
by: Voelcker, Claas A, et al.
Published: (2024)
Why Do Transformers Fail to Forecast Time Series In-Context?
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
by: Xue, Anton, et al.
Published: (2024)
by: Xue, Anton, et al.
Published: (2024)
Collaborative Prediction: Tractable Information Aggregation via Agreement
by: Collina, Natalie, et al.
Published: (2025)
by: Collina, Natalie, et al.
Published: (2025)
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026)
by: Kearns, Ryan Othniel
Published: (2026)
Auditing Language Model Unlearning via Information Decomposition
by: Goel, Anmol, et al.
Published: (2026)
by: Goel, Anmol, et al.
Published: (2026)
World Model Robustness via Surprise Recognition
by: Zollicoffer, Geigh, et al.
Published: (2025)
by: Zollicoffer, Geigh, et al.
Published: (2025)
MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
by: Voelcker, Claas A, et al.
Published: (2024)
by: Voelcker, Claas A, et al.
Published: (2024)
Toward Human-AI Complementarity Across Diverse Tasks
by: Xu, Yuzheng, et al.
Published: (2026)
by: Xu, Yuzheng, et al.
Published: (2026)
Influence functions and regularity tangents for efficient active learning
by: Eaton, Frederik
Published: (2024)
by: Eaton, Frederik
Published: (2024)
A Theory of Learning with Autoregressive Chain of Thought
by: Joshi, Nirmit, et al.
Published: (2025)
by: Joshi, Nirmit, et al.
Published: (2025)
Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation
by: Sovrano, Francesco, et al.
Published: (2026)
by: Sovrano, Francesco, et al.
Published: (2026)
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
by: Eaton, Kenneth, et al.
Published: (2024)
by: Eaton, Kenneth, et al.
Published: (2024)
Using Analytics on Student Created Data to Content Validate Pedagogical Tools
by: Kos, John, et al.
Published: (2023)
by: Kos, John, et al.
Published: (2023)
Sycophantic Anchors: Localizing and Quantifying User Agreement in Reasoning Models
by: Duszenko, Jacek
Published: (2026)
by: Duszenko, Jacek
Published: (2026)
Significativity Indices for Agreement Values
by: Casagrande, Alberto, et al.
Published: (2025)
by: Casagrande, Alberto, et al.
Published: (2025)
Representation Without Reward: A JEPA Audit for LLM Fine-Tuning
by: Sengupta, Biswa
Published: (2026)
by: Sengupta, Biswa
Published: (2026)
JPmHC Dynamical Isometry via Orthogonal Hyper-Connections
by: Sengupta, Biswa, et al.
Published: (2026)
by: Sengupta, Biswa, et al.
Published: (2026)
Improving LLM Group Fairness on Tabular Data via In-Context Learning
by: Cherepanova, Valeriia, et al.
Published: (2024)
by: Cherepanova, Valeriia, et al.
Published: (2024)
Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse Agents
by: Kiyani, Shayan, et al.
Published: (2025)
by: Kiyani, Shayan, et al.
Published: (2025)
Robust Decision Making with Partially Calibrated Forecasts
by: Kiyani, Shayan, et al.
Published: (2025)
by: Kiyani, Shayan, et al.
Published: (2025)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
by: Goel, Ayush, et al.
Published: (2026)
by: Goel, Ayush, et al.
Published: (2026)
Emergent Alignment via Competition
by: Collina, Natalie, et al.
Published: (2025)
by: Collina, Natalie, et al.
Published: (2025)
Conformal Language Model Reasoning with Coherent Factuality
by: Rubin-Toles, Maxon, et al.
Published: (2025)
by: Rubin-Toles, Maxon, et al.
Published: (2025)
ADPO: Anchored Direct Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Balanced Filtering via Disclosure-Controlled Proxies
by: Deng, Siqi, et al.
Published: (2023)
by: Deng, Siqi, et al.
Published: (2023)
Can Active Label Correction Improve LLM-based Modular AI Systems?
by: Taneja, Karan, et al.
Published: (2024)
by: Taneja, Karan, et al.
Published: (2024)
Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI
by: Aijaz, Aisha, et al.
Published: (2026)
by: Aijaz, Aisha, et al.
Published: (2026)
HARBOR: Automated Harness Optimization
by: Sengupta, Biswa, et al.
Published: (2026)
by: Sengupta, Biswa, et al.
Published: (2026)
Similar Items
-
Oracle-Efficient Reinforcement Learning for Max Value Ensembles
by: Hussing, Marcel, et al.
Published: (2024) -
Replicable Reinforcement Learning with Linear Function Approximation
by: Eaton, Eric, et al.
Published: (2025) -
Intersectional Fairness in Reinforcement Learning with Large State and Constraint Spaces
by: Eaton, Eric, et al.
Published: (2025) -
Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
by: Hussing, Marcel, et al.
Published: (2024) -
Robotic Manipulation Datasets for Offline Compositional Reinforcement Learning
by: Hussing, Marcel, et al.
Published: (2023)