RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yangyang, Li, Yi-Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
von: Kubica, Dominick, et al.
Veröffentlicht: (2025)
von: Kubica, Dominick, et al.
Veröffentlicht: (2025)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
von: Shafique, Muhammad Ali, et al.
Veröffentlicht: (2026)
von: Shafique, Muhammad Ali, et al.
Veröffentlicht: (2026)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
von: Lan, Guangchen, et al.
Veröffentlicht: (2026)
von: Lan, Guangchen, et al.
Veröffentlicht: (2026)
Machine Unlearning for Masked Diffusion Language Models
von: Lee, Georu, et al.
Veröffentlicht: (2026)
von: Lee, Georu, et al.
Veröffentlicht: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
von: Young, Richard J.
Veröffentlicht: (2026)
von: Young, Richard J.
Veröffentlicht: (2026)
Truth as a Compression Artifact in Language Model Training
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
von: Vera, Patricio
Veröffentlicht: (2026)
von: Vera, Patricio
Veröffentlicht: (2026)
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law
von: Shportun, Nazarii
Veröffentlicht: (2026)
von: Shportun, Nazarii
Veröffentlicht: (2026)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models
von: Zhang, Yuxuan
Veröffentlicht: (2025)
von: Zhang, Yuxuan
Veröffentlicht: (2025)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Exemplar Retrieval Without Overhypothesis Induction: Limits of Distributional Sequence Learning in Early Word Learning
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
Eyla: Toward an Identity-Anchored LLM Architecture with Integrated Biological Priors -- Vision, Implementation Attempt, and Lessons from AI-Assisted Development
von: Aditto, Arif
Veröffentlicht: (2026)
von: Aditto, Arif
Veröffentlicht: (2026)
KAConvText: Novel Approach to Burmese Sentence Classification using Kolmogorov-Arnold Convolution
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
A Hierarchical Error Framework for Reliable Automated Coding in Communication Research: Applications to Health and Political Communication
von: Zhao, Zhilong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhilong, et al.
Veröffentlicht: (2025)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
MeMo: Towards Language Models with Associative Memory Mechanisms
von: Zanzotto, Fabio Massimo, et al.
Veröffentlicht: (2025)
von: Zanzotto, Fabio Massimo, et al.
Veröffentlicht: (2025)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
von: Liao, Jianxing, et al.
Veröffentlicht: (2025)
von: Liao, Jianxing, et al.
Veröffentlicht: (2025)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
von: Wu, Shuai, et al.
Veröffentlicht: (2026)
von: Wu, Shuai, et al.
Veröffentlicht: (2026)
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
von: Sauter, Andreas, et al.
Veröffentlicht: (2026)
von: Sauter, Andreas, et al.
Veröffentlicht: (2026)
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
von: Koc, Vincent
Veröffentlicht: (2025)
von: Koc, Vincent
Veröffentlicht: (2025)
EasyMath: A 0-shot Math Benchmark for SLMs
von: Karki, Drishya, et al.
Veröffentlicht: (2025)
von: Karki, Drishya, et al.
Veröffentlicht: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
von: Xie, Pengzhen, et al.
Veröffentlicht: (2025)
von: Xie, Pengzhen, et al.
Veröffentlicht: (2025)
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
von: Cui, Sasha, et al.
Veröffentlicht: (2025)
von: Cui, Sasha, et al.
Veröffentlicht: (2025)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
von: Kugler, Kai
Veröffentlicht: (2025)
von: Kugler, Kai
Veröffentlicht: (2025)
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
von: Heydari, Sina, et al.
Veröffentlicht: (2026)
von: Heydari, Sina, et al.
Veröffentlicht: (2026)
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
Ähnliche Einträge
-
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025) -
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
von: Kubica, Dominick, et al.
Veröffentlicht: (2025) -
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
von: Shafique, Muhammad Ali, et al.
Veröffentlicht: (2026) -
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
von: Lan, Guangchen, et al.
Veröffentlicht: (2026) -
Machine Unlearning for Masked Diffusion Language Models
von: Lee, Georu, et al.
Veröffentlicht: (2026)