Saved in:
| Main Author: | Zang, Jianxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.02618 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modeling Selective Feature Attention for Representation-based Siamese Text Matching
by: Zang, Jianxiang, et al.
Published: (2024)
by: Zang, Jianxiang, et al.
Published: (2024)
Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
by: Zang, Jianxiang, et al.
Published: (2025)
by: Zang, Jianxiang, et al.
Published: (2025)
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
by: Zang, Jianxiang, et al.
Published: (2025)
by: Zang, Jianxiang, et al.
Published: (2025)
RRM: Robust Reward Model Training Mitigates Reward Hacking
by: Liu, Tianqi, et al.
Published: (2024)
by: Liu, Tianqi, et al.
Published: (2024)
Explanation based Bias Decoupling Regularization for Natural Language Inference
by: Zang, Jianxiang, et al.
Published: (2024)
by: Zang, Jianxiang, et al.
Published: (2024)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
by: Turpin, Miles, et al.
Published: (2025)
by: Turpin, Miles, et al.
Published: (2025)
Spontaneous Reward Hacking in Iterative Self-Refinement
by: Pan, Jane, et al.
Published: (2024)
by: Pan, Jane, et al.
Published: (2024)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
by: Wang, Songtao, et al.
Published: (2026)
by: Wang, Songtao, et al.
Published: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
by: Deng, Wenlong, et al.
Published: (2026)
by: Deng, Wenlong, et al.
Published: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
by: Pan, Alexander, et al.
Published: (2024)
by: Pan, Alexander, et al.
Published: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
by: Ferreira, Pedro, et al.
Published: (2025)
by: Ferreira, Pedro, et al.
Published: (2025)
Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation
by: Liu, Yifeng, et al.
Published: (2026)
by: Liu, Yifeng, et al.
Published: (2026)
Discriminative Policy Optimization for Token-Level Reward Models
by: Chen, Hongzhan, et al.
Published: (2025)
by: Chen, Hongzhan, et al.
Published: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
by: Wu, Rui, et al.
Published: (2026)
by: Wu, Rui, et al.
Published: (2026)
Monitoring Emergent Reward Hacking During Generation via Internal Activations
by: Wilhelm, Patrick, et al.
Published: (2026)
by: Wilhelm, Patrick, et al.
Published: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
S2Sent: Nested Selectivity Aware Sentence Representation Learning
by: Zang, Jianxiang, et al.
Published: (2025)
by: Zang, Jianxiang, et al.
Published: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
by: Ackermann, Johannes, et al.
Published: (2026)
by: Ackermann, Johannes, et al.
Published: (2026)
Pre-Trained Policy Discriminators are General Reward Models
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference Data
by: Guo, Siqi, et al.
Published: (2025)
by: Guo, Siqi, et al.
Published: (2025)
Self-Calibrating Language Models via Test-Time Discriminative Distillation
by: Hedna, Mohamed Rissal, et al.
Published: (2026)
by: Hedna, Mohamed Rissal, et al.
Published: (2026)
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
by: Schulhoff, Sander, et al.
Published: (2023)
by: Schulhoff, Sander, et al.
Published: (2023)
What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA
by: Ji, Yuelyu, et al.
Published: (2026)
by: Ji, Yuelyu, et al.
Published: (2026)
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
Tailoring Self-Rationalizers with Multi-Reward Distillation
by: Ramnath, Sahana, et al.
Published: (2023)
by: Ramnath, Sahana, et al.
Published: (2023)
Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information
by: Chen, Yao, et al.
Published: (2026)
by: Chen, Yao, et al.
Published: (2026)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
by: Peng, Hao, et al.
Published: (2026)
by: Peng, Hao, et al.
Published: (2026)
LLMR: Knowledge Distillation with a Large Language Model-Induced Reward
by: Li, Dongheng, et al.
Published: (2024)
by: Li, Dongheng, et al.
Published: (2024)
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
by: Zhu, Zining, et al.
Published: (2025)
by: Zhu, Zining, et al.
Published: (2025)
Similar Items
-
Modeling Selective Feature Attention for Representation-based Siamese Text Matching
by: Zang, Jianxiang, et al.
Published: (2024) -
Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
by: Zang, Jianxiang, et al.
Published: (2025) -
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
by: Zang, Jianxiang, et al.
Published: (2025) -
RRM: Robust Reward Model Training Mitigates Reward Hacking
by: Liu, Tianqi, et al.
Published: (2024) -
Explanation based Bias Decoupling Regularization for Natural Language Inference
by: Zang, Jianxiang, et al.
Published: (2024)