GCPO: When Contrast Fails, Go Gold
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Hao, Liu, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
por: Liu, Muxin, et al.
Publicado: (2026)
por: Liu, Muxin, et al.
Publicado: (2026)
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
por: Li, Xiaoxiao
Publicado: (2026)
por: Li, Xiaoxiao
Publicado: (2026)
When Mean CE Fails: Median CE Can Better Track Language Model Quality
por: Guo, Hao, et al.
Publicado: (2026)
por: Guo, Hao, et al.
Publicado: (2026)
When Absolute State Fails: Evaluating Proprioceptive Encodings for Robust Manipulation
por: Alvarez, Maxime, et al.
Publicado: (2026)
por: Alvarez, Maxime, et al.
Publicado: (2026)
When Counterfactual Reasoning Fails: Chaos and Real-World Complexity
por: Aalaila, Yahya, et al.
Publicado: (2025)
por: Aalaila, Yahya, et al.
Publicado: (2025)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
por: Wang, Yongxin, et al.
Publicado: (2025)
por: Wang, Yongxin, et al.
Publicado: (2025)
When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series
por: Park, Min-Yeong, et al.
Publicado: (2025)
por: Park, Min-Yeong, et al.
Publicado: (2025)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
por: Yin, Hao, et al.
Publicado: (2025)
por: Yin, Hao, et al.
Publicado: (2025)
When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL
por: Wang, Youting, et al.
Publicado: (2026)
por: Wang, Youting, et al.
Publicado: (2026)
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
por: Rakshit, Sushrita, et al.
Publicado: (2026)
por: Rakshit, Sushrita, et al.
Publicado: (2026)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
por: Mehrafarin, Houman, et al.
Publicado: (2026)
por: Mehrafarin, Houman, et al.
Publicado: (2026)
When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE
por: Erol, Melihcan, et al.
Publicado: (2026)
por: Erol, Melihcan, et al.
Publicado: (2026)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
por: Huang, Donghao, et al.
Publicado: (2026)
por: Huang, Donghao, et al.
Publicado: (2026)
Frequency Matters: When Time Series Foundation Models Fail Under Spectral Shift
por: Wang, Tianze, et al.
Publicado: (2025)
por: Wang, Tianze, et al.
Publicado: (2025)
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
por: Lu, Ruofan, et al.
Publicado: (2025)
por: Lu, Ruofan, et al.
Publicado: (2025)
When Sensors Fail: Temporal Sequence Models for Robust PPO under Sensor Drift
por: Vogt-Lowell, Kevin, et al.
Publicado: (2026)
por: Vogt-Lowell, Kevin, et al.
Publicado: (2026)
When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
por: Kim, Hanna, et al.
Publicado: (2024)
por: Kim, Hanna, et al.
Publicado: (2024)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
por: Khullar, Dipika, et al.
Publicado: (2026)
por: Khullar, Dipika, et al.
Publicado: (2026)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
por: Schnabel, Tobias, et al.
Publicado: (2025)
por: Schnabel, Tobias, et al.
Publicado: (2025)
When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
por: Mao, Ruixiang, et al.
Publicado: (2026)
por: Mao, Ruixiang, et al.
Publicado: (2026)
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
por: Moure, Pehuén, et al.
Publicado: (2026)
por: Moure, Pehuén, et al.
Publicado: (2026)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
por: Yan, Yuping, et al.
Publicado: (2025)
por: Yan, Yuping, et al.
Publicado: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
por: Hadeliya, Tsimur, et al.
Publicado: (2025)
por: Hadeliya, Tsimur, et al.
Publicado: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
por: Rehman, Mohammad Abdul, et al.
Publicado: (2025)
por: Rehman, Mohammad Abdul, et al.
Publicado: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
por: Landesberg, Eddie
Publicado: (2026)
por: Landesberg, Eddie
Publicado: (2026)
More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration
por: Yadav, Advait, et al.
Publicado: (2026)
por: Yadav, Advait, et al.
Publicado: (2026)
When Agents Go Quiet: Output Generation Capacity and Format-Cost Separation for LLM Document Synthesis
por: Agyemang, Justice Owusu, et al.
Publicado: (2026)
por: Agyemang, Justice Owusu, et al.
Publicado: (2026)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
por: Wang, Zehao, et al.
Publicado: (2026)
por: Wang, Zehao, et al.
Publicado: (2026)
Solving 7x7 Killall-Go with Seki Database
por: Tsai, Yun-Jui, et al.
Publicado: (2024)
por: Tsai, Yun-Jui, et al.
Publicado: (2024)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
When Learning Rates Go Wrong: Early Structural Signals in PPO Actor-Critic
por: Fernández-Hernández, Alberto, et al.
Publicado: (2026)
por: Fernández-Hernández, Alberto, et al.
Publicado: (2026)
Where LLM Agents Fail and How They can Learn From Failures
por: Zhu, Kunlun, et al.
Publicado: (2025)
por: Zhu, Kunlun, et al.
Publicado: (2025)
When Validation Fails: Cross-Institutional Blood Pressure Prediction and the Limits of Electronic Health Record-Based Models
por: Azam, Md Basit, et al.
Publicado: (2025)
por: Azam, Md Basit, et al.
Publicado: (2025)
Moving On, Even When You're Broken: Fail-Active Trajectory Generation via Diffusion Policies Conditioned on Embodiment and Task
por: Briscoe-Martinez, Gilberto G., et al.
Publicado: (2026)
por: Briscoe-Martinez, Gilberto G., et al.
Publicado: (2026)
Learning Through Noise: Why Subliminal Learning Works and When It Fails
por: Brockers, Vincent C., et al.
Publicado: (2026)
por: Brockers, Vincent C., et al.
Publicado: (2026)
When Stability Fails: Hidden Failure Modes Of LLMS in Data-Constrained Scientific Decision-Making
por: Riasat, Nazia
Publicado: (2026)
por: Riasat, Nazia
Publicado: (2026)
Ascent Fails to Forget
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2025)
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2025)
The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
por: Cohen, Ruth, et al.
Publicado: (2026)
por: Cohen, Ruth, et al.
Publicado: (2026)
When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies
por: Popchanovska, Evgenija, et al.
Publicado: (2026)
por: Popchanovska, Evgenija, et al.
Publicado: (2026)
When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
por: Chen, Shenyang, et al.
Publicado: (2026)
por: Chen, Shenyang, et al.
Publicado: (2026)
Ejemplares similares
-
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
por: Liu, Muxin, et al.
Publicado: (2026) -
When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
por: Li, Xiaoxiao
Publicado: (2026) -
When Mean CE Fails: Median CE Can Better Track Language Model Quality
por: Guo, Hao, et al.
Publicado: (2026) -
When Absolute State Fails: Evaluating Proprioceptive Encodings for Robust Manipulation
por: Alvarez, Maxime, et al.
Publicado: (2026) -
When Counterfactual Reasoning Fails: Chaos and Real-World Complexity
por: Aalaila, Yahya, et al.
Publicado: (2025)