What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Dong, Liang, Jian, Wang, Yanbo, Lu, Shuo, He, Ran, Tan, Tieniu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
von: Liang, Jian, et al.
Veröffentlicht: (2023)
von: Liang, Jian, et al.
Veröffentlicht: (2023)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation
von: Wang, Zhengbo, et al.
Veröffentlicht: (2026)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2026)
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
von: Liu, Yexiang, et al.
Veröffentlicht: (2025)
von: Liu, Yexiang, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
von: Yan, Dong, et al.
Veröffentlicht: (2026)
von: Yan, Dong, et al.
Veröffentlicht: (2026)
TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning
von: Wang, Yinuo, et al.
Veröffentlicht: (2026)
von: Wang, Yinuo, et al.
Veröffentlicht: (2026)
The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
von: Sheng, Lijun, et al.
Veröffentlicht: (2025)
von: Sheng, Lijun, et al.
Veröffentlicht: (2025)
Grid and Road Expressions Are Complementary for Trajectory Representation Learning
von: Zhou, Silin, et al.
Veröffentlicht: (2024)
von: Zhou, Silin, et al.
Veröffentlicht: (2024)
Frustratingly Easy Feature Reconstruction for Out-of-Distribution Detection
von: Wang, Yingsheng, et al.
Veröffentlicht: (2025)
von: Wang, Yingsheng, et al.
Veröffentlicht: (2025)
Repetitive Contrastive Learning Enhances Mamba's Selectivity in Time Series Prediction
von: Yan, Wenbo, et al.
Veröffentlicht: (2025)
von: Yan, Wenbo, et al.
Veröffentlicht: (2025)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models
von: Tan, Lit Sin, et al.
Veröffentlicht: (2026)
von: Tan, Lit Sin, et al.
Veröffentlicht: (2026)
Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning
von: Lu, Jian
Veröffentlicht: (2025)
von: Lu, Jian
Veröffentlicht: (2025)
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
von: Ma, Lu, et al.
Veröffentlicht: (2025)
von: Ma, Lu, et al.
Veröffentlicht: (2025)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
von: Hübotter, Jonas, et al.
Veröffentlicht: (2025)
von: Hübotter, Jonas, et al.
Veröffentlicht: (2025)
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
von: Li, Yibo, et al.
Veröffentlicht: (2026)
von: Li, Yibo, et al.
Veröffentlicht: (2026)
Reaching Consensus in Cooperative Multi-Agent Reinforcement Learning with Goal Imagination
von: Wang, Liangzhou, et al.
Veröffentlicht: (2024)
von: Wang, Liangzhou, et al.
Veröffentlicht: (2024)
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
von: Liu, Jia, et al.
Veröffentlicht: (2025)
von: Liu, Jia, et al.
Veröffentlicht: (2025)
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Conjunction Subspaces Test for Conformal and Selective Classification
von: He, Zengyou, et al.
Veröffentlicht: (2024)
von: He, Zengyou, et al.
Veröffentlicht: (2024)
Maximum Entropy Reinforcement Learning with Diffusion Policy
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2025)
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
von: Zhao, Chu, et al.
Veröffentlicht: (2026)
von: Zhao, Chu, et al.
Veröffentlicht: (2026)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
von: Feng, Lang, et al.
Veröffentlicht: (2026)
von: Feng, Lang, et al.
Veröffentlicht: (2026)
QSpec: Speculative Decoding with Complementary Quantization Schemes
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Machine Learning Engineering Agents
von: Yang, Sherry, et al.
Veröffentlicht: (2025)
von: Yang, Sherry, et al.
Veröffentlicht: (2025)
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting
von: Lin, Hongxiang, et al.
Veröffentlicht: (2026)
von: Lin, Hongxiang, et al.
Veröffentlicht: (2026)
Reinforcement Learning Teachers of Test Time Scaling
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
Counterfactual Explanations for Continuous Action Reinforcement Learning
von: Dong, Shuyang, et al.
Veröffentlicht: (2025)
von: Dong, Shuyang, et al.
Veröffentlicht: (2025)
Sufficient and Necessary Explanations (and What Lies in Between)
von: Bharti, Beepul, et al.
Veröffentlicht: (2024)
von: Bharti, Beepul, et al.
Veröffentlicht: (2024)
HTMformer: Hybrid Time and Multivariate Transformer for Time Series Forecasting
von: Wang, Tan, et al.
Veröffentlicht: (2025)
von: Wang, Tan, et al.
Veröffentlicht: (2025)
LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation
von: Tan, Heng, et al.
Veröffentlicht: (2025)
von: Tan, Heng, et al.
Veröffentlicht: (2025)
Conservative Distributional Reinforcement Learning with Safety Constraints
von: Zhang, Hengrui, et al.
Veröffentlicht: (2022)
von: Zhang, Hengrui, et al.
Veröffentlicht: (2022)
Understanding What Affects the Generalization Gap in Visual Reinforcement Learning: Theory and Empirical Evidence
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
Less is More: Pseudo-Label Filtering for Continual Test-Time Adaptation
von: Tan, Jiayao, et al.
Veröffentlicht: (2024)
von: Tan, Jiayao, et al.
Veröffentlicht: (2024)
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
von: Sokota, Samuel, et al.
Veröffentlicht: (2025)
von: Sokota, Samuel, et al.
Veröffentlicht: (2025)
Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning
von: Han, Seungyub, et al.
Veröffentlicht: (2026)
von: Han, Seungyub, et al.
Veröffentlicht: (2026)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
von: Liang, Jian, et al.
Veröffentlicht: (2023) -
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024) -
Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation
von: Wang, Zhengbo, et al.
Veröffentlicht: (2026) -
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024) -
Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
von: Liu, Yexiang, et al.
Veröffentlicht: (2025)