On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Rosie, Shah, Anshul, Zhu, Xiaoyu, Deng, Xinke, Jiang, Zhongyu, Yang, Yang, Liebelt, Joerg, Mondal, Arnab |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Driven Prolog-based Chain-of-Thought
by: Tan, Xiaoyu, et al.
Published: (2024)
by: Tan, Xiaoyu, et al.
Published: (2024)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
by: Ubukata, Shunsuke
Published: (2026)
by: Ubukata, Shunsuke
Published: (2026)
Random Scaling of Emergent Capabilities
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
by: Wei, Kaiwen, et al.
Published: (2025)
by: Wei, Kaiwen, et al.
Published: (2025)
Efficient Reasoning via Thought-Training and Thought-Free Inference
by: Wu, Canhui, et al.
Published: (2025)
by: Wu, Canhui, et al.
Published: (2025)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
by: Klöser, Lars, et al.
Published: (2024)
by: Klöser, Lars, et al.
Published: (2024)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
by: Zhao, Yong, et al.
Published: (2025)
by: Zhao, Yong, et al.
Published: (2025)
A Likelihood Ratio Test of Genetic Relationship among Languages
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
Chain of Draft: Thinking Faster by Writing Less
by: Xu, Silei, et al.
Published: (2025)
by: Xu, Silei, et al.
Published: (2025)
Overcoming the Generalization Limits of SLM Finetuning for Shape-Based Extraction of Datatype and Object Properties
by: Ringwald, Célian, et al.
Published: (2025)
by: Ringwald, Célian, et al.
Published: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
CauESC: A Causal Aware Model for Emotional Support Conversation
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
by: Hong, Colin, et al.
Published: (2025)
by: Hong, Colin, et al.
Published: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
by: Chen, Xinjie, et al.
Published: (2026)
by: Chen, Xinjie, et al.
Published: (2026)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
by: Nikankin, Yaniv, et al.
Published: (2025)
by: Nikankin, Yaniv, et al.
Published: (2025)
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
by: Aslam, Nazia, et al.
Published: (2026)
by: Aslam, Nazia, et al.
Published: (2026)
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
by: Le, Nguyen-Khang, et al.
Published: (2025)
by: Le, Nguyen-Khang, et al.
Published: (2025)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Large Language Models Can Better Understand Knowledge Graphs Than We Thought
by: Dai, Xinbang, et al.
Published: (2024)
by: Dai, Xinbang, et al.
Published: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
Evaluation Before Generation: A Paradigm for Robust Multimodal Sentiment Analysis with Missing Modalities
by: Chen, Rongfei, et al.
Published: (2026)
by: Chen, Rongfei, et al.
Published: (2026)
The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
by: Dugan, Liam, et al.
Published: (2024)
by: Dugan, Liam, et al.
Published: (2024)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
by: Bajpai, Ashutosh, et al.
Published: (2025)
by: Bajpai, Ashutosh, et al.
Published: (2025)
MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr
by: Yang, Xuwen
Published: (2025)
by: Yang, Xuwen
Published: (2025)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
by: Chen, Yanbing, et al.
Published: (2024)
by: Chen, Yanbing, et al.
Published: (2024)
Robustness of Large Language Models to Perturbations in Text
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Low-Resource Court Judgment Summarization for Common Law Systems
by: Liu, Shuaiqi, et al.
Published: (2024)
by: Liu, Shuaiqi, et al.
Published: (2024)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2025)
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2025)
SLM Finetuning for Natural Language to Domain Specific Code Generation in Production
by: Nair, Renjini R., et al.
Published: (2026)
by: Nair, Renjini R., et al.
Published: (2026)
Technical Report of TeleChat2, TeleChat2.5 and T1
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Similar Items
-
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025) -
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025) -
Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Driven Prolog-based Chain-of-Thought
by: Tan, Xiaoyu, et al.
Published: (2024) -
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024) -
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
by: Ubukata, Shunsuke
Published: (2026)