Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Yiming, Yang, Yujiu, Wang, Junjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
RaVL: Discovering and Mitigating Spurious Correlations in Fine-Tuned Vision-Language Models
von: Varma, Maya, et al.
Veröffentlicht: (2024)
von: Varma, Maya, et al.
Veröffentlicht: (2024)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
von: Yu, Bo, et al.
Veröffentlicht: (2026)
von: Yu, Bo, et al.
Veröffentlicht: (2026)
GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
von: Wang, Chun, et al.
Veröffentlicht: (2025)
von: Wang, Chun, et al.
Veröffentlicht: (2025)
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Hierarchy-Aware Fine-Tuning of Vision-Language Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
von: Wu, Xiangyang, et al.
Veröffentlicht: (2025)
von: Wu, Xiangyang, et al.
Veröffentlicht: (2025)
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
von: Zhang, Yizhen, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhen, et al.
Veröffentlicht: (2025)
Prompting Large Vision-Language Models for Compositional Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
Towards Calibrated Robust Fine-Tuning of Vision-Language Models
von: Oh, Changdae, et al.
Veröffentlicht: (2023)
von: Oh, Changdae, et al.
Veröffentlicht: (2023)
Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship Classification
von: Lan, Long, et al.
Veröffentlicht: (2024)
von: Lan, Long, et al.
Veröffentlicht: (2024)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
von: Ratzlaff, Neale, et al.
Veröffentlicht: (2024)
von: Ratzlaff, Neale, et al.
Veröffentlicht: (2024)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
von: Yang, Yiming, et al.
Veröffentlicht: (2025)
von: Yang, Yiming, et al.
Veröffentlicht: (2025)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
von: Yan, Qiao, et al.
Veröffentlicht: (2025)
von: Yan, Qiao, et al.
Veröffentlicht: (2025)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
Fine-Tuning Vision-Language Model for Automated Engineering Drawing Information Extraction
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2024)
Robust Depth Enhancement via Polarization Prompt Fusion Tuning
von: Ikemura, Kei, et al.
Veröffentlicht: (2024)
von: Ikemura, Kei, et al.
Veröffentlicht: (2024)
Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation
von: Liu, Yaohua, et al.
Veröffentlicht: (2025)
von: Liu, Yaohua, et al.
Veröffentlicht: (2025)
Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy
von: Yang, Yiting, et al.
Veröffentlicht: (2025)
von: Yang, Yiting, et al.
Veröffentlicht: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
von: Chen, Jun, et al.
Veröffentlicht: (2024)
von: Chen, Jun, et al.
Veröffentlicht: (2024)
Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2026)
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2026)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
von: Cai, Xinyan, et al.
Veröffentlicht: (2025)
von: Cai, Xinyan, et al.
Veröffentlicht: (2025)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
von: Sun, Haoyuan, et al.
Veröffentlicht: (2025)
von: Sun, Haoyuan, et al.
Veröffentlicht: (2025)
SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
von: Luo, Junwei, et al.
Veröffentlicht: (2024)
von: Luo, Junwei, et al.
Veröffentlicht: (2024)
Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
von: Hu, Rui, et al.
Veröffentlicht: (2024)
von: Hu, Rui, et al.
Veröffentlicht: (2024)
Adversarial Prompt Tuning for Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
von: Zhu, Wenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Wenhui, et al.
Veröffentlicht: (2025)
Confounder-Aware Medical Data Selection for Fine-Tuning Pretrained Vision Models
von: Ji, Anyang, et al.
Veröffentlicht: (2025)
von: Ji, Anyang, et al.
Veröffentlicht: (2025)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Improving Medical Visual Reinforcement Fine-Tuning via Perception and Reasoning Augmentation
von: Yang, Guangjing, et al.
Veröffentlicht: (2026)
von: Yang, Guangjing, et al.
Veröffentlicht: (2026)
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights
von: Tsolissou, Daphne, et al.
Veröffentlicht: (2025)
von: Tsolissou, Daphne, et al.
Veröffentlicht: (2025)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
Fine-tuning Pre-trained Vision-Language Models in a Human-Annotation-Free Manner
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
von: Tan, Huajie, et al.
Veröffentlicht: (2025) -
RaVL: Discovering and Mitigating Spurious Correlations in Fine-Tuned Vision-Language Models
von: Varma, Maya, et al.
Veröffentlicht: (2024) -
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
von: Yu, Bo, et al.
Veröffentlicht: (2026) -
GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
von: Wang, Chun, et al.
Veröffentlicht: (2025) -
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)