Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ratzlaff, Neale, Luo, Man, Su, Xin, Lal, Vasudev, Howard, Phillip |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
por: Madasu, Avinash, et al.
Publicado: (2025)
por: Madasu, Avinash, et al.
Publicado: (2025)
DPO Learning with LLMs-Judge Signal for Computer Use Agents
por: Luo, Man, et al.
Publicado: (2025)
por: Luo, Man, et al.
Publicado: (2025)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
por: Madasu, Avinash, et al.
Publicado: (2025)
por: Madasu, Avinash, et al.
Publicado: (2025)
SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples
por: Howard, Phillip, et al.
Publicado: (2023)
por: Howard, Phillip, et al.
Publicado: (2023)
Quantifying and Enabling the Interpretability of CLIP-like Models
por: Madasu, Avinash, et al.
Publicado: (2024)
por: Madasu, Avinash, et al.
Publicado: (2024)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
por: Ratzlaff, Neale, et al.
Publicado: (2024)
por: Ratzlaff, Neale, et al.
Publicado: (2024)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
por: Su, Xin, et al.
Publicado: (2024)
por: Su, Xin, et al.
Publicado: (2024)
Probing the Representational Power of Sparse Autoencoders in Vision Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
por: Ratzlaff, Neale, et al.
Publicado: (2024)
por: Ratzlaff, Neale, et al.
Publicado: (2024)
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
por: Aflalo, Estelle, et al.
Publicado: (2024)
por: Aflalo, Estelle, et al.
Publicado: (2024)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
por: Madasu, Avinash, et al.
Publicado: (2023)
por: Madasu, Avinash, et al.
Publicado: (2023)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
por: Wu, Changti, et al.
Publicado: (2026)
por: Wu, Changti, et al.
Publicado: (2026)
Cross-Cultural Value Awareness in Large Vision-Language Models
por: Howard, Phillip, et al.
Publicado: (2026)
por: Howard, Phillip, et al.
Publicado: (2026)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
por: Kundu, Souvik, et al.
Publicado: (2025)
por: Kundu, Souvik, et al.
Publicado: (2025)
Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation
por: Hao, Chao, et al.
Publicado: (2026)
por: Hao, Chao, et al.
Publicado: (2026)
mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval
por: Kim, Kyeong Seon, et al.
Publicado: (2026)
por: Kim, Kyeong Seon, et al.
Publicado: (2026)
Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
por: Hu, Rui, et al.
Publicado: (2024)
por: Hu, Rui, et al.
Publicado: (2024)
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
por: Guo, Haiyang, et al.
Publicado: (2025)
por: Guo, Haiyang, et al.
Publicado: (2025)
PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology
por: Wu, Xiaomin, et al.
Publicado: (2024)
por: Wu, Xiaomin, et al.
Publicado: (2024)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
por: Ren, Yiming, et al.
Publicado: (2026)
por: Ren, Yiming, et al.
Publicado: (2026)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
por: Yan, Qiao, et al.
Publicado: (2025)
por: Yan, Qiao, et al.
Publicado: (2025)
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
por: Zeng, Xingchen, et al.
Publicado: (2024)
por: Zeng, Xingchen, et al.
Publicado: (2024)
Steering Large Language Models to Evaluate and Amplify Creativity
por: Olson, Matthew Lyle, et al.
Publicado: (2024)
por: Olson, Matthew Lyle, et al.
Publicado: (2024)
Instruction-tuned Self-Questioning Framework for Multimodal Reasoning
por: Jang, You-Won, et al.
Publicado: (2025)
por: Jang, You-Won, et al.
Publicado: (2025)
Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
por: Lin, Chenchen, et al.
Publicado: (2026)
por: Lin, Chenchen, et al.
Publicado: (2026)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
por: Lakhanpal, Sanyam, et al.
Publicado: (2024)
por: Lakhanpal, Sanyam, et al.
Publicado: (2024)
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
por: Xi, Gongli, et al.
Publicado: (2026)
por: Xi, Gongli, et al.
Publicado: (2026)
BAMI: Training-Free Bias Mitigation in GUI Grounding
por: Zhang, Borui, et al.
Publicado: (2026)
por: Zhang, Borui, et al.
Publicado: (2026)
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
por: Dong, Mingkang, et al.
Publicado: (2026)
por: Dong, Mingkang, et al.
Publicado: (2026)
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
por: Li, Pengteng, et al.
Publicado: (2025)
por: Li, Pengteng, et al.
Publicado: (2025)
Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
por: Jung, Ji Hyeok, et al.
Publicado: (2024)
por: Jung, Ji Hyeok, et al.
Publicado: (2024)
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
por: Sun, Haoyuan, et al.
Publicado: (2025)
por: Sun, Haoyuan, et al.
Publicado: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
por: Jin, Hongbo, et al.
Publicado: (2025)
por: Jin, Hongbo, et al.
Publicado: (2025)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
por: Ou, Siqu, et al.
Publicado: (2025)
por: Ou, Siqu, et al.
Publicado: (2025)
Uncovering Bias in Large Vision-Language Models with Counterfactuals
por: Howard, Phillip, et al.
Publicado: (2024)
por: Howard, Phillip, et al.
Publicado: (2024)
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
por: Fang, Yiyang, et al.
Publicado: (2026)
por: Fang, Yiyang, et al.
Publicado: (2026)
Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models
por: Peng, Ruiying, et al.
Publicado: (2026)
por: Peng, Ruiying, et al.
Publicado: (2026)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
por: Liu, Junming, et al.
Publicado: (2025)
por: Liu, Junming, et al.
Publicado: (2025)
Ejemplares similares
-
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
por: Madasu, Avinash, et al.
Publicado: (2025) -
DPO Learning with LLMs-Judge Signal for Computer Use Agents
por: Luo, Man, et al.
Publicado: (2025) -
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
por: Madasu, Avinash, et al.
Publicado: (2025) -
SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples
por: Howard, Phillip, et al.
Publicado: (2023) -
Quantifying and Enabling the Interpretability of CLIP-like Models
por: Madasu, Avinash, et al.
Publicado: (2024)