Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khayatan, Pegah, Shukor, Mustafa, Parekh, Jayneel, Dapogny, Arnaud, Cord, Matthieu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning to Steer: Input-dependent Steering for Multimodal LLMs
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
A Concept-Based Explainability Framework for Large Multimodal Models
von: Parekh, Jayneel, et al.
Veröffentlicht: (2024)
von: Parekh, Jayneel, et al.
Veröffentlicht: (2024)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
Skipping Computations in Multimodal LLMs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
What Makes Multimodal In-Context Learning Work?
von: Baldassini, Folco Bertini, et al.
Veröffentlicht: (2024)
von: Baldassini, Folco Bertini, et al.
Veröffentlicht: (2024)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2024)
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2024)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
von: Shukor, Mustafa, et al.
Veröffentlicht: (2023)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2023)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
von: Caselles-Dupré, Hugo, et al.
Veröffentlicht: (2026)
von: Caselles-Dupré, Hugo, et al.
Veröffentlicht: (2026)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
von: Yu, Hao, et al.
Veröffentlicht: (2025)
von: Yu, Hao, et al.
Veröffentlicht: (2025)
ReGentS: Real-World Safety-Critical Driving Scenario Generation Made Stable
von: Yin, Yuan, et al.
Veröffentlicht: (2024)
von: Yin, Yuan, et al.
Veröffentlicht: (2024)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
von: Mayladan, Ali, et al.
Veröffentlicht: (2023)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
Steering the Verifiability of Multimodal AI Hallucinations
von: Pang, Jianhong, et al.
Veröffentlicht: (2026)
von: Pang, Jianhong, et al.
Veröffentlicht: (2026)
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
von: Zhu, Zhihao, et al.
Veröffentlicht: (2026)
von: Zhu, Zhihao, et al.
Veröffentlicht: (2026)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
von: Li, Yunxin, et al.
Veröffentlicht: (2023)
von: Li, Yunxin, et al.
Veröffentlicht: (2023)
LLMs Can Compensate for Deficiencies in Visual Representations
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
Scientific Reasoning: Assessment of Multimodal Generative LLMs
von: Dreyer, Florian, et al.
Veröffentlicht: (2025)
von: Dreyer, Florian, et al.
Veröffentlicht: (2025)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
von: Wang, Zihang, et al.
Veröffentlicht: (2026)
von: Wang, Zihang, et al.
Veröffentlicht: (2026)
Restyling Unsupervised Concept Based Interpretable Networks with Generative Models
von: Parekh, Jayneel, et al.
Veröffentlicht: (2024)
von: Parekh, Jayneel, et al.
Veröffentlicht: (2024)
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
von: Couairon, Paul, et al.
Veröffentlicht: (2024)
von: Couairon, Paul, et al.
Veröffentlicht: (2024)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
Moment Sampling in Video LLMs for Long-Form Video QA
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
von: Horawalavithana, Sameera, et al.
Veröffentlicht: (2026)
von: Horawalavithana, Sameera, et al.
Veröffentlicht: (2026)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
von: Sarto, Sara, et al.
Veröffentlicht: (2025)
von: Sarto, Sara, et al.
Veröffentlicht: (2025)
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
von: Onsu, Murat Arda, et al.
Veröffentlicht: (2025)
von: Onsu, Murat Arda, et al.
Veröffentlicht: (2025)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
Veagle: Advancements in Multimodal Representation Learning
von: Chawla, Rajat, et al.
Veröffentlicht: (2024)
von: Chawla, Rajat, et al.
Veröffentlicht: (2024)
Orthogonal Finetuning Made Scalable
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
Scaling Laws for Native Multimodal Models
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning to Steer: Input-dependent Steering for Multimodal LLMs
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025) -
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026) -
A Concept-Based Explainability Framework for Large Multimodal Models
von: Parekh, Jayneel, et al.
Veröffentlicht: (2024) -
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024) -
Skipping Computations in Multimodal LLMs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)