M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yanshu, Cao, Yi, He, Hongyang, Cheng, Qisen, Fu, Xiang, Xiao, Xi, Wang, Tianyang, Tang, Ruixiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Towards Generalizable Implicit In-Context Learning with Attention Routing
von: Li, Jiaqian, et al.
Veröffentlicht: (2025)
von: Li, Jiaqian, et al.
Veröffentlicht: (2025)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
von: Li, Yanshu
Veröffentlicht: (2025)
von: Li, Yanshu
Veröffentlicht: (2025)
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
von: Li, Jiaqian, et al.
Veröffentlicht: (2026)
von: Li, Jiaqian, et al.
Veröffentlicht: (2026)
Enhancing Fine-grained Object Detection in Aerial Images via Orthogonal Mapping
von: Zhu, Haoran, et al.
Veröffentlicht: (2024)
von: Zhu, Haoran, et al.
Veröffentlicht: (2024)
Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions
von: Ma, Chenrui, et al.
Veröffentlicht: (2025)
von: Ma, Chenrui, et al.
Veröffentlicht: (2025)
SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
von: Yi, Huahui, et al.
Veröffentlicht: (2025)
von: Yi, Huahui, et al.
Veröffentlicht: (2025)
Learning Unsupervised Semantic Document Representation for Fine-grained Aspect-based Sentiment Analysis
von: Fu, Hao-Ming, et al.
Veröffentlicht: (2024)
von: Fu, Hao-Ming, et al.
Veröffentlicht: (2024)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
von: Wu, Rui, et al.
Veröffentlicht: (2026)
von: Wu, Rui, et al.
Veröffentlicht: (2026)
Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
von: Wu, Rui, et al.
Veröffentlicht: (2025)
von: Wu, Rui, et al.
Veröffentlicht: (2025)
M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
von: Zeng, Yongjie, et al.
Veröffentlicht: (2025)
von: Zeng, Yongjie, et al.
Veröffentlicht: (2025)
Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation
von: Xiao, Xi, et al.
Veröffentlicht: (2026)
von: Xiao, Xi, et al.
Veröffentlicht: (2026)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
von: Zeng, Runjia, et al.
Veröffentlicht: (2026)
von: Zeng, Runjia, et al.
Veröffentlicht: (2026)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
von: Feng, Ruixiang, et al.
Veröffentlicht: (2025)
von: Feng, Ruixiang, et al.
Veröffentlicht: (2025)
Identity-Aware U-Net: Fine-grained Cell Segmentation via Identity-Aware Representation Learning
von: Xiao, Rui
Veröffentlicht: (2026)
von: Xiao, Rui
Veröffentlicht: (2026)
Multimodal Fine-grained Reasoning for Post Quality Evaluation
von: Guo, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoxu, et al.
Veröffentlicht: (2025)
Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
von: Jia, Zhenqi, et al.
Veröffentlicht: (2025)
von: Jia, Zhenqi, et al.
Veröffentlicht: (2025)
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
von: Yang, Jianjiang, et al.
Veröffentlicht: (2025)
von: Yang, Jianjiang, et al.
Veröffentlicht: (2025)
Towards Context-Robust LLMs: A Gated Representation Fine-tuning Approach
von: Zeng, Shenglai, et al.
Veröffentlicht: (2025)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2025)
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
Accessible Fine-grained Data Representation via Spatial Audio
von: Liu, Can, et al.
Veröffentlicht: (2026)
von: Liu, Can, et al.
Veröffentlicht: (2026)
EvaNet: Towards More Efficient and Consistent Infrared and Visible Image Fusion Assessment
von: Cheng, Chunyang, et al.
Veröffentlicht: (2026)
von: Cheng, Chunyang, et al.
Veröffentlicht: (2026)
Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual Representation
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2025)
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2025)
FLAIR: VLM with Fine-grained Language-informed Image Representations
von: Xiao, Rui, et al.
Veröffentlicht: (2024)
von: Xiao, Rui, et al.
Veröffentlicht: (2024)
Self-Constructed Context Decompilation with Fined-grained Alignment Enhancement
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
Learning Straight Flows: Variational Flow Matching for Efficient Generation
von: Ma, Chenrui, et al.
Veröffentlicht: (2025)
von: Ma, Chenrui, et al.
Veröffentlicht: (2025)
MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts
von: Li, Jiatong, et al.
Veröffentlicht: (2024)
von: Li, Jiatong, et al.
Veröffentlicht: (2024)
MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2024)
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2024)
LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation
von: Zhou, Tianyu, et al.
Veröffentlicht: (2025)
von: Zhou, Tianyu, et al.
Veröffentlicht: (2025)
ContextNav: Towards Agentic Multimodal In-Context Learning
von: Fu, Honghao, et al.
Veröffentlicht: (2025)
von: Fu, Honghao, et al.
Veröffentlicht: (2025)
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
Unifying Prediction and Explanation in Time-Series Transformers via Shapley-based Pretraining
von: Cheng, Qisen, et al.
Veröffentlicht: (2025)
von: Cheng, Qisen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025) -
Towards Generalizable Implicit In-Context Learning with Attention Routing
von: Li, Jiaqian, et al.
Veröffentlicht: (2025) -
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025) -
TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
von: Li, Yanshu, et al.
Veröffentlicht: (2025) -
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
von: Li, Yanshu
Veröffentlicht: (2025)