ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Wenjie, Wu, Hao, Qiu, Xin, Wang, Xudong, Fan, Yingqi, Zhang, Yihan, Zhao, Anhao, Ma, Yunpu, Shen, Xiaoyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StreamingThinker: Large Language Models Can Think While Reading
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
by: Fan, Yingqi, et al.
Published: (2025)
by: Fan, Yingqi, et al.
Published: (2025)
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
by: Fan, Yingqi, et al.
Published: (2026)
by: Fan, Yingqi, et al.
Published: (2026)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
ViCA-NeRF: View-Consistency-Aware 3D Editing of Neural Radiance Fields
by: Dong, Jiahua, et al.
Published: (2024)
by: Dong, Jiahua, et al.
Published: (2024)
Rethinking the Role of LLMs in Time Series Forecasting
by: Qiu, Xin, et al.
Published: (2026)
by: Qiu, Xin, et al.
Published: (2026)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
by: Zhao, Anhao, et al.
Published: (2025)
by: Zhao, Anhao, et al.
Published: (2025)
The Few Govern the Many:Unveiling Few-Layer Dominance for Time Series Models
by: Qiu, Xin, et al.
Published: (2025)
by: Qiu, Xin, et al.
Published: (2025)
Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices
by: Lin, Junyan, et al.
Published: (2025)
by: Lin, Junyan, et al.
Published: (2025)
Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism
by: Zhao, Anhao, et al.
Published: (2024)
by: Zhao, Anhao, et al.
Published: (2024)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
by: Zhang, Jialiang, et al.
Published: (2026)
by: Zhang, Jialiang, et al.
Published: (2026)
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs
by: Gao, Xinyu, et al.
Published: (2026)
by: Gao, Xinyu, et al.
Published: (2026)
Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning
by: Chen, Yanjun, et al.
Published: (2025)
by: Chen, Yanjun, et al.
Published: (2025)
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM
by: Wang, Yujun, et al.
Published: (2025)
by: Wang, Yujun, et al.
Published: (2025)
CA-YOLO: Cross Attention Empowered YOLO for Biomimetic Localization
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
by: Gedik, Hakan Emre, et al.
Published: (2025)
by: Gedik, Hakan Emre, et al.
Published: (2025)
MedSAM-CA: A CNN-Augmented ViT with Attention-Enhanced Multi-Scale Fusion for Medical Image Segmentation
by: Tian, Peiting, et al.
Published: (2025)
by: Tian, Peiting, et al.
Published: (2025)
Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
by: Chen, Xinghao, et al.
Published: (2025)
by: Chen, Xinghao, et al.
Published: (2025)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
by: Kopiczko, Dawid J., et al.
Published: (2024)
by: Kopiczko, Dawid J., et al.
Published: (2024)
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
by: Xie, Qi, et al.
Published: (2025)
by: Xie, Qi, et al.
Published: (2025)
FairViT: Fair Vision Transformer via Adaptive Masking
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
ViSymRe: Vision Multimodal Symbolic Regression
by: Li, Da, et al.
Published: (2024)
by: Li, Da, et al.
Published: (2024)
Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics
by: Wang, Jiashuo, et al.
Published: (2026)
by: Wang, Jiashuo, et al.
Published: (2026)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
by: Chen, Boqi, et al.
Published: (2026)
by: Chen, Boqi, et al.
Published: (2026)
You Only Need Less Attention at Each Stage in Vision Transformers
by: Zhang, Shuoxi, et al.
Published: (2024)
by: Zhang, Shuoxi, et al.
Published: (2024)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
StochCA: A Novel Approach for Exploiting Pretrained Models with Cross-Attention
by: Seo, Seungwon, et al.
Published: (2024)
by: Seo, Seungwon, et al.
Published: (2024)
Similar Items
-
StreamingThinker: Large Language Models Can Think While Reading
by: Tong, Junlong, et al.
Published: (2025) -
$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
by: Fan, Yingqi, et al.
Published: (2025) -
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
by: Fan, Yingqi, et al.
Published: (2026) -
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
by: Zhao, Anhao, et al.
Published: (2026) -
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
by: Wu, Hao, et al.
Published: (2026)