CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Shilin, Han, Jiaming, Tsai, Joey, Xue, Hongwei, Fang, Rongyao, Hong, Lingyi, Guo, Ziyu, Zhang, Ray |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
Beyond Scattered Acceptance: Fast and Coherent Inference for DLMs via Longest Stable Prefixes
by: Li, Pengxiang, et al.
Published: (2026)
by: Li, Pengxiang, et al.
Published: (2026)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
WiFi-based Cross-Domain Gesture Recognition Using Attention Mechanism
by: Liu, Ruijing, et al.
Published: (2025)
by: Liu, Ruijing, et al.
Published: (2025)
LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue
by: Li, Chaoyue, et al.
Published: (2026)
by: Li, Chaoyue, et al.
Published: (2026)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
by: Guo, Zhenyu, et al.
Published: (2025)
by: Guo, Zhenyu, et al.
Published: (2025)
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
by: Peng, Yingzhe, et al.
Published: (2025)
by: Peng, Yingzhe, et al.
Published: (2025)
SCASeg: Strip Cross-Attention for Efficient Semantic Segmentation
by: Xu, Guoan, et al.
Published: (2024)
by: Xu, Guoan, et al.
Published: (2024)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
by: Fei, Jiajun, et al.
Published: (2024)
by: Fei, Jiajun, et al.
Published: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
by: He, Bo, et al.
Published: (2024)
by: He, Bo, et al.
Published: (2024)
PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation
by: Yan, Shilin, et al.
Published: (2023)
by: Yan, Shilin, et al.
Published: (2023)
S$^3$Attention: Improving Long Sequence Attention with Smoothed Skeleton Sketching
by: Wang, Xue, et al.
Published: (2024)
by: Wang, Xue, et al.
Published: (2024)
Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
by: Ghadiya, Ayush, et al.
Published: (2024)
by: Ghadiya, Ayush, et al.
Published: (2024)
Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation
by: Dai, Ji, et al.
Published: (2026)
by: Dai, Ji, et al.
Published: (2026)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Dual Cross-Attention for Medical Image Segmentation
by: Ates, Gorkem Can, et al.
Published: (2023)
by: Ates, Gorkem Can, et al.
Published: (2023)
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs
by: Yamao, Sosuke, et al.
Published: (2024)
by: Yamao, Sosuke, et al.
Published: (2024)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
by: Kim, Ji-Hoon, et al.
Published: (2024)
by: Kim, Ji-Hoon, et al.
Published: (2024)
Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers
by: Hiller, Markus, et al.
Published: (2024)
by: Hiller, Markus, et al.
Published: (2024)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
by: Li, Ruibin, et al.
Published: (2026)
by: Li, Ruibin, et al.
Published: (2026)
CaDA: Cross-Problem Routing Solver with Constraint-Aware Dual-Attention
by: Li, Han, et al.
Published: (2024)
by: Li, Han, et al.
Published: (2024)
DMWM: Dual-Mind World Model with Long-Term Imagination
by: Wang, Lingyi, et al.
Published: (2025)
by: Wang, Lingyi, et al.
Published: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
by: Qi, Ji, et al.
Published: (2025)
by: Qi, Ji, et al.
Published: (2025)
Exploring the Potential of Encoder-free Architectures in 3D LMMs
by: Tang, Yiwen, et al.
Published: (2025)
by: Tang, Yiwen, et al.
Published: (2025)
AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
by: Li, Jieyu, et al.
Published: (2025)
by: Li, Jieyu, et al.
Published: (2025)
LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
by: Xia, Shuhan, et al.
Published: (2025)
by: Xia, Shuhan, et al.
Published: (2025)
Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrieval
by: Yang, Jing, et al.
Published: (2026)
by: Yang, Jing, et al.
Published: (2026)
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
by: Jiao, Pengkun, et al.
Published: (2025)
by: Jiao, Pengkun, et al.
Published: (2025)
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
D-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference
by: Daher, Leen, et al.
Published: (2025)
by: Daher, Leen, et al.
Published: (2025)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024)
by: Bharadwaj, Rohit, et al.
Published: (2024)
Provable Differentially Private Computation of the Cross-Attention Mechanism
by: Ke, Yekun, et al.
Published: (2024)
by: Ke, Yekun, et al.
Published: (2024)
Similar Items
-
LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
by: Wang, Jiarui, et al.
Published: (2025) -
Beyond Scattered Acceptance: Fast and Coherent Inference for DLMs via Longest Stable Prefixes
by: Li, Pengxiang, et al.
Published: (2026) -
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
by: Khattak, Muhammad Uzair, et al.
Published: (2024) -
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
by: Xu, Zitong, et al.
Published: (2025) -
Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
by: Li, Pengxiang, et al.
Published: (2025)