RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer
Fuente:
arXiv
Salvato in:
| Autori principali: | Ni, Haotian, Wei, Yake, Liu, Hang, Chen, Gong, Peng, Chong, Lin, Hao, Hu, Di |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
di: Wei, Yake, et al.
Pubblicazione: (2024)
di: Wei, Yake, et al.
Pubblicazione: (2024)
MIBench: Evaluating LMMs on Multimodal Interaction
di: Miao, Yu, et al.
Pubblicazione: (2026)
di: Miao, Yu, et al.
Pubblicazione: (2026)
MokA: Multimodal Low-Rank Adaptation for MLLMs
di: Wei, Yake, et al.
Pubblicazione: (2025)
di: Wei, Yake, et al.
Pubblicazione: (2025)
Diagnosing and Re-learning for Balanced Multimodal Learning
di: Wei, Yake, et al.
Pubblicazione: (2024)
di: Wei, Yake, et al.
Pubblicazione: (2024)
Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception
di: Peng, Ruotian, et al.
Pubblicazione: (2025)
di: Peng, Ruotian, et al.
Pubblicazione: (2025)
On-the-fly Modulation for Balanced Multimodal Learning
di: Wei, Yake, et al.
Pubblicazione: (2024)
di: Wei, Yake, et al.
Pubblicazione: (2024)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
di: Yang, Zequn, et al.
Pubblicazione: (2024)
di: Yang, Zequn, et al.
Pubblicazione: (2024)
Enhancing multimodal cooperation via sample-level modality valuation
di: Wei, Yake, et al.
Pubblicazione: (2023)
di: Wei, Yake, et al.
Pubblicazione: (2023)
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
di: Cao, Jianjian, et al.
Pubblicazione: (2024)
di: Cao, Jianjian, et al.
Pubblicazione: (2024)
Reviving Undersampling for Long-Tailed Learning
di: Yu, Hao, et al.
Pubblicazione: (2024)
di: Yu, Hao, et al.
Pubblicazione: (2024)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
di: Yang, Tiancheng, et al.
Pubblicazione: (2025)
di: Yang, Tiancheng, et al.
Pubblicazione: (2025)
Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs
di: Pahuja, Vardaan, et al.
Pubblicazione: (2023)
di: Pahuja, Vardaan, et al.
Pubblicazione: (2023)
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval
di: Yang, Wenjie, et al.
Pubblicazione: (2026)
di: Yang, Wenjie, et al.
Pubblicazione: (2026)
CATNet: Collaborative Alignment and Transformation Network for Cooperative Perception
di: Chen, Gong, et al.
Pubblicazione: (2026)
di: Chen, Gong, et al.
Pubblicazione: (2026)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
di: Peng, Kunyu, et al.
Pubblicazione: (2026)
di: Peng, Kunyu, et al.
Pubblicazione: (2026)
InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
PRevivor: Reviving Ancient Chinese Paintings using Prior-Guided Color Transformers
di: Tang, Tan, et al.
Pubblicazione: (2025)
di: Tang, Tan, et al.
Pubblicazione: (2025)
Single Image Rolling Shutter Removal with Diffusion Models
di: Yang, Zhanglei, et al.
Pubblicazione: (2024)
di: Yang, Zhanglei, et al.
Pubblicazione: (2024)
DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
di: Zhuang, Jiedong, et al.
Pubblicazione: (2026)
di: Zhuang, Jiedong, et al.
Pubblicazione: (2026)
In Defense and Revival of Bayesian Filtering for Thermal Infrared Object Tracking
di: Gao, Peng, et al.
Pubblicazione: (2024)
di: Gao, Peng, et al.
Pubblicazione: (2024)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
di: Sun, Hao, et al.
Pubblicazione: (2024)
di: Sun, Hao, et al.
Pubblicazione: (2024)
DPNet: Dynamic Pooling Network for Tiny Object Detection
di: Gong, Luqi, et al.
Pubblicazione: (2025)
di: Gong, Luqi, et al.
Pubblicazione: (2025)
DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
di: Zhao, Yuzhong, et al.
Pubblicazione: (2024)
di: Zhao, Yuzhong, et al.
Pubblicazione: (2024)
Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder
di: Peng, Chong, et al.
Pubblicazione: (2024)
di: Peng, Chong, et al.
Pubblicazione: (2024)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
di: Liu, Kunhao, et al.
Pubblicazione: (2025)
di: Liu, Kunhao, et al.
Pubblicazione: (2025)
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
di: Hu, Yu, et al.
Pubblicazione: (2025)
di: Hu, Yu, et al.
Pubblicazione: (2025)
A Spatial-Spectral-Frequency Interactive Network for Multimodal Remote Sensing Classification
di: Liu, Hao, et al.
Pubblicazione: (2025)
di: Liu, Hao, et al.
Pubblicazione: (2025)
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
di: Gou, Yunhao, et al.
Pubblicazione: (2024)
di: Gou, Yunhao, et al.
Pubblicazione: (2024)
Cohort-Individual Cooperative Learning for Multimodal Cancer Survival Analysis
di: Zhou, Huajun, et al.
Pubblicazione: (2024)
di: Zhou, Huajun, et al.
Pubblicazione: (2024)
Unbiased Dynamic Multimodal Fusion
di: Wei, Shicai, et al.
Pubblicazione: (2026)
di: Wei, Shicai, et al.
Pubblicazione: (2026)
Cross-Modal RGB-D Fusion Transformer for 6D Pose Estimation of Non-Cooperative Spacecraft with Stereo-Derived Depth
di: Zhen, Yongliang, et al.
Pubblicazione: (2026)
di: Zhen, Yongliang, et al.
Pubblicazione: (2026)
Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
di: Dong, Haotian, et al.
Pubblicazione: (2025)
di: Dong, Haotian, et al.
Pubblicazione: (2025)
MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
di: Sun, Shengyang, et al.
Pubblicazione: (2024)
di: Sun, Shengyang, et al.
Pubblicazione: (2024)
Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks
di: Cai, Yuliang, et al.
Pubblicazione: (2024)
di: Cai, Yuliang, et al.
Pubblicazione: (2024)
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
di: Chen, Cheng, et al.
Pubblicazione: (2026)
di: Chen, Cheng, et al.
Pubblicazione: (2026)
RSL-BA: Rolling Shutter Line Bundle Adjustment
di: Zhang, Yongcong, et al.
Pubblicazione: (2024)
di: Zhang, Yongcong, et al.
Pubblicazione: (2024)
CoopDiff: A Diffusion-Guided Approach for Cooperation under Corruptions
di: Chen, Gong, et al.
Pubblicazione: (2026)
di: Chen, Gong, et al.
Pubblicazione: (2026)
CompetitorFormer: Competitor Transformer for 3D Instance Segmentation
di: Wang, Duanchu, et al.
Pubblicazione: (2024)
di: Wang, Duanchu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
di: Wei, Yake, et al.
Pubblicazione: (2024) -
MIBench: Evaluating LMMs on Multimodal Interaction
di: Miao, Yu, et al.
Pubblicazione: (2026) -
MokA: Multimodal Low-Rank Adaptation for MLLMs
di: Wei, Yake, et al.
Pubblicazione: (2025) -
Diagnosing and Re-learning for Balanced Multimodal Learning
di: Wei, Yake, et al.
Pubblicazione: (2024) -
Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception
di: Peng, Ruotian, et al.
Pubblicazione: (2025)