Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Boyu, Chen, Siran, Li, Kunchang, Xu, Qinglin, Qiao, Yu, Wang, Yali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
by: Chen, Boyu, et al.
Published: (2024)
by: Chen, Boyu, et al.
Published: (2024)
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024)
by: Li, Kunchang, et al.
Published: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
by: Ding, Yanbo, et al.
Published: (2024)
by: Ding, Yanbo, et al.
Published: (2024)
Low-Resolution Action Recognition for Tiny Actions Challenge
by: Chen, Boyu, et al.
Published: (2022)
by: Chen, Boyu, et al.
Published: (2022)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
by: Guo, Yiwei, et al.
Published: (2024)
by: Guo, Yiwei, et al.
Published: (2024)
Harvest Video Foundation Models via Efficient Post-Pretraining
by: Li, Yizhuo, et al.
Published: (2023)
by: Li, Yizhuo, et al.
Published: (2023)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
by: Wang, Zikang, et al.
Published: (2025)
by: Wang, Zikang, et al.
Published: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
by: Yue, Zhengrong, et al.
Published: (2025)
by: Yue, Zhengrong, et al.
Published: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
by: Zhang, Hongjie, et al.
Published: (2023)
by: Zhang, Hongjie, et al.
Published: (2023)
VideoChat: Chat-Centric Video Understanding
by: Li, KunChang, et al.
Published: (2023)
by: Li, KunChang, et al.
Published: (2023)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
by: Zeng, Xiangyu, et al.
Published: (2024)
by: Zeng, Xiangyu, et al.
Published: (2024)
Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration
by: Liao, Chih-Ting, et al.
Published: (2025)
by: Liao, Chih-Ting, et al.
Published: (2025)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Vlogger: Make Your Dream A Vlog
by: Zhuang, Shaobin, et al.
Published: (2024)
by: Zhuang, Shaobin, et al.
Published: (2024)
Recursive Generalization Transformer for Image Super-Resolution
by: Chen, Zheng, et al.
Published: (2023)
by: Chen, Zheng, et al.
Published: (2023)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
by: Li, Handong, et al.
Published: (2025)
by: Li, Handong, et al.
Published: (2025)
TimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in Vision
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation
by: Zhang, Yuxuan, et al.
Published: (2023)
by: Zhang, Yuxuan, et al.
Published: (2023)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
by: Yan, Ziang, et al.
Published: (2024)
by: Yan, Ziang, et al.
Published: (2024)
Make Your Training Flexible: Towards Deployment-Efficient Video Models
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
by: Bai, Purui, et al.
Published: (2026)
by: Bai, Purui, et al.
Published: (2026)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex Degradations
by: Li, Xiaohui, et al.
Published: (2025)
by: Li, Xiaohui, et al.
Published: (2025)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Rethinking Diffusion Model-Based Video Super-Resolution: Leveraging Dense Guidance from Aligned Features
by: Xu, Jingyi, et al.
Published: (2025)
by: Xu, Jingyi, et al.
Published: (2025)
Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities
by: Song, Peibo, et al.
Published: (2026)
by: Song, Peibo, et al.
Published: (2026)
ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
by: Wu, Xiaoxue, et al.
Published: (2025)
by: Wu, Xiaoxue, et al.
Published: (2025)
Similar Items
-
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
by: Chen, Boyu, et al.
Published: (2024) -
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
by: Chen, Siran, et al.
Published: (2025) -
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025) -
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024) -
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)