MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Minghao, Wang, Zhengpu, Hu, Mengxian, Dang, Ronghao, Lin, Xiao, Zhou, Xun, Liu, Chengju, Chen, Qijun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
by: Zhu, Minghao, et al.
Published: (2023)
by: Zhu, Minghao, et al.
Published: (2023)
CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
Efficient Text-driven Motion Generation via Latent Consistency Training
by: Hu, Mengxian, et al.
Published: (2024)
by: Hu, Mengxian, et al.
Published: (2024)
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2024)
by: Wang, Liuyi, et al.
Published: (2024)
Vision-and-Language Navigation via Causal Learning
by: Wang, Liuyi, et al.
Published: (2024)
by: Wang, Liuyi, et al.
Published: (2024)
A Dual Semantic-Aware Recurrent Global-Adaptive Network For Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2023)
by: Wang, Liuyi, et al.
Published: (2023)
CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation
by: Lin, Xiao, et al.
Published: (2025)
by: Lin, Xiao, et al.
Published: (2025)
TransPose: 6D Object Pose Estimation with Geometry-Aware Transformer
by: Lin, Xiao, et al.
Published: (2023)
by: Lin, Xiao, et al.
Published: (2023)
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
by: Dang, Ronghao, et al.
Published: (2023)
by: Dang, Ronghao, et al.
Published: (2023)
Realizing Text-Driven Motion Generation on NAO Robot: A Reinforcement Learning-Optimized Control Pipeline
by: Xu, Zihan, et al.
Published: (2025)
by: Xu, Zihan, et al.
Published: (2025)
Beyond instruction-conditioning, MoTE: Mixture of Task Experts for Multi-task Embedding Models
by: Romero, Miguel, et al.
Published: (2025)
by: Romero, Miguel, et al.
Published: (2025)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
MoTE: Mixture of Task-specific Experts for Pre-Trained ModelBased Class-incremental Learning
by: Li, Linjie, et al.
Published: (2025)
by: Li, Linjie, et al.
Published: (2025)
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
by: He, Zongtao, et al.
Published: (2025)
by: He, Zongtao, et al.
Published: (2025)
SAM-LAD: Segment Anything Model Meets Zero-Shot Logic Anomaly Detection
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2024)
by: Wang, Liuyi, et al.
Published: (2024)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
by: He, Zongtao, et al.
Published: (2023)
by: He, Zongtao, et al.
Published: (2023)
DTL: Disentangled Transfer Learning for Visual Recognition
by: Fu, Minghao, et al.
Published: (2023)
by: Fu, Minghao, et al.
Published: (2023)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
by: Fan, Rong, et al.
Published: (2026)
by: Fan, Rong, et al.
Published: (2026)
PASTS: Progress-Aware Spatio-Temporal Transformer Speaker For Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2023)
by: Wang, Liuyi, et al.
Published: (2023)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
Artificial Intelligence for Biomedical Video Generation
by: Li, Linyuan, et al.
Published: (2024)
by: Li, Linyuan, et al.
Published: (2024)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
CLASH: Collaborative Large-Small Hierarchical Framework for Continuous Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation
by: Lu, Yujie, et al.
Published: (2026)
by: Lu, Yujie, et al.
Published: (2026)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
by: Sheng, Kai, et al.
Published: (2026)
by: Sheng, Kai, et al.
Published: (2026)
Are Video Generation Models Geographically Fair? An Attraction-Centric Evaluation of Global Visual Knowledge
by: Liu, Xiao, et al.
Published: (2026)
by: Liu, Xiao, et al.
Published: (2026)
RoadFormer+: Delivering RGB-X Scene Parsing through Scale-Aware Information Decoupling and Advanced Heterogeneous Feature Fusion
by: Huang, Jianxin, et al.
Published: (2024)
by: Huang, Jianxin, et al.
Published: (2024)
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
by: Xie, Qi, et al.
Published: (2025)
by: Xie, Qi, et al.
Published: (2025)
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
by: Ma, Yingjie, et al.
Published: (2024)
by: Ma, Yingjie, et al.
Published: (2024)
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
by: Lin, Ronghao, et al.
Published: (2025)
by: Lin, Ronghao, et al.
Published: (2025)
MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
by: Zhu, Chenyang, et al.
Published: (2026)
by: Zhu, Chenyang, et al.
Published: (2026)
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Diffusion as Reasoning: Enhancing Object Navigation via Diffusion Model Conditioned on LLM-based Object-Room Knowledge
by: Ji, Yiming, et al.
Published: (2024)
by: Ji, Yiming, et al.
Published: (2024)
Visual Large Language Models for Generalized and Specialized Applications
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Knowledge-enhanced Visual-Language Pretraining for Computational Pathology
by: Zhou, Xiao, et al.
Published: (2024)
by: Zhou, Xiao, et al.
Published: (2024)
Similar Items
-
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
by: Zhu, Minghao, et al.
Published: (2023) -
CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
by: Lin, Xiao, et al.
Published: (2024) -
Efficient Text-driven Motion Generation via Latent Consistency Training
by: Hu, Mengxian, et al.
Published: (2024) -
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
by: Wang, Hongyu, et al.
Published: (2025) -
Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2024)