Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Tianyi, Li, Yiming, Wang, Wenqian, Wang, Jiaojiao, Cai, Chen, Wang, Yi, Yap, Kim-Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset
von: Li, Yiming, et al.
Veröffentlicht: (2026)
von: Li, Yiming, et al.
Veröffentlicht: (2026)
Open World Object Detection: A Survey
von: Li, Yiming, et al.
Veröffentlicht: (2024)
von: Li, Yiming, et al.
Veröffentlicht: (2024)
Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification
von: Gao, Qinghao, et al.
Veröffentlicht: (2025)
von: Gao, Qinghao, et al.
Veröffentlicht: (2025)
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment
von: Xu, Huangbiao, et al.
Veröffentlicht: (2025)
von: Xu, Huangbiao, et al.
Veröffentlicht: (2025)
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
von: Liu, Wenyang, et al.
Veröffentlicht: (2024)
von: Liu, Wenyang, et al.
Veröffentlicht: (2024)
Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained Experts
von: Xu, Yangyang, et al.
Veröffentlicht: (2025)
von: Xu, Yangyang, et al.
Veröffentlicht: (2025)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
von: Jain, Gagan, et al.
Veröffentlicht: (2024)
von: Jain, Gagan, et al.
Veröffentlicht: (2024)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
Fine-Grained Generalization via Structuralizing Concept and Feature Space into Commonality, Specificity and Confounding
von: Wang, Zhen, et al.
Veröffentlicht: (2026)
von: Wang, Zhen, et al.
Veröffentlicht: (2026)
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
von: Hong, Lingyi, et al.
Veröffentlicht: (2026)
von: Hong, Lingyi, et al.
Veröffentlicht: (2026)
On Learning Discriminative Features from Synthesized Data for Self-Supervised Fine-Grained Visual Recognition
von: Wang, Zihu, et al.
Veröffentlicht: (2024)
von: Wang, Zihu, et al.
Veröffentlicht: (2024)
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action Recognition
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
von: Kuchibhotla, Hari Chandana, et al.
Veröffentlicht: (2025)
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition
von: Cho, Seungyeon, et al.
Veröffentlicht: (2025)
von: Cho, Seungyeon, et al.
Veröffentlicht: (2025)
Token Activation Map to Visually Explain Multimodal LLMs
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking
von: Zhu, Yabin, et al.
Veröffentlicht: (2026)
von: Zhu, Yabin, et al.
Veröffentlicht: (2026)
Top-Down Framework for Weakly-supervised Grounded Image Captioning
von: Cai, Chen, et al.
Veröffentlicht: (2023)
von: Cai, Chen, et al.
Veröffentlicht: (2023)
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
von: Sun, Baoli, et al.
Veröffentlicht: (2025)
von: Sun, Baoli, et al.
Veröffentlicht: (2025)
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
SEMC: Structure-Enhanced Mixture-of-Experts Contrastive Learning for Ultrasound Standard Plane Recognition
von: Cai, Qing, et al.
Veröffentlicht: (2025)
von: Cai, Qing, et al.
Veröffentlicht: (2025)
Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
von: Ge, Chendi, et al.
Veröffentlicht: (2025)
von: Ge, Chendi, et al.
Veröffentlicht: (2025)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
Efficient Mixture-of-Expert for Video-based Driver State and Physiological Multi-task Estimation in Conditional Autonomous Driving
von: Wang, Jiyao, et al.
Veröffentlicht: (2024)
von: Wang, Jiyao, et al.
Veröffentlicht: (2024)
Neural Dynamics Model of Visual Decision-Making: Learning from Human Experts
von: Su, Jie, et al.
Veröffentlicht: (2024)
von: Su, Jie, et al.
Veröffentlicht: (2024)
Hierarchically Decoupled Mixture-of-Experts for Robust Traffic Sign Recognition in Complex Driving Scenarios
von: Wang, Mingxiao, et al.
Veröffentlicht: (2026)
von: Wang, Mingxiao, et al.
Veröffentlicht: (2026)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
von: Cai, Kaitong, et al.
Veröffentlicht: (2025)
Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
von: Zeng, Zhen, et al.
Veröffentlicht: (2024)
Novel Class Discovery for Ultra-Fine-Grained Visual Categorization
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Free-Grained Hierarchical Visual Recognition
von: Park, Seulki, et al.
Veröffentlicht: (2025)
von: Park, Seulki, et al.
Veröffentlicht: (2025)
BHaRNet: Reliability-Aware Body-Hand Modality Expertized Networks for Fine-grained Skeleton Action Recognition
von: Cho, Seungyeon, et al.
Veröffentlicht: (2026)
von: Cho, Seungyeon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024) -
CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024) -
DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset
von: Li, Yiming, et al.
Veröffentlicht: (2026) -
Open World Object Detection: A Survey
von: Li, Yiming, et al.
Veröffentlicht: (2024) -
Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification
von: Gao, Qinghao, et al.
Veröffentlicht: (2025)