RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiao, Pengkun, Wu, Xinlan, Zhu, Bin, Chen, Jingjing, Ngo, Chong-Wah, Jiang, Yugang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
von: Jiao, Pengkun, et al.
Veröffentlicht: (2025)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2025)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
von: Yin, Yuehao, et al.
Veröffentlicht: (2023)
Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
von: Qi, Huiyan, et al.
Veröffentlicht: (2025)
von: Qi, Huiyan, et al.
Veröffentlicht: (2025)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Interpretable Embedding for Ad-hoc Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Retrieval Augmented Recipe Generation
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
von: Wu, Xinlan, et al.
Veröffentlicht: (2025)
von: Wu, Xinlan, et al.
Veröffentlicht: (2025)
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking
von: Zhu, Yabin, et al.
Veröffentlicht: (2026)
von: Zhu, Yabin, et al.
Veröffentlicht: (2026)
From Canteen Food to Daily Meals: Generalizing Food Recognition to More Practical Scenarios
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
von: Liu, Guoshan, et al.
Veröffentlicht: (2024)
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts
von: Zhu, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhu, Chenyang, et al.
Veröffentlicht: (2024)
Domain Expansion and Boundary Growth for Open-Set Single-Source Domain Generalization
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Mixture of Style Experts for Diverse Image Stylization
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
von: Sun, Zhen, et al.
Veröffentlicht: (2025)
von: Sun, Zhen, et al.
Veröffentlicht: (2025)
PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
von: Hao, Yanbin, et al.
Veröffentlicht: (2024)
von: Hao, Yanbin, et al.
Veröffentlicht: (2024)
NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification
von: Li, Shihao, et al.
Veröffentlicht: (2025)
von: Li, Shihao, et al.
Veröffentlicht: (2025)
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation
von: Bi, Hanbo, et al.
Veröffentlicht: (2025)
von: Bi, Hanbo, et al.
Veröffentlicht: (2025)
DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
Mr. DETR++: Instructive Multi-Route Training for Detection Transformers with Mixture-of-Experts
von: Zhang, Chang-Bin, et al.
Veröffentlicht: (2024)
von: Zhang, Chang-Bin, et al.
Veröffentlicht: (2024)
MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment
von: Xu, Huangbiao, et al.
Veröffentlicht: (2025)
von: Xu, Huangbiao, et al.
Veröffentlicht: (2025)
LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification
von: Gao, Qinghao, et al.
Veröffentlicht: (2025)
von: Gao, Qinghao, et al.
Veröffentlicht: (2025)
Rectifying Magnitude Neglect in Linear Attention
von: Fan, Qihang, et al.
Veröffentlicht: (2025)
von: Fan, Qihang, et al.
Veröffentlicht: (2025)
Multi-Task Dense Prediction via Mixture of Low-Rank Experts
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts
von: Guo, Hang, et al.
Veröffentlicht: (2023)
von: Guo, Hang, et al.
Veröffentlicht: (2023)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformers
von: Yang, Xiang, et al.
Veröffentlicht: (2026)
von: Yang, Xiang, et al.
Veröffentlicht: (2026)
Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow
von: Gupta, Prerit, et al.
Veröffentlicht: (2025)
von: Gupta, Prerit, et al.
Veröffentlicht: (2025)
MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding
von: Chopra, Shivang, et al.
Veröffentlicht: (2025)
von: Chopra, Shivang, et al.
Veröffentlicht: (2025)
Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance Learning
von: Shao, Daniel, et al.
Veröffentlicht: (2026)
von: Shao, Daniel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
von: Jiao, Pengkun, et al.
Veröffentlicht: (2025) -
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024) -
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
von: Yin, Yuehao, et al.
Veröffentlicht: (2023) -
Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
von: Qi, Huiyan, et al.
Veröffentlicht: (2025) -
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)