DeepSVU: Towards In-depth Security-oriented Video Understanding via Unified Physical-world Regularized MoE
Fuente:
arXiv
Salvato in:
| Autori principali: | Jin, Yujie, Zhang, Wenxin, Wang, Jingjing, Zhou, Guodong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
di: Luo, Jiamin, et al.
Pubblicazione: (2025)
di: Luo, Jiamin, et al.
Pubblicazione: (2025)
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
di: Shi, Yangming, et al.
Pubblicazione: (2026)
di: Shi, Yangming, et al.
Pubblicazione: (2026)
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
di: Zhao, Xinyuan, et al.
Pubblicazione: (2026)
di: Zhao, Xinyuan, et al.
Pubblicazione: (2026)
Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
di: Zheng, Youwei, et al.
Pubblicazione: (2025)
di: Zheng, Youwei, et al.
Pubblicazione: (2025)
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
di: Ma, Yingjie, et al.
Pubblicazione: (2024)
di: Ma, Yingjie, et al.
Pubblicazione: (2024)
MiM-DiT: MoE in MoE with Diffusion Transformers for All-in-One Image Restoration
di: Kong, Lingshun, et al.
Pubblicazione: (2026)
di: Kong, Lingshun, et al.
Pubblicazione: (2026)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
di: Jin, In-Hwan, et al.
Pubblicazione: (2025)
di: Jin, In-Hwan, et al.
Pubblicazione: (2025)
Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLM
di: Ma, Junxiao, et al.
Pubblicazione: (2025)
di: Ma, Junxiao, et al.
Pubblicazione: (2025)
Towards Event-oriented Long Video Understanding
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
di: Liu, Xiangyue, et al.
Pubblicazione: (2026)
di: Liu, Xiangyue, et al.
Pubblicazione: (2026)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
di: Xu, Yu, et al.
Pubblicazione: (2026)
di: Xu, Yu, et al.
Pubblicazione: (2026)
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
di: Wei, Yujie, et al.
Pubblicazione: (2025)
di: Wei, Yujie, et al.
Pubblicazione: (2025)
UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
di: Zhao, Yuan, et al.
Pubblicazione: (2025)
di: Zhao, Yuan, et al.
Pubblicazione: (2025)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
di: Li, Yu, et al.
Pubblicazione: (2025)
di: Li, Yu, et al.
Pubblicazione: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach
di: Chen, Yunuo, et al.
Pubblicazione: (2025)
di: Chen, Yunuo, et al.
Pubblicazione: (2025)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
Nucleus-Image: Sparse MoE for Image Generation
di: Akiti, Chandan, et al.
Pubblicazione: (2026)
di: Akiti, Chandan, et al.
Pubblicazione: (2026)
Towards Accurate Unified Anomaly Segmentation
di: Ma, Wenxin, et al.
Pubblicazione: (2025)
di: Ma, Wenxin, et al.
Pubblicazione: (2025)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE
di: Bao, Wei, et al.
Pubblicazione: (2026)
di: Bao, Wei, et al.
Pubblicazione: (2026)
MReg: A Novel Regression Model with MoE-based Video Feature Mining for Mitral Regurgitation Diagnosis
di: Liu, Zhe, et al.
Pubblicazione: (2025)
di: Liu, Zhe, et al.
Pubblicazione: (2025)
SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation
di: Wang, Guoan, et al.
Pubblicazione: (2024)
di: Wang, Guoan, et al.
Pubblicazione: (2024)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
di: Huang, Jincai, et al.
Pubblicazione: (2026)
di: Huang, Jincai, et al.
Pubblicazione: (2026)
LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
di: Shang, Yu, et al.
Pubblicazione: (2025)
di: Shang, Yu, et al.
Pubblicazione: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
di: Chen, Zeren, et al.
Pubblicazione: (2023)
di: Chen, Zeren, et al.
Pubblicazione: (2023)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
di: Wu, Yinghao, et al.
Pubblicazione: (2026)
di: Wu, Yinghao, et al.
Pubblicazione: (2026)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
di: Min, Chengxi, et al.
Pubblicazione: (2025)
di: Min, Chengxi, et al.
Pubblicazione: (2025)
eMoE-Tracker: Environmental MoE-based Transformer for Robust Event-guided Object Tracking
di: Chen, Yucheng, et al.
Pubblicazione: (2024)
di: Chen, Yucheng, et al.
Pubblicazione: (2024)
UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
di: Pan, Hewen, et al.
Pubblicazione: (2025)
di: Pan, Hewen, et al.
Pubblicazione: (2025)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
di: Li, Yunxin, et al.
Pubblicazione: (2025)
di: Li, Yunxin, et al.
Pubblicazione: (2025)
Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models
di: Wang, Peiran, et al.
Pubblicazione: (2025)
di: Wang, Peiran, et al.
Pubblicazione: (2025)
Semi-MoE: Mixture-of-Experts meets Semi-Supervised Histopathology Segmentation
di: Vu, Nguyen Lan Vi, et al.
Pubblicazione: (2025)
di: Vu, Nguyen Lan Vi, et al.
Pubblicazione: (2025)
GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts
di: Liao, Minwen, et al.
Pubblicazione: (2025)
di: Liao, Minwen, et al.
Pubblicazione: (2025)
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
di: Bo, Zi-Hao, et al.
Pubblicazione: (2026)
di: Bo, Zi-Hao, et al.
Pubblicazione: (2026)
Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
di: Chen, Xuweiyi, et al.
Pubblicazione: (2025)
di: Chen, Xuweiyi, et al.
Pubblicazione: (2025)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
di: Jiang, Songtao, et al.
Pubblicazione: (2024)
di: Jiang, Songtao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
di: Luo, Jiamin, et al.
Pubblicazione: (2025) -
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
di: Shi, Yangming, et al.
Pubblicazione: (2026) -
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
di: Zhao, Xinyuan, et al.
Pubblicazione: (2026) -
Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
di: Zheng, Youwei, et al.
Pubblicazione: (2025) -
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
di: Ma, Yingjie, et al.
Pubblicazione: (2024)