Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jain, Gagan, Hegde, Nidhi, Kusupati, Aditya, Nagrani, Arsha, Buch, Shyamal, Jain, Prateek, Arnab, Anurag, Paul, Sujoy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Streaming Dense Video Captioning
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
ELT: Elastic Looped Transformers for Visual Generation
von: Goyal, Sahil, et al.
Veröffentlicht: (2026)
von: Goyal, Sahil, et al.
Veröffentlicht: (2026)
Masked Generative Nested Transformers with Decode Time Scaling
von: Goyal, Sahil, et al.
Veröffentlicht: (2025)
von: Goyal, Sahil, et al.
Veröffentlicht: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
von: Min, Juhong, et al.
Veröffentlicht: (2024)
von: Min, Juhong, et al.
Veröffentlicht: (2024)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
LookupViT: Compressing visual information to a limited number of tokens
von: Koner, Rajat, et al.
Veröffentlicht: (2024)
von: Koner, Rajat, et al.
Veröffentlicht: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
von: Nagrani, Arsha, et al.
Veröffentlicht: (2026)
von: Nagrani, Arsha, et al.
Veröffentlicht: (2026)
VoCap: Video Object Captioning and Segmentation from Any Prompt
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
MatFormer: Nested Transformer for Elastic Inference
von: Devvrit, et al.
Veröffentlicht: (2023)
von: Devvrit, et al.
Veröffentlicht: (2023)
MINERVA: Evaluating Complex Video Reasoning
von: Nagrani, Arsha, et al.
Veröffentlicht: (2025)
von: Nagrani, Arsha, et al.
Veröffentlicht: (2025)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
AutoAD III: The Prequel -- Back to the Pixels
von: Han, Tengda, et al.
Veröffentlicht: (2024)
von: Han, Tengda, et al.
Veröffentlicht: (2024)
GazeLT: Visual attention-guided long-tailed disease classification in chest radiographs
von: Bhattacharya, Moinak, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Moinak, et al.
Veröffentlicht: (2025)
Matryoshka Representation Learning
von: Kusupati, Aditya, et al.
Veröffentlicht: (2022)
von: Kusupati, Aditya, et al.
Veröffentlicht: (2022)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
von: Xie, Junyu, et al.
Veröffentlicht: (2024)
von: Xie, Junyu, et al.
Veröffentlicht: (2024)
CAViAR: Critic-Augmented Video Agentic Reasoning
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
Principles of Visual Tokens for Efficient Video Understanding
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
von: Xie, Junyu, et al.
Veröffentlicht: (2025)
von: Xie, Junyu, et al.
Veröffentlicht: (2025)
FaceCloak: Learning to Protect Face Templates
von: Banerjee, Sudipta, et al.
Veröffentlicht: (2025)
von: Banerjee, Sudipta, et al.
Veröffentlicht: (2025)
RadGazeGen: Radiomics and Gaze-guided Medical Image Generation using Diffusion Models
von: Bhattacharya, Moinak, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Moinak, et al.
Veröffentlicht: (2024)
Survival Modeling from Whole Slide Images via Patch-Level Graph Clustering and Mixture Density Experts
von: Sekhar, Ardhendu, et al.
Veröffentlicht: (2025)
von: Sekhar, Ardhendu, et al.
Veröffentlicht: (2025)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
von: Singh, Darshan, et al.
Veröffentlicht: (2026)
von: Singh, Darshan, et al.
Veröffentlicht: (2026)
Streaming Detection of Queried Event Start
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2024)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2024)
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
von: Liu, Tianyi, et al.
Veröffentlicht: (2026)
von: Liu, Tianyi, et al.
Veröffentlicht: (2026)
Frequency-Adaptive Pan-Sharpening with Mixture of Experts
von: He, Xuanhua, et al.
Veröffentlicht: (2024)
von: He, Xuanhua, et al.
Veröffentlicht: (2024)
Dynamic Mixture-of-Experts for Visual Autoregressive Model
von: Vincenti, Jort, et al.
Veröffentlicht: (2025)
von: Vincenti, Jort, et al.
Veröffentlicht: (2025)
Unsupervised Video Highlight Detection by Learning from Audio and Visual Recurrence
von: Islam, Zahidul, et al.
Veröffentlicht: (2024)
von: Islam, Zahidul, et al.
Veröffentlicht: (2024)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
YOLO Meets Mixture-of-Experts: Adaptive Expert Routing for Robust Object Detection
von: Meiraz, Ori, et al.
Veröffentlicht: (2025)
von: Meiraz, Ori, et al.
Veröffentlicht: (2025)
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
von: Hong, Lingyi, et al.
Veröffentlicht: (2026)
von: Hong, Lingyi, et al.
Veröffentlicht: (2026)
SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation
von: Lu, Yujie, et al.
Veröffentlicht: (2026)
von: Lu, Yujie, et al.
Veröffentlicht: (2026)
Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
von: Yang, Longrong, et al.
Veröffentlicht: (2024)
von: Yang, Longrong, et al.
Veröffentlicht: (2024)
MatMamba: A Matryoshka State Space Model
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
Triad: Empowering LMM-based Anomaly Detection with Vision Expert-guided Visual Tokenizer and Manufacturing Process
von: Li, Yuanze, et al.
Veröffentlicht: (2025)
von: Li, Yuanze, et al.
Veröffentlicht: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Streaming Dense Video Captioning
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024) -
ELT: Elastic Looped Transformers for Visual Generation
von: Goyal, Sahil, et al.
Veröffentlicht: (2026) -
Masked Generative Nested Transformers with Decode Time Scaling
von: Goyal, Sahil, et al.
Veröffentlicht: (2025) -
VicTR: Video-conditioned Text Representations for Activity Recognition
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023) -
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
von: Min, Juhong, et al.
Veröffentlicht: (2024)