Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Tangtartharakul, Gene, Storrs, Katherine R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Language Models show widespread visual deficits on neuropsychological tests
by: Tangtartharakul, Gene, et al.
Published: (2025)
by: Tangtartharakul, Gene, et al.
Published: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
by: Rudman, William, et al.
Published: (2026)
by: Rudman, William, et al.
Published: (2026)
Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection
by: Wang, Gaojian, et al.
Published: (2025)
by: Wang, Gaojian, et al.
Published: (2025)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
by: Khurdula, Harsha Vardhan, et al.
Published: (2024)
by: Khurdula, Harsha Vardhan, et al.
Published: (2024)
Adapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners
by: Narasinghe, N. K. B. M. P. K. B., et al.
Published: (2025)
by: Narasinghe, N. K. B. M. P. K. B., et al.
Published: (2025)
Object-centric proto-symbolic behavioural reasoning from pixels
by: van Bergen, Ruben, et al.
Published: (2024)
by: van Bergen, Ruben, et al.
Published: (2024)
E = T*H/(O+B): A Dimensionless Control Parameter for Mixture-of-Experts Ecology
by: Zhang, Qingjun
Published: (2026)
by: Zhang, Qingjun
Published: (2026)
Statistical Analysis of the Impact of Quaternion Components in Convolutional Neural Networks
by: Altamirano-Gómez, Gerardo, et al.
Published: (2024)
by: Altamirano-Gómez, Gerardo, et al.
Published: (2024)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026)
by: Dong, Mingkang, et al.
Published: (2026)
Sample as You Infer: Predictive Coding With Langevin Dynamics
by: Zahid, Umais, et al.
Published: (2023)
by: Zahid, Umais, et al.
Published: (2023)
Invariant Representation via Decoupling Style and Spurious Features from Images
by: Li, Ruimeng, et al.
Published: (2023)
by: Li, Ruimeng, et al.
Published: (2023)
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
by: Bu, Weijue, et al.
Published: (2025)
by: Bu, Weijue, et al.
Published: (2025)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
by: Wang, Zhaohui, et al.
Published: (2025)
by: Wang, Zhaohui, et al.
Published: (2025)
KGTN-ens: Few-Shot Image Classification with Knowledge Graph Ensembles
by: Filipiak, Dominik, et al.
Published: (2022)
by: Filipiak, Dominik, et al.
Published: (2022)
MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
by: Fan, Qiannan, et al.
Published: (2025)
by: Fan, Qiannan, et al.
Published: (2025)
Block Expanded DINORET: Adapting Natural Domain Foundation Models for Retinal Imaging Without Catastrophic Forgetting
by: Zoellin, Jay, et al.
Published: (2024)
by: Zoellin, Jay, et al.
Published: (2024)
BlanketGen2-Fit3D: Synthetic Blanket Augmentation Towards Improving Real-World In-Bed Blanket Occluded Human Pose Estimation
by: Karácsony, Tamás, et al.
Published: (2025)
by: Karácsony, Tamás, et al.
Published: (2025)
Don't Forget your Inverse DDIM for Image Editing
by: Gomez-Trenado, Guillermo, et al.
Published: (2025)
by: Gomez-Trenado, Guillermo, et al.
Published: (2025)
Genflow Ad Studio: A Compound AI Architecture for Brand-Aligned, Self-Correcting Video Generation
by: Das, Debanshu, et al.
Published: (2026)
by: Das, Debanshu, et al.
Published: (2026)
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
by: Ziakas, Christos, et al.
Published: (2025)
by: Ziakas, Christos, et al.
Published: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026)
by: Liu, Zhi
Published: (2026)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
by: Safdar, Aon, et al.
Published: (2025)
by: Safdar, Aon, et al.
Published: (2025)
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
by: Grosbard, Idan Daniel, et al.
Published: (2026)
by: Grosbard, Idan Daniel, et al.
Published: (2026)
Fashion Florence: Fine-Tuning Florence-2 for Structured Fashion Attribute Extraction
by: Berlia, Anushree
Published: (2026)
by: Berlia, Anushree
Published: (2026)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
by: Dokme, Atahan, et al.
Published: (2026)
by: Dokme, Atahan, et al.
Published: (2026)
Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments
by: Gonzalez, Laura Alejandra Encinar, et al.
Published: (2025)
by: Gonzalez, Laura Alejandra Encinar, et al.
Published: (2025)
Deep Learning methodology for the identification of wood species using high-resolution macroscopic images
by: Herrera-Poyatos, David, et al.
Published: (2024)
by: Herrera-Poyatos, David, et al.
Published: (2024)
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
by: Wu, Qingyu, et al.
Published: (2026)
by: Wu, Qingyu, et al.
Published: (2026)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
by: Oliveira, Daniel A. P., et al.
Published: (2024)
by: Oliveira, Daniel A. P., et al.
Published: (2024)
ESCAPE: Energy-based Selective Adaptive Correction for Out-of-distribution 3D Human Pose Estimation
by: Bidulka, Luke, et al.
Published: (2024)
by: Bidulka, Luke, et al.
Published: (2024)
Unsupervised Decomposition and Recombination with Discriminator-Driven Diffusion Models
by: Wang, Archer, et al.
Published: (2026)
by: Wang, Archer, et al.
Published: (2026)
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
by: Huang, Yuhang, et al.
Published: (2025)
by: Huang, Yuhang, et al.
Published: (2025)
AI-Dentify: Deep learning for proximal caries detection on bitewing x-ray -- HUNT4 Oral Health Study
by: de Frutos, Javier Pérez, et al.
Published: (2023)
by: de Frutos, Javier Pérez, et al.
Published: (2023)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Dynamic Residual Encoding with Slide-Level Contrastive Learning for End-to-End Whole Slide Image Representation
by: Jin, Jing, et al.
Published: (2025)
by: Jin, Jing, et al.
Published: (2025)
On the Limitations of Vision-Language Models in Understanding Image Transforms
by: Anis, Ahmad Mustafa, et al.
Published: (2025)
by: Anis, Ahmad Mustafa, et al.
Published: (2025)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
by: Khorrami, Khazar, et al.
Published: (2021)
by: Khorrami, Khazar, et al.
Published: (2021)
From Deception to Perception: The Surprising Benefits of Deepfakes for Detecting, Measuring, and Mitigating Bias
by: Liu, Yizhi, et al.
Published: (2025)
by: Liu, Yizhi, et al.
Published: (2025)
Beyond Few-shot Object Detection: A Detailed Survey
by: Chudasama, Vishal, et al.
Published: (2024)
by: Chudasama, Vishal, et al.
Published: (2024)
Similar Items
-
Visual Language Models show widespread visual deficits on neuropsychological tests
by: Tangtartharakul, Gene, et al.
Published: (2025) -
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
by: Rudman, William, et al.
Published: (2026) -
Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection
by: Wang, Gaojian, et al.
Published: (2025) -
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
by: Khurdula, Harsha Vardhan, et al.
Published: (2024) -
Adapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners
by: Narasinghe, N. K. B. M. P. K. B., et al.
Published: (2025)