Navigating Efficiency in MobileViT through Gaussian Process on Global Architecture Factors
Fuente:
arXiv
Guardado en:
| Autores principales: | Meng, Ke, Chen, Kai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
por: Gebremedhin, Tekleab G., et al.
Publicado: (2025)
por: Gebremedhin, Tekleab G., et al.
Publicado: (2025)
Efficient Few-Shot Learning for Edge AI via Knowledge Distillation on MobileViT
por: Tsuyuki, Shuhei, et al.
Publicado: (2026)
por: Tsuyuki, Shuhei, et al.
Publicado: (2026)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
por: Tonmoy, Moshiur Rahman, et al.
Publicado: (2025)
por: Tonmoy, Moshiur Rahman, et al.
Publicado: (2025)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
por: Wu, Fengyi, et al.
Publicado: (2025)
por: Wu, Fengyi, et al.
Publicado: (2025)
ViRED: Prediction of Visual Relations in Engineering Drawings
por: Gu, Chao, et al.
Publicado: (2024)
por: Gu, Chao, et al.
Publicado: (2024)
UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
por: Dai, Guangzhao, et al.
Publicado: (2024)
por: Dai, Guangzhao, et al.
Publicado: (2024)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
por: Tong, Haoyu, et al.
Publicado: (2026)
por: Tong, Haoyu, et al.
Publicado: (2026)
ViPO: Visual Preference Optimization at Scale
por: Li, Ming, et al.
Publicado: (2026)
por: Li, Ming, et al.
Publicado: (2026)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
por: Kim, Ye-eun, et al.
Publicado: (2025)
por: Kim, Ye-eun, et al.
Publicado: (2025)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
por: Wu, Linquan, et al.
Publicado: (2026)
por: Wu, Linquan, et al.
Publicado: (2026)
Self-Attention through Kernel-Eigen Pair Sparse Variational Gaussian Processes
por: Chen, Yingyi, et al.
Publicado: (2024)
por: Chen, Yingyi, et al.
Publicado: (2024)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
por: Bartoccioni, Florent, et al.
Publicado: (2025)
por: Bartoccioni, Florent, et al.
Publicado: (2025)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
por: Shah, Arya, et al.
Publicado: (2025)
por: Shah, Arya, et al.
Publicado: (2025)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
por: Li, Zhengang, et al.
Publicado: (2024)
por: Li, Zhengang, et al.
Publicado: (2024)
Feature-EndoGaussian: Feature Distilled Gaussian Splatting in Surgical Deformable Scene Reconstruction
por: Li, Kai, et al.
Publicado: (2025)
por: Li, Kai, et al.
Publicado: (2025)
ViT-Lens: Towards Omni-modal Representations
por: Lei, Weixian, et al.
Publicado: (2023)
por: Lei, Weixian, et al.
Publicado: (2023)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
por: Zhang, Mengchen, et al.
Publicado: (2025)
por: Zhang, Mengchen, et al.
Publicado: (2025)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
por: Xu, Xuwei, et al.
Publicado: (2023)
por: Xu, Xuwei, et al.
Publicado: (2023)
MobileDenseAttn:A Dual-Stream Architecture for Accurate and Interpretable Brain Tumor Detection
por: Banik, Shudipta, et al.
Publicado: (2025)
por: Banik, Shudipta, et al.
Publicado: (2025)
Gaussian Grouping: Segment and Edit Anything in 3D Scenes
por: Ye, Mingqiao, et al.
Publicado: (2023)
por: Ye, Mingqiao, et al.
Publicado: (2023)
SFMViT: SlowFast Meet ViT in Chaotic World
por: Lin, Jiaying, et al.
Publicado: (2024)
por: Lin, Jiaying, et al.
Publicado: (2024)
7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting
por: Gao, Zhongpai, et al.
Publicado: (2025)
por: Gao, Zhongpai, et al.
Publicado: (2025)
Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning
por: Qin, Wenda, et al.
Publicado: (2025)
por: Qin, Wenda, et al.
Publicado: (2025)
Enhancing Layer Attention Efficiency through Pruning Redundant Retrievals
por: Li, Hanze, et al.
Publicado: (2025)
por: Li, Hanze, et al.
Publicado: (2025)
GI-NAS: Boosting Gradient Inversion Attacks Through Adaptive Neural Architecture Search
por: Yu, Wenbo, et al.
Publicado: (2024)
por: Yu, Wenbo, et al.
Publicado: (2024)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
por: Li, Kailing, et al.
Publicado: (2025)
por: Li, Kailing, et al.
Publicado: (2025)
MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos
por: Sun, Jian, et al.
Publicado: (2023)
por: Sun, Jian, et al.
Publicado: (2023)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
por: Feng, Wenfeng, et al.
Publicado: (2025)
por: Feng, Wenfeng, et al.
Publicado: (2025)
6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering
por: Gao, Zhongpai, et al.
Publicado: (2024)
por: Gao, Zhongpai, et al.
Publicado: (2024)
NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps
por: Zhan, Dijia, et al.
Publicado: (2026)
por: Zhan, Dijia, et al.
Publicado: (2026)
HydraViT: Stacking Heads for a Scalable ViT
por: Haberer, Janek, et al.
Publicado: (2024)
por: Haberer, Janek, et al.
Publicado: (2024)
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
por: Wang, Hui, et al.
Publicado: (2026)
por: Wang, Hui, et al.
Publicado: (2026)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
por: Tang, Tianqi, et al.
Publicado: (2024)
por: Tang, Tianqi, et al.
Publicado: (2024)
LoViT: Long Video Transformer for Surgical Phase Recognition
por: Liu, Yang, et al.
Publicado: (2023)
por: Liu, Yang, et al.
Publicado: (2023)
ViLLa: A Neuro-Symbolic approach for Animal Monitoring
por: Koduri, Harsha
Publicado: (2025)
por: Koduri, Harsha
Publicado: (2025)
Sub-token ViT Embedding via Stochastic Resonance Transformers
por: Lao, Dong, et al.
Publicado: (2023)
por: Lao, Dong, et al.
Publicado: (2023)
Global Intervention and Distillation for Federated Out-of-Distribution Generalization
por: Qi, Zhuang, et al.
Publicado: (2025)
por: Qi, Zhuang, et al.
Publicado: (2025)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
por: Sheng, Kai, et al.
Publicado: (2026)
por: Sheng, Kai, et al.
Publicado: (2026)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
por: Zou, Dongyun, et al.
Publicado: (2026)
por: Zou, Dongyun, et al.
Publicado: (2026)
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
por: Xu, Xuwei, et al.
Publicado: (2025)
por: Xu, Xuwei, et al.
Publicado: (2025)
Ejemplares similares
-
Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
por: Gebremedhin, Tekleab G., et al.
Publicado: (2025) -
Efficient Few-Shot Learning for Edge AI via Knowledge Distillation on MobileViT
por: Tsuyuki, Shuhei, et al.
Publicado: (2026) -
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
por: Tonmoy, Moshiur Rahman, et al.
Publicado: (2025) -
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
por: Wu, Fengyi, et al.
Publicado: (2025) -
ViRED: Prediction of Visual Relations in Engineering Drawings
por: Gu, Chao, et al.
Publicado: (2024)