Salvato in:
| Autori principali: | Meng, Ke, Chen, Kai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2406.04820 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
di: Gebremedhin, Tekleab G., et al.
Pubblicazione: (2025)
di: Gebremedhin, Tekleab G., et al.
Pubblicazione: (2025)
Efficient Few-Shot Learning for Edge AI via Knowledge Distillation on MobileViT
di: Tsuyuki, Shuhei, et al.
Pubblicazione: (2026)
di: Tsuyuki, Shuhei, et al.
Pubblicazione: (2026)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
di: Tonmoy, Moshiur Rahman, et al.
Pubblicazione: (2025)
di: Tonmoy, Moshiur Rahman, et al.
Pubblicazione: (2025)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
di: Wu, Fengyi, et al.
Pubblicazione: (2025)
di: Wu, Fengyi, et al.
Pubblicazione: (2025)
ViRED: Prediction of Visual Relations in Engineering Drawings
di: Gu, Chao, et al.
Pubblicazione: (2024)
di: Gu, Chao, et al.
Pubblicazione: (2024)
ViPO: Visual Preference Optimization at Scale
di: Li, Ming, et al.
Pubblicazione: (2026)
di: Li, Ming, et al.
Pubblicazione: (2026)
UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
di: Dai, Guangzhao, et al.
Pubblicazione: (2024)
di: Dai, Guangzhao, et al.
Pubblicazione: (2024)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
di: Tong, Haoyu, et al.
Pubblicazione: (2026)
di: Tong, Haoyu, et al.
Pubblicazione: (2026)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
di: Wu, Linquan, et al.
Pubblicazione: (2026)
di: Wu, Linquan, et al.
Pubblicazione: (2026)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
di: Kim, Ye-eun, et al.
Pubblicazione: (2025)
di: Kim, Ye-eun, et al.
Pubblicazione: (2025)
Self-Attention through Kernel-Eigen Pair Sparse Variational Gaussian Processes
di: Chen, Yingyi, et al.
Pubblicazione: (2024)
di: Chen, Yingyi, et al.
Pubblicazione: (2024)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
di: Bartoccioni, Florent, et al.
Pubblicazione: (2025)
di: Bartoccioni, Florent, et al.
Pubblicazione: (2025)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
di: Li, Zhengang, et al.
Pubblicazione: (2024)
di: Li, Zhengang, et al.
Pubblicazione: (2024)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
di: Shah, Arya, et al.
Pubblicazione: (2025)
di: Shah, Arya, et al.
Pubblicazione: (2025)
ViT-Lens: Towards Omni-modal Representations
di: Lei, Weixian, et al.
Pubblicazione: (2023)
di: Lei, Weixian, et al.
Pubblicazione: (2023)
MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos
di: Sun, Jian, et al.
Pubblicazione: (2023)
di: Sun, Jian, et al.
Pubblicazione: (2023)
Feature-EndoGaussian: Feature Distilled Gaussian Splatting in Surgical Deformable Scene Reconstruction
di: Li, Kai, et al.
Pubblicazione: (2025)
di: Li, Kai, et al.
Pubblicazione: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
HydraViT: Stacking Heads for a Scalable ViT
di: Haberer, Janek, et al.
Pubblicazione: (2024)
di: Haberer, Janek, et al.
Pubblicazione: (2024)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
di: Xu, Xuwei, et al.
Pubblicazione: (2023)
di: Xu, Xuwei, et al.
Pubblicazione: (2023)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
di: Li, Kailing, et al.
Pubblicazione: (2025)
di: Li, Kailing, et al.
Pubblicazione: (2025)
Gaussian Grouping: Segment and Edit Anything in 3D Scenes
di: Ye, Mingqiao, et al.
Pubblicazione: (2023)
di: Ye, Mingqiao, et al.
Pubblicazione: (2023)
MobileDenseAttn:A Dual-Stream Architecture for Accurate and Interpretable Brain Tumor Detection
di: Banik, Shudipta, et al.
Pubblicazione: (2025)
di: Banik, Shudipta, et al.
Pubblicazione: (2025)
SFMViT: SlowFast Meet ViT in Chaotic World
di: Lin, Jiaying, et al.
Pubblicazione: (2024)
di: Lin, Jiaying, et al.
Pubblicazione: (2024)
7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
GI-NAS: Boosting Gradient Inversion Attacks Through Adaptive Neural Architecture Search
di: Yu, Wenbo, et al.
Pubblicazione: (2024)
di: Yu, Wenbo, et al.
Pubblicazione: (2024)
OmniPatch: A Universal Adversarial Patch for ViT-CNN Cross-Architecture Transfer in Semantic Segmentation
di: Aggarwal, Aarush, et al.
Pubblicazione: (2026)
di: Aggarwal, Aarush, et al.
Pubblicazione: (2026)
6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
Enhancing Layer Attention Efficiency through Pruning Redundant Retrievals
di: Li, Hanze, et al.
Pubblicazione: (2025)
di: Li, Hanze, et al.
Pubblicazione: (2025)
Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning
di: Qin, Wenda, et al.
Pubblicazione: (2025)
di: Qin, Wenda, et al.
Pubblicazione: (2025)
Global Intervention and Distillation for Federated Out-of-Distribution Generalization
di: Qi, Zhuang, et al.
Pubblicazione: (2025)
di: Qi, Zhuang, et al.
Pubblicazione: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
di: Feng, Wenfeng, et al.
Pubblicazione: (2025)
di: Feng, Wenfeng, et al.
Pubblicazione: (2025)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
di: Tang, Tianqi, et al.
Pubblicazione: (2024)
di: Tang, Tianqi, et al.
Pubblicazione: (2024)
LoViT: Long Video Transformer for Surgical Phase Recognition
di: Liu, Yang, et al.
Pubblicazione: (2023)
di: Liu, Yang, et al.
Pubblicazione: (2023)
ViLLa: A Neuro-Symbolic approach for Animal Monitoring
di: Koduri, Harsha
Pubblicazione: (2025)
di: Koduri, Harsha
Pubblicazione: (2025)
Sub-token ViT Embedding via Stochastic Resonance Transformers
di: Lao, Dong, et al.
Pubblicazione: (2023)
di: Lao, Dong, et al.
Pubblicazione: (2023)
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
di: Wang, Hui, et al.
Pubblicazione: (2026)
di: Wang, Hui, et al.
Pubblicazione: (2026)
P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation
di: Sheng, Kai, et al.
Pubblicazione: (2026)
di: Sheng, Kai, et al.
Pubblicazione: (2026)
NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps
di: Zhan, Dijia, et al.
Pubblicazione: (2026)
di: Zhan, Dijia, et al.
Pubblicazione: (2026)
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
di: He, Zongtao, et al.
Pubblicazione: (2025)
di: He, Zongtao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
di: Gebremedhin, Tekleab G., et al.
Pubblicazione: (2025) -
Efficient Few-Shot Learning for Edge AI via Knowledge Distillation on MobileViT
di: Tsuyuki, Shuhei, et al.
Pubblicazione: (2026) -
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
di: Tonmoy, Moshiur Rahman, et al.
Pubblicazione: (2025) -
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
di: Wu, Fengyi, et al.
Pubblicazione: (2025) -
ViRED: Prediction of Visual Relations in Engineering Drawings
di: Gu, Chao, et al.
Pubblicazione: (2024)