Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Hwang, Dongyoon, Lee, Byungkun, Lee, Hojoon, Kim, Hyunseung, Choo, Jaegul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
di: Kim, Donghu, et al.
Pubblicazione: (2024)
di: Kim, Donghu, et al.
Pubblicazione: (2024)
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
di: Kim, Hyunseung, et al.
Pubblicazione: (2024)
di: Kim, Hyunseung, et al.
Pubblicazione: (2024)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control
di: Cui, Zichen Jeff, et al.
Pubblicazione: (2024)
di: Cui, Zichen Jeff, et al.
Pubblicazione: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
di: Chattopadhyay, Nandish, et al.
Pubblicazione: (2026)
di: Chattopadhyay, Nandish, et al.
Pubblicazione: (2026)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
di: Chen, Lu, et al.
Pubblicazione: (2025)
di: Chen, Lu, et al.
Pubblicazione: (2025)
PHUMA: Physically-Grounded Humanoid Locomotion Dataset
di: Lee, Kyungmin, et al.
Pubblicazione: (2025)
di: Lee, Kyungmin, et al.
Pubblicazione: (2025)
Token Cropr: Faster ViTs for Quite a Few Tasks
di: Bergner, Benjamin, et al.
Pubblicazione: (2024)
di: Bergner, Benjamin, et al.
Pubblicazione: (2024)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
di: Heo, Jung Hwan, et al.
Pubblicazione: (2023)
di: Heo, Jung Hwan, et al.
Pubblicazione: (2023)
Octic Vision Transformers: Quicker ViTs Through Equivariance
di: Nordström, David, et al.
Pubblicazione: (2025)
di: Nordström, David, et al.
Pubblicazione: (2025)
Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
di: Yunusa, Haruna, et al.
Pubblicazione: (2024)
di: Yunusa, Haruna, et al.
Pubblicazione: (2024)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
di: Huang, Lianghuan, et al.
Pubblicazione: (2025)
di: Huang, Lianghuan, et al.
Pubblicazione: (2025)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
di: Han, Donghoon, et al.
Pubblicazione: (2023)
di: Han, Donghoon, et al.
Pubblicazione: (2023)
MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation
di: Lancaster, Patrick, et al.
Pubblicazione: (2023)
di: Lancaster, Patrick, et al.
Pubblicazione: (2023)
Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
di: Lappe, Alexander, et al.
Pubblicazione: (2025)
di: Lappe, Alexander, et al.
Pubblicazione: (2025)
RL makes MLLMs see better than SFT
di: Song, Junha, et al.
Pubblicazione: (2025)
di: Song, Junha, et al.
Pubblicazione: (2025)
TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers
di: Wang, Zhibo, et al.
Pubblicazione: (2026)
di: Wang, Zhibo, et al.
Pubblicazione: (2026)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
di: Alvetreti, Federico, et al.
Pubblicazione: (2025)
di: Alvetreti, Federico, et al.
Pubblicazione: (2025)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
di: Salzmann, Tim, et al.
Pubblicazione: (2024)
di: Salzmann, Tim, et al.
Pubblicazione: (2024)
Elastic ViTs from Pretrained Models without Retraining
di: Simoncini, Walter, et al.
Pubblicazione: (2025)
di: Simoncini, Walter, et al.
Pubblicazione: (2025)
Pretrained ViTs Yield Versatile Representations For Medical Images
di: Matsoukas, Christos, et al.
Pubblicazione: (2023)
di: Matsoukas, Christos, et al.
Pubblicazione: (2023)
ViTaSCOPE: Visuo-tactile Implicit Representation for In-hand Pose and Extrinsic Contact Estimation
di: Lee, Jayjun, et al.
Pubblicazione: (2025)
di: Lee, Jayjun, et al.
Pubblicazione: (2025)
ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing
di: Scherl, Alessandro, et al.
Pubblicazione: (2025)
di: Scherl, Alessandro, et al.
Pubblicazione: (2025)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2024)
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2024)
SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
di: Papais, Sandro, et al.
Pubblicazione: (2026)
di: Papais, Sandro, et al.
Pubblicazione: (2026)
TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
di: Fekri, Pedram, et al.
Pubblicazione: (2025)
di: Fekri, Pedram, et al.
Pubblicazione: (2025)
Understanding Particles From Video: Property Estimation of Granular Materials via Visuo-Haptic Learning
di: Zhang, Zeqing, et al.
Pubblicazione: (2024)
di: Zhang, Zeqing, et al.
Pubblicazione: (2024)
AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
di: Feng, Ruoxuan, et al.
Pubblicazione: (2025)
di: Feng, Ruoxuan, et al.
Pubblicazione: (2025)
Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
di: Elisha, Yehonatan, et al.
Pubblicazione: (2026)
di: Elisha, Yehonatan, et al.
Pubblicazione: (2026)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
di: Hwang, Dongyoon, et al.
Pubblicazione: (2025)
di: Hwang, Dongyoon, et al.
Pubblicazione: (2025)
ViT-2SPN: Vision Transformer-based Dual-Stream Self-Supervised Pretraining Networks for Retinal OCT Classification
di: Saraei, Mohammadreza, et al.
Pubblicazione: (2025)
di: Saraei, Mohammadreza, et al.
Pubblicazione: (2025)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
di: Chen, Zerui, et al.
Pubblicazione: (2024)
di: Chen, Zerui, et al.
Pubblicazione: (2024)
ViTCAE: ViT-based Class-conditioned Autoencoder
di: Jebraeeli, Vahid, et al.
Pubblicazione: (2025)
di: Jebraeeli, Vahid, et al.
Pubblicazione: (2025)
Visuo-Acoustic Hand Pose and Contact Estimation
di: Mao, Yuemin, et al.
Pubblicazione: (2025)
di: Mao, Yuemin, et al.
Pubblicazione: (2025)
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs
di: Wang, Xiyao, et al.
Pubblicazione: (2026)
di: Wang, Xiyao, et al.
Pubblicazione: (2026)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
di: Shah, Arya, et al.
Pubblicazione: (2025)
di: Shah, Arya, et al.
Pubblicazione: (2025)
P3-PO: Prescriptive Point Priors for Visuo-Spatial Generalization of Robot Policies
di: Levy, Mara, et al.
Pubblicazione: (2024)
di: Levy, Mara, et al.
Pubblicazione: (2024)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
di: Zhong, Yunshan, et al.
Pubblicazione: (2023)
di: Zhong, Yunshan, et al.
Pubblicazione: (2023)
Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning
di: Giannakakis, Nikos, et al.
Pubblicazione: (2025)
di: Giannakakis, Nikos, et al.
Pubblicazione: (2025)
GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
di: Lee, Hyunju, et al.
Pubblicazione: (2025)
di: Lee, Hyunju, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
di: Kim, Donghu, et al.
Pubblicazione: (2024) -
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
di: Kim, Hyunseung, et al.
Pubblicazione: (2024) -
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
di: Kim, Donghyun, et al.
Pubblicazione: (2024) -
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control
di: Cui, Zichen Jeff, et al.
Pubblicazione: (2024) -
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
di: Chattopadhyay, Nandish, et al.
Pubblicazione: (2026)