A Sensorimotor Vision Transformer
Fuente:
arXiv
Salvato in:
| Autori principali: | Gadzicki, Konrad, Schill, Kerstin, Zetzsche, Christoph |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
di: Galella, Santiago, et al.
Pubblicazione: (2026)
di: Galella, Santiago, et al.
Pubblicazione: (2026)
Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition
di: Carreira, Luiz G F, et al.
Pubblicazione: (2026)
di: Carreira, Luiz G F, et al.
Pubblicazione: (2026)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
di: Peng, Hongxing, et al.
Pubblicazione: (2025)
di: Peng, Hongxing, et al.
Pubblicazione: (2025)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
di: Rajapaksha, Uchitha, et al.
Pubblicazione: (2024)
di: Rajapaksha, Uchitha, et al.
Pubblicazione: (2024)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
di: Diller, Christian, et al.
Pubblicazione: (2023)
di: Diller, Christian, et al.
Pubblicazione: (2023)
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification
di: Ho, Darryl, et al.
Pubblicazione: (2025)
di: Ho, Darryl, et al.
Pubblicazione: (2025)
PhysicsNeRF: Physics-Guided 3D Reconstruction from Sparse Views
di: Barhdadi, Mohamed Rayan, et al.
Pubblicazione: (2025)
di: Barhdadi, Mohamed Rayan, et al.
Pubblicazione: (2025)
TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition
di: Hassan, Imtiaz Ul, et al.
Pubblicazione: (2026)
di: Hassan, Imtiaz Ul, et al.
Pubblicazione: (2026)
Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data
di: Thomas, Ancymol, et al.
Pubblicazione: (2026)
di: Thomas, Ancymol, et al.
Pubblicazione: (2026)
Data Augmentation with Diffusion Models for Colon Polyp Localization on the Low Data Regime: How much real data is enough?
di: Tormos, Adrian, et al.
Pubblicazione: (2024)
di: Tormos, Adrian, et al.
Pubblicazione: (2024)
DMFourLLIE: Dual-Stage and Multi-Branch Fourier Network for Low-Light Image Enhancement
di: Zhang, Tongshun, et al.
Pubblicazione: (2024)
di: Zhang, Tongshun, et al.
Pubblicazione: (2024)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
di: Wu, Jason, et al.
Pubblicazione: (2026)
di: Wu, Jason, et al.
Pubblicazione: (2026)
Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
di: Jiang, Xintong, et al.
Pubblicazione: (2025)
di: Jiang, Xintong, et al.
Pubblicazione: (2025)
SH17: A Dataset for Human Safety and Personal Protective Equipment Detection in Manufacturing Industry
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
Sign language recognition based on deep learning and low-cost handcrafted descriptors
di: Carneiro, Alvaro Leandro Cavalcante, et al.
Pubblicazione: (2024)
di: Carneiro, Alvaro Leandro Cavalcante, et al.
Pubblicazione: (2024)
Capacity Constraint Analysis Using Object Detection for Smart Manufacturing
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations
di: Diller, Christian, et al.
Pubblicazione: (2022)
di: Diller, Christian, et al.
Pubblicazione: (2022)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
di: Brovko, D. V.
Pubblicazione: (2025)
di: Brovko, D. V.
Pubblicazione: (2025)
FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
di: Wang, Zixing, et al.
Pubblicazione: (2025)
di: Wang, Zixing, et al.
Pubblicazione: (2025)
UGOD: Uncertainty-Guided Differentiable Opacity and Soft Dropout for Enhanced Sparse-View 3DGS
di: Guo, Zhihao, et al.
Pubblicazione: (2025)
di: Guo, Zhihao, et al.
Pubblicazione: (2025)
Beyond still images: Temporal features and input variance resilience
di: Fadaei, Amir Hosein, et al.
Pubblicazione: (2023)
di: Fadaei, Amir Hosein, et al.
Pubblicazione: (2023)
Implementing Adaptations for Vision AutoRegressive Model
di: Shaikh, Kaif, et al.
Pubblicazione: (2025)
di: Shaikh, Kaif, et al.
Pubblicazione: (2025)
Detecting AI-Generated Videos with Spiking Neural Networks
di: Jang, Minsuk, et al.
Pubblicazione: (2026)
di: Jang, Minsuk, et al.
Pubblicazione: (2026)
CARScenes: Semantic VLM Dataset for Safe Autonomous Driving
di: He, Yuankai, et al.
Pubblicazione: (2025)
di: He, Yuankai, et al.
Pubblicazione: (2025)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
di: Lin, Xiaojian, et al.
Pubblicazione: (2025)
di: Lin, Xiaojian, et al.
Pubblicazione: (2025)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
di: Korolkov, Vasilii
Pubblicazione: (2025)
di: Korolkov, Vasilii
Pubblicazione: (2025)
Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
di: Gurbindo, Unai, et al.
Pubblicazione: (2025)
di: Gurbindo, Unai, et al.
Pubblicazione: (2025)
IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
di: Mishra, Shashank, et al.
Pubblicazione: (2025)
di: Mishra, Shashank, et al.
Pubblicazione: (2025)
LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation
di: Wei, Hualiang, et al.
Pubblicazione: (2026)
di: Wei, Hualiang, et al.
Pubblicazione: (2026)
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
di: Baxevanakis, Spiros, et al.
Pubblicazione: (2026)
di: Baxevanakis, Spiros, et al.
Pubblicazione: (2026)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
di: Qesaraku, Bjorna, et al.
Pubblicazione: (2025)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
di: Ji, Binbin, et al.
Pubblicazione: (2025)
di: Ji, Binbin, et al.
Pubblicazione: (2025)
A deep learning approach to track eye movements based on events
di: Seth, Chirag, et al.
Pubblicazione: (2025)
di: Seth, Chirag, et al.
Pubblicazione: (2025)
A novel approach towards the classification of Bone Fracture from Musculoskeletal Radiography images using Attention Based Transfer Learning
di: Ruhi, Sayeda Sanzida Ferdous, et al.
Pubblicazione: (2024)
di: Ruhi, Sayeda Sanzida Ferdous, et al.
Pubblicazione: (2024)
POC-SLT: Partial Object Completion with SDF Latent Transformers
di: Zakeri, Faezeh, et al.
Pubblicazione: (2024)
di: Zakeri, Faezeh, et al.
Pubblicazione: (2024)
Single-sample image-fusion upsampling of fluorescence lifetime images
di: Kapitány, Valentin, et al.
Pubblicazione: (2024)
di: Kapitány, Valentin, et al.
Pubblicazione: (2024)
Salient Concept-Aware Generative Data Augmentation
di: Zhao, Tianchen, et al.
Pubblicazione: (2025)
di: Zhao, Tianchen, et al.
Pubblicazione: (2025)
DNRSelect: Active Best View Selection for Deferred Neural Rendering
di: Wu, Dongli, et al.
Pubblicazione: (2025)
di: Wu, Dongli, et al.
Pubblicazione: (2025)
Understanding colors of Dufaycolor: Can we recover them using historical colorimetric and spectral data?
di: Hubička, Jan, et al.
Pubblicazione: (2025)
di: Hubička, Jan, et al.
Pubblicazione: (2025)
Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study
di: Nie, Shiheng, et al.
Pubblicazione: (2026)
di: Nie, Shiheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
di: Galella, Santiago, et al.
Pubblicazione: (2026) -
Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition
di: Carreira, Luiz G F, et al.
Pubblicazione: (2026) -
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
di: Peng, Hongxing, et al.
Pubblicazione: (2025) -
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
di: Rajapaksha, Uchitha, et al.
Pubblicazione: (2024) -
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
di: Diller, Christian, et al.
Pubblicazione: (2023)