MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Galella, Santiago, Osuna-Vargas, Pamela, Wehrheim, Maren, Vilas, Martina G., Roig, Gemma, Kaschube, Matthias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Sensorimotor Vision Transformer
von: Gadzicki, Konrad, et al.
Veröffentlicht: (2025)
von: Gadzicki, Konrad, et al.
Veröffentlicht: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
von: Wu, Jason, et al.
Veröffentlicht: (2026)
von: Wu, Jason, et al.
Veröffentlicht: (2026)
CARScenes: Semantic VLM Dataset for Safe Autonomous Driving
von: He, Yuankai, et al.
Veröffentlicht: (2025)
von: He, Yuankai, et al.
Veröffentlicht: (2025)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
von: Korolkov, Vasilii
Veröffentlicht: (2025)
von: Korolkov, Vasilii
Veröffentlicht: (2025)
TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition
von: Hassan, Imtiaz Ul, et al.
Veröffentlicht: (2026)
von: Hassan, Imtiaz Ul, et al.
Veröffentlicht: (2026)
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification
von: Ho, Darryl, et al.
Veröffentlicht: (2025)
von: Ho, Darryl, et al.
Veröffentlicht: (2025)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
von: Peng, Hongxing, et al.
Veröffentlicht: (2025)
von: Peng, Hongxing, et al.
Veröffentlicht: (2025)
PhysicsNeRF: Physics-Guided 3D Reconstruction from Sparse Views
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
SH17: A Dataset for Human Safety and Personal Protective Equipment Detection in Manufacturing Industry
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
Implementing Adaptations for Vision AutoRegressive Model
von: Shaikh, Kaif, et al.
Veröffentlicht: (2025)
von: Shaikh, Kaif, et al.
Veröffentlicht: (2025)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
von: Rajapaksha, Uchitha, et al.
Veröffentlicht: (2024)
von: Rajapaksha, Uchitha, et al.
Veröffentlicht: (2024)
Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
von: Jiang, Xintong, et al.
Veröffentlicht: (2025)
von: Jiang, Xintong, et al.
Veröffentlicht: (2025)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
von: Diller, Christian, et al.
Veröffentlicht: (2023)
von: Diller, Christian, et al.
Veröffentlicht: (2023)
Sign language recognition based on deep learning and low-cost handcrafted descriptors
von: Carneiro, Alvaro Leandro Cavalcante, et al.
Veröffentlicht: (2024)
von: Carneiro, Alvaro Leandro Cavalcante, et al.
Veröffentlicht: (2024)
Data Augmentation with Diffusion Models for Colon Polyp Localization on the Low Data Regime: How much real data is enough?
von: Tormos, Adrian, et al.
Veröffentlicht: (2024)
von: Tormos, Adrian, et al.
Veröffentlicht: (2024)
LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation
von: Wei, Hualiang, et al.
Veröffentlicht: (2026)
von: Wei, Hualiang, et al.
Veröffentlicht: (2026)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
von: Brovko, D. V.
Veröffentlicht: (2025)
von: Brovko, D. V.
Veröffentlicht: (2025)
Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition
von: Carreira, Luiz G F, et al.
Veröffentlicht: (2026)
von: Carreira, Luiz G F, et al.
Veröffentlicht: (2026)
Beyond still images: Temporal features and input variance resilience
von: Fadaei, Amir Hosein, et al.
Veröffentlicht: (2023)
von: Fadaei, Amir Hosein, et al.
Veröffentlicht: (2023)
FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
von: Wang, Zixing, et al.
Veröffentlicht: (2025)
von: Wang, Zixing, et al.
Veröffentlicht: (2025)
UGOD: Uncertainty-Guided Differentiable Opacity and Soft Dropout for Enhanced Sparse-View 3DGS
von: Guo, Zhihao, et al.
Veröffentlicht: (2025)
von: Guo, Zhihao, et al.
Veröffentlicht: (2025)
Neural Attention: A Novel Mechanism for Enhanced Expressive Power in Transformer Models
von: DiGiugno, Andrew, et al.
Veröffentlicht: (2025)
von: DiGiugno, Andrew, et al.
Veröffentlicht: (2025)
Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data
von: Thomas, Ancymol, et al.
Veröffentlicht: (2026)
von: Thomas, Ancymol, et al.
Veröffentlicht: (2026)
Introspection in Learned Semantic Scene Graph Localisation
von: Bissessur, Manshika Charvi, et al.
Veröffentlicht: (2025)
von: Bissessur, Manshika Charvi, et al.
Veröffentlicht: (2025)
Capacity Constraint Analysis Using Object Detection for Smart Manufacturing
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
von: Ahmad, Hafiz Mughees, et al.
Veröffentlicht: (2024)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
von: Lin, Xiaojian, et al.
Veröffentlicht: (2025)
von: Lin, Xiaojian, et al.
Veröffentlicht: (2025)
Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
von: Gurbindo, Unai, et al.
Veröffentlicht: (2025)
von: Gurbindo, Unai, et al.
Veröffentlicht: (2025)
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations
von: Diller, Christian, et al.
Veröffentlicht: (2022)
von: Diller, Christian, et al.
Veröffentlicht: (2022)
DMFourLLIE: Dual-Stage and Multi-Branch Fourier Network for Low-Light Image Enhancement
von: Zhang, Tongshun, et al.
Veröffentlicht: (2024)
von: Zhang, Tongshun, et al.
Veröffentlicht: (2024)
IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
von: Mishra, Shashank, et al.
Veröffentlicht: (2025)
von: Mishra, Shashank, et al.
Veröffentlicht: (2025)
Detecting AI-Generated Videos with Spiking Neural Networks
von: Jang, Minsuk, et al.
Veröffentlicht: (2026)
von: Jang, Minsuk, et al.
Veröffentlicht: (2026)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
Have We Mastered Scale in Deep Monocular Visual SLAM? The ScaleMaster Dataset and Benchmark
von: Ju, Hyoseok, et al.
Veröffentlicht: (2026)
von: Ju, Hyoseok, et al.
Veröffentlicht: (2026)
Objaverse++: Curated 3D Object Dataset with Quality Annotations
von: Lin, Chendi, et al.
Veröffentlicht: (2025)
von: Lin, Chendi, et al.
Veröffentlicht: (2025)
Salient Concept-Aware Generative Data Augmentation
von: Zhao, Tianchen, et al.
Veröffentlicht: (2025)
von: Zhao, Tianchen, et al.
Veröffentlicht: (2025)
A deep learning approach to track eye movements based on events
von: Seth, Chirag, et al.
Veröffentlicht: (2025)
von: Seth, Chirag, et al.
Veröffentlicht: (2025)
Single-sample image-fusion upsampling of fluorescence lifetime images
von: Kapitány, Valentin, et al.
Veröffentlicht: (2024)
von: Kapitány, Valentin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Sensorimotor Vision Transformer
von: Gadzicki, Konrad, et al.
Veröffentlicht: (2025) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
von: Wu, Jason, et al.
Veröffentlicht: (2026) -
CARScenes: Semantic VLM Dataset for Safe Autonomous Driving
von: He, Yuankai, et al.
Veröffentlicht: (2025) -
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
von: Korolkov, Vasilii
Veröffentlicht: (2025)