AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Vasa, Santosh, Ramadwar, Aditi, Darabattula, Jnana Rama Krishna, Anwar, Md Zafar, Antol, Stanislaw, Vatavu, Andrei, Monninger, Thomas, Ding, Sihao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AugMapNet: Improving Spatial Latent Structure via BEV Grid Augmentation for Enhanced Vectorized Online HD Map Construction
by: Monninger, Thomas, et al.
Published: (2025)
by: Monninger, Thomas, et al.
Published: (2025)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
by: Ding, Sihao, et al.
Published: (2025)
by: Ding, Sihao, et al.
Published: (2025)
MapDiffusion: Generative Diffusion for Vectorized Online HD Map Construction and Uncertainty Estimation in Autonomous Driving
by: Monninger, Thomas, et al.
Published: (2025)
by: Monninger, Thomas, et al.
Published: (2025)
TempBEV: Improving Learned BEV Encoders with Combined Image and BEV Space Temporal Aggregation
by: Monninger, Thomas, et al.
Published: (2024)
by: Monninger, Thomas, et al.
Published: (2024)
NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction
by: Monninger, Thomas, et al.
Published: (2025)
by: Monninger, Thomas, et al.
Published: (2025)
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
by: Monninger, Thomas, et al.
Published: (2026)
by: Monninger, Thomas, et al.
Published: (2026)
LMT-Net: Lane Model Transformer Network for Automated HD Mapping from Sparse Vehicle Observations
by: Mink, Michael, et al.
Published: (2024)
by: Mink, Michael, et al.
Published: (2024)
AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
by: Sun, Xiaolou, et al.
Published: (2026)
by: Sun, Xiaolou, et al.
Published: (2026)
A Systematic Literature Review on Deep Learning-based Depth Estimation in Computer Vision
by: Rohan, Ali, et al.
Published: (2025)
by: Rohan, Ali, et al.
Published: (2025)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
One Agent to Guide Them All: Empowering MLLMs for Vision-and-Language Navigation via Explicit World Representation
by: Li, Zerui, et al.
Published: (2026)
by: Li, Zerui, et al.
Published: (2026)
Integration of Computer Vision with Adaptive Control for Autonomous Driving Using ADORE
by: Ahammed, Abu Shad, et al.
Published: (2025)
by: Ahammed, Abu Shad, et al.
Published: (2025)
Automated Parking Planning with Vision-Based BEV Approach
by: Zhao, Yuxuan
Published: (2024)
by: Zhao, Yuxuan
Published: (2024)
Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts
by: Chen, Hongyi, et al.
Published: (2024)
by: Chen, Hongyi, et al.
Published: (2024)
Deformable Radar Polygon: A Lightweight and Predictable Occupancy Representation for Short-range Collision Avoidance
by: Xiangyu, Gao, et al.
Published: (2022)
by: Xiangyu, Gao, et al.
Published: (2022)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
by: Ding, Hongyu, et al.
Published: (2025)
by: Ding, Hongyu, et al.
Published: (2025)
Vision-based Perception System for Automated Delivery Robot-Pedestrians Interactions
by: Tushe, Ergi, et al.
Published: (2025)
by: Tushe, Ergi, et al.
Published: (2025)
An Open Source Computer Vision and Machine Learning Framework for Affordable Life Science Robotic Automation
by: Logan, Zachary, et al.
Published: (2026)
by: Logan, Zachary, et al.
Published: (2026)
Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving
by: Samak, Chinmay Vilas, et al.
Published: (2025)
by: Samak, Chinmay Vilas, et al.
Published: (2025)
DTactive: A Vision-Based Tactile Sensor with Active Surface
by: Xu, Jikai, et al.
Published: (2024)
by: Xu, Jikai, et al.
Published: (2024)
VL-Explore: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots
by: Zhang, Yuxuan, et al.
Published: (2025)
by: Zhang, Yuxuan, et al.
Published: (2025)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
by: Li, Pengteng, et al.
Published: (2026)
by: Li, Pengteng, et al.
Published: (2026)
MonoMPC: Monocular Vision Based Navigation with Learned Collision Model and Risk-Aware Model Predictive Control
by: Sharma, Basant, et al.
Published: (2025)
by: Sharma, Basant, et al.
Published: (2025)
Mechanical Automation with Vision: A Design for Rubik's Cube Solver
by: Chalise, Abhinav, et al.
Published: (2025)
by: Chalise, Abhinav, et al.
Published: (2025)
AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory
by: Lan, Zhiqian, et al.
Published: (2025)
by: Lan, Zhiqian, et al.
Published: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
by: Deng, Shengliang, et al.
Published: (2025)
by: Deng, Shengliang, et al.
Published: (2025)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
by: Song, Kun, et al.
Published: (2025)
by: Song, Kun, et al.
Published: (2025)
Autonomous Aggregate Sorting in Construction and Mining via Computer Vision-Aided Robotic Arm Systems
by: Shawon, Md. Taherul Islam, et al.
Published: (2025)
by: Shawon, Md. Taherul Islam, et al.
Published: (2025)
Vision Controlled Orthotic Hand Exoskeleton
by: Blais, Connor, et al.
Published: (2025)
by: Blais, Connor, et al.
Published: (2025)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
RobotDesignGPT: Automated Robot Design Synthesis using Vision Language Models
by: Sontakke, Nitish, et al.
Published: (2026)
by: Sontakke, Nitish, et al.
Published: (2026)
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
AutoDSL: Automated domain-specific language design for structural representation of procedures with constraints
by: Shi, Yu-Zhe, et al.
Published: (2024)
by: Shi, Yu-Zhe, et al.
Published: (2024)
Goal-Based Vision-Language Driving
by: Patapati, Santosh, et al.
Published: (2025)
by: Patapati, Santosh, et al.
Published: (2025)
AutoDrive-QA: A Multiple-Choice Benchmark for Vision-Language Evaluation in Urban Autonomous Driving
by: Khalili, Boshra, et al.
Published: (2025)
by: Khalili, Boshra, et al.
Published: (2025)
TacScope: A Miniaturized Vision‐Based Tactile Sensor for Surgical Applications
by: Md Rakibul Islam Prince, et al.
Published: (2025)
by: Md Rakibul Islam Prince, et al.
Published: (2025)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Similar Items
-
AugMapNet: Improving Spatial Latent Structure via BEV Grid Augmentation for Enhanced Vectorized Online HD Map Construction
by: Monninger, Thomas, et al.
Published: (2025) -
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
by: Ding, Sihao, et al.
Published: (2025) -
MapDiffusion: Generative Diffusion for Vectorized Online HD Map Construction and Uncertainty Estimation in Autonomous Driving
by: Monninger, Thomas, et al.
Published: (2025) -
TempBEV: Improving Learned BEV Encoders with Combined Image and BEV Space Temporal Aggregation
by: Monninger, Thomas, et al.
Published: (2024) -
NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction
by: Monninger, Thomas, et al.
Published: (2025)