Saved in:
| Main Authors: | Etchegaray, Djamahl, Fu, Yuxia, Huang, Zi, Luo, Yadan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.00525 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments
by: Etchegaray, Djamahl, et al.
Published: (2024)
by: Etchegaray, Djamahl, et al.
Published: (2024)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
by: Wang, Lu, et al.
Published: (2025)
by: Wang, Lu, et al.
Published: (2025)
SCORE: Soft Label Compression-Centric Dataset Condensation via Coding Rate Optimization
by: Yuan, Bowen, et al.
Published: (2025)
by: Yuan, Bowen, et al.
Published: (2025)
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection
by: Chen, Zhuoxiao, et al.
Published: (2024)
by: Chen, Zhuoxiao, et al.
Published: (2024)
Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru
by: Cusipuma, Dunant, et al.
Published: (2025)
by: Cusipuma, Dunant, et al.
Published: (2025)
Theoretically Achieving Continuous Representation of Oriented Bounding Boxes
by: Xiao, Zi-Kai, et al.
Published: (2024)
by: Xiao, Zi-Kai, et al.
Published: (2024)
Box-Free Model Watermarks Are Prone to Black-Box Removal Attacks
by: An, Haonan, et al.
Published: (2024)
by: An, Haonan, et al.
Published: (2024)
Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models
by: Huang, Zi-Xuan, et al.
Published: (2024)
by: Huang, Zi-Xuan, et al.
Published: (2024)
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
by: Qian, Zekun, et al.
Published: (2026)
by: Qian, Zekun, et al.
Published: (2026)
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
by: Wang, Zixin, et al.
Published: (2024)
by: Wang, Zixin, et al.
Published: (2024)
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
by: Zhang, Zhizhen, et al.
Published: (2025)
by: Zhang, Zhizhen, et al.
Published: (2025)
Open-CRB: Towards Open World Active Learning for 3D Object Detection
by: Chen, Zhuoxiao, et al.
Published: (2023)
by: Chen, Zhuoxiao, et al.
Published: (2023)
GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Multiple Different Black Box Explanations for Image Classifiers
by: Chockler, Hana, et al.
Published: (2023)
by: Chockler, Hana, et al.
Published: (2023)
Kvasir-VQA: A Text-Image Pair GI Tract Dataset
by: Gautam, Sushant, et al.
Published: (2024)
by: Gautam, Sushant, et al.
Published: (2024)
EMT: A Visual Multi-Task Benchmark Dataset for Autonomous Driving
by: Madjid, Nadya Abdel, et al.
Published: (2025)
by: Madjid, Nadya Abdel, et al.
Published: (2025)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
by: Mo, Ye, et al.
Published: (2025)
by: Mo, Ye, et al.
Published: (2025)
Box6D : Zero-shot Category-level 6D Pose Estimation of Warehouse Boxes
by: Ma, Yintao, et al.
Published: (2025)
by: Ma, Yintao, et al.
Published: (2025)
MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems
by: Yang, Peiru, et al.
Published: (2025)
by: Yang, Peiru, et al.
Published: (2025)
CodeMerge: Codebook-Guided Model Merging for Robust Test-Time Adaptation in Autonomous Driving
by: Yang, Huitong, et al.
Published: (2025)
by: Yang, Huitong, et al.
Published: (2025)
TopoStreamer: Temporal Lane Segment Topology Reasoning in Autonomous Driving
by: Yang, Yiming, et al.
Published: (2025)
by: Yang, Yiming, et al.
Published: (2025)
Improving Black-Box Generative Attacks via Generator Semantic Consistency
by: Jeong, Jongoh, et al.
Published: (2025)
by: Jeong, Jongoh, et al.
Published: (2025)
NBBOX: Noisy Bounding Box Improves Remote Sensing Object Detection
by: Kim, Yechan, et al.
Published: (2024)
by: Kim, Yechan, et al.
Published: (2024)
MPDIoU: A Loss for Efficient and Accurate Bounding Box Regression
by: Ma, Siliang, et al.
Published: (2023)
by: Ma, Siliang, et al.
Published: (2023)
Certified Zeroth-order Black-Box Defense with Robust UNet Denoiser
by: Verma, Astha, et al.
Published: (2023)
by: Verma, Astha, et al.
Published: (2023)
Cross-Stage Coherence in Hierarchical Driving VQA: Explicit Baselines and Learned Gated Context Projectors
by: Jain, Gautam Kumar, et al.
Published: (2026)
by: Jain, Gautam Kumar, et al.
Published: (2026)
DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
by: Fu, Yongjie, et al.
Published: (2024)
by: Fu, Yongjie, et al.
Published: (2024)
Robust Box Prompt based SAM for Medical Image Segmentation
by: Huang, Yuhao, et al.
Published: (2024)
by: Huang, Yuhao, et al.
Published: (2024)
BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
by: Li, Feiran, et al.
Published: (2026)
by: Li, Feiran, et al.
Published: (2026)
ADBA:Approximation Decision Boundary Approach for Black-Box Adversarial Attacks
by: Wang, Feiyang, et al.
Published: (2024)
by: Wang, Feiyang, et al.
Published: (2024)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling
by: Zou, Hongjian, et al.
Published: (2026)
by: Zou, Hongjian, et al.
Published: (2026)
Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
by: Wu, Aodi, et al.
Published: (2025)
by: Wu, Aodi, et al.
Published: (2025)
Finding needles in a haystack: A Black-Box Approach to Invisible Watermark Detection
by: Pan, Minzhou, et al.
Published: (2024)
by: Pan, Minzhou, et al.
Published: (2024)
FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments
by: Malone, Connor, et al.
Published: (2026)
by: Malone, Connor, et al.
Published: (2026)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
by: Han, Hongyong, et al.
Published: (2025)
by: Han, Hongyong, et al.
Published: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
by: Tuong, Nguyen Anh, et al.
Published: (2026)
by: Tuong, Nguyen Anh, et al.
Published: (2026)
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
by: Roy, Parthib, et al.
Published: (2024)
by: Roy, Parthib, et al.
Published: (2024)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
by: Wang, Xingyue, et al.
Published: (2026)
by: Wang, Xingyue, et al.
Published: (2026)
Similar Items
-
Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments
by: Etchegaray, Djamahl, et al.
Published: (2024) -
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
by: Wang, Lu, et al.
Published: (2025) -
SCORE: Soft Label Compression-Centric Dataset Condensation via Coding Rate Optimization
by: Yuan, Bowen, et al.
Published: (2025) -
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection
by: Chen, Zhuoxiao, et al.
Published: (2024) -
Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru
by: Cusipuma, Dunant, et al.
Published: (2025)