Road Obstacle Video Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Rai, Shyam Nandan, Karthik, Shyamgopal, Georgescu, Mariana-Iuliana, Caputo, Barbara, Masone, Carlo, Akata, Zeynep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
FLAIR: VLM with Fine-grained Language-informed Image Representations
by: Xiao, Rui, et al.
Published: (2024)
by: Xiao, Rui, et al.
Published: (2024)
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
by: Eyring, Luca, et al.
Published: (2024)
by: Eyring, Luca, et al.
Published: (2024)
PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation
by: Rosi, Gabriele, et al.
Published: (2026)
by: Rosi, Gabriele, et al.
Published: (2026)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
by: Eyring, Luca, et al.
Published: (2025)
by: Eyring, Luca, et al.
Published: (2025)
Scalable Ranked Preference Optimization for Text-to-Image Generation
by: Karthik, Shyamgopal, et al.
Published: (2024)
by: Karthik, Shyamgopal, et al.
Published: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
X-Aligner: Composed Visual Retrieval without the Bells and Whistles
by: Zheng, Yuqian, et al.
Published: (2026)
by: Zheng, Yuqian, et al.
Published: (2026)
Towards Safer and Understandable Driver Intention Prediction
by: Karuppasamy, Mukilan, et al.
Published: (2025)
by: Karuppasamy, Mukilan, et al.
Published: (2025)
JIST: Joint Image and Sequence Training for Sequential Visual Place Recognition
by: Berton, Gabriele, et al.
Published: (2024)
by: Berton, Gabriele, et al.
Published: (2024)
EarthLoc: Astronaut Photography Localization by Indexing Earth from Space
by: Berton, Gabriele, et al.
Published: (2024)
by: Berton, Gabriele, et al.
Published: (2024)
The Unreasonable Effectiveness of Pre-Trained Features for Camera Pose Refinement
by: Trivigno, Gabriele, et al.
Published: (2024)
by: Trivigno, Gabriele, et al.
Published: (2024)
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
Post-hoc Probabilistic Vision-Language Models
by: Baumann, Anton, et al.
Published: (2024)
by: Baumann, Anton, et al.
Published: (2024)
EarthMatch: Iterative Coregistration for Fine-grained Localization of Astronaut Photography
by: Berton, Gabriele, et al.
Published: (2024)
by: Berton, Gabriele, et al.
Published: (2024)
MeshVPR: Citywide Visual Place Recognition Using 3D Meshes
by: Berton, Gabriele, et al.
Published: (2024)
by: Berton, Gabriele, et al.
Published: (2024)
Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
by: Dalmonte, Francesco, et al.
Published: (2025)
by: Dalmonte, Francesco, et al.
Published: (2025)
LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
by: Cuttano, Claudia, et al.
Published: (2024)
by: Cuttano, Claudia, et al.
Published: (2024)
Subspace-Boosted Model Merging
by: Skorobogat, Ronald, et al.
Published: (2025)
by: Skorobogat, Ronald, et al.
Published: (2025)
MegaLoc: One Retrieval to Place Them All
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
by: Cuttano, Claudia, et al.
Published: (2025)
by: Cuttano, Claudia, et al.
Published: (2025)
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Simplifying Knowledge Transfer in Pretrained Models
by: Jain, Siddharth, et al.
Published: (2025)
by: Jain, Siddharth, et al.
Published: (2025)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
AstroLoc: Robust Space to Ground Image Localizer
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
All You Need to Know About Training Image Retrieval Models
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
INSID3: Training-Free In-Context Segmentation with DINOv3
by: Cuttano, Claudia, et al.
Published: (2026)
by: Cuttano, Claudia, et al.
Published: (2026)
Audiovisual Masked Autoencoders
by: Georgescu, Mariana-Iuliana, et al.
Published: (2022)
by: Georgescu, Mariana-Iuliana, et al.
Published: (2022)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
To Match or Not to Match: Revisiting Image Matching for Reliable Visual Place Recognition
by: Sferrazza, Davide, et al.
Published: (2025)
by: Sferrazza, Davide, et al.
Published: (2025)
MARCO: Navigating the Unseen Space of Semantic Correspondence
by: Cuttano, Claudia, et al.
Published: (2026)
by: Cuttano, Claudia, et al.
Published: (2026)
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025)
by: Kim, Jae Myung, et al.
Published: (2025)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
Segment-Level Road Obstacle Detection Using Visual Foundation Model Priors and Likelihood Ratios
by: Shoeb, Youssef, et al.
Published: (2024)
by: Shoeb, Youssef, et al.
Published: (2024)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
Similar Items
-
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024) -
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023) -
FLAIR: VLM with Fine-grained Language-informed Image Representations
by: Xiao, Rui, et al.
Published: (2024) -
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
by: Eyring, Luca, et al.
Published: (2024) -
PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation
by: Rosi, Gabriele, et al.
Published: (2026)