Improving Open-World Object Localization by Discovering Background
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Ashish, Jones, Michael J., Peng, Kuan-Chuan, Cherian, Anoop, Chatterjee, Moitreya, Learned-Miller, Erik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Guided Agentic Object Detection for Open-World Understanding
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
by: Li, Danrui, et al.
Published: (2026)
by: Li, Danrui, et al.
Published: (2026)
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
by: Xiang, Xinhao, et al.
Published: (2025)
by: Xiang, Xinhao, et al.
Published: (2025)
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
ComplexVAD: Detecting Interaction Anomalies in Video
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models
by: Mumcu, Furkan, et al.
Published: (2026)
by: Mumcu, Furkan, et al.
Published: (2026)
DCVNet: Dilated Cost Volume Networks for Fast Optical Flow
by: Jiang, Huaizu, et al.
Published: (2021)
by: Jiang, Huaizu, et al.
Published: (2021)
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
by: Cherian, Anoop, et al.
Published: (2025)
by: Cherian, Anoop, et al.
Published: (2025)
MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
by: Kogashi, Kaen, et al.
Published: (2025)
by: Kogashi, Kaen, et al.
Published: (2025)
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2024)
by: Li, Jiaming, et al.
Published: (2024)
Multimodal 3D Object Detection on Unseen Domains
by: Hegde, Deepti, et al.
Published: (2024)
by: Hegde, Deepti, et al.
Published: (2024)
Equivariant Spatio-Temporal Self-Supervision for LiDAR Object Detection
by: Hegde, Deepti, et al.
Published: (2024)
by: Hegde, Deepti, et al.
Published: (2024)
Programmatic Video Prediction Using Large Language Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Joint Training of Image Generator and Detector for Road Defect Detection
by: Peng, Kuan-Chuan
Published: (2025)
by: Peng, Kuan-Chuan
Published: (2025)
Towards Zero-shot 3D Anomaly Localization
by: Wang, Yizhou, et al.
Published: (2024)
by: Wang, Yizhou, et al.
Published: (2024)
A Probability-guided Sampler for Neural Implicit Surface Rendering
by: Pais, Gonçalo Dias, et al.
Published: (2025)
by: Pais, Gonçalo Dias, et al.
Published: (2025)
FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations
by: Sawada, Naoko, et al.
Published: (2025)
by: Sawada, Naoko, et al.
Published: (2025)
Auto-Vocabulary 3D Object Detection
by: Zhang, Haomeng, et al.
Published: (2025)
by: Zhang, Haomeng, et al.
Published: (2025)
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
by: Kundu, Sanjoy, et al.
Published: (2023)
by: Kundu, Sanjoy, et al.
Published: (2023)
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
by: Ding, Tianye, et al.
Published: (2025)
by: Ding, Tianye, et al.
Published: (2025)
The Spatio-Temporal Poisson Point Process: A Simple Model for the Alignment of Event Camera Data
by: Gu, Cheng, et al.
Published: (2021)
by: Gu, Cheng, et al.
Published: (2021)
Semi-supervised Open-World Object Detection
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Improving Pre-trained Self-Supervised Embeddings Through Effective Entropy Maximization
by: Chakraborty, Deep, et al.
Published: (2024)
by: Chakraborty, Deep, et al.
Published: (2024)
Boosting Open-Vocabulary Object Detection by Handling Background Samples
by: Zeng, Ruizhe, et al.
Published: (2024)
by: Zeng, Ruizhe, et al.
Published: (2024)
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
by: Ni, Yao, et al.
Published: (2025)
by: Ni, Yao, et al.
Published: (2025)
FLIGHT: Fibonacci Lattice-based Inference for Geometric Heading in real-Time
by: Dirnfeld, David, et al.
Published: (2026)
by: Dirnfeld, David, et al.
Published: (2026)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
by: Lai, Yung-Hsuan, et al.
Published: (2025)
by: Lai, Yung-Hsuan, et al.
Published: (2025)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
by: Zhang, Jiahao, et al.
Published: (2023)
by: Zhang, Jiahao, et al.
Published: (2023)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
Open World Object Detection: A Survey
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Developing Gridded Emission Inventory from High-Resolution Satellite Object Detection for Improved Air Quality Forecasts
by: Ghosal, Shubham, et al.
Published: (2024)
by: Ghosal, Shubham, et al.
Published: (2024)
Unsupervised Open-Vocabulary Object Localization in Videos
by: Fan, Ke, et al.
Published: (2023)
by: Fan, Ke, et al.
Published: (2023)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Long-Tailed Anomaly Detection with Learnable Class Names
by: Ho, Chih-Hui, et al.
Published: (2024)
by: Ho, Chih-Hui, et al.
Published: (2024)
Toward Long-Tailed Online Anomaly Detection through Class-Agnostic Concepts
by: Yang, Chiao-An, et al.
Published: (2025)
by: Yang, Chiao-An, et al.
Published: (2025)
Similar Items
-
LLM-Guided Agentic Object Detection for Open-World Understanding
by: Mumcu, Furkan, et al.
Published: (2025) -
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
by: Li, Danrui, et al.
Published: (2026) -
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
by: Xiang, Xinhao, et al.
Published: (2025) -
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
by: Mumcu, Furkan, et al.
Published: (2025) -
ComplexVAD: Detecting Interaction Anomalies in Video
by: Mumcu, Furkan, et al.
Published: (2025)