Sampling Bag of Views for Open-Vocabulary Object Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Choi, Hojun, Choe, Junsuk, Shim, Hyunjung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
di: Kang, Inha, et al.
Pubblicazione: (2025)
di: Kang, Inha, et al.
Pubblicazione: (2025)
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
di: Choi, Hojun, et al.
Pubblicazione: (2025)
di: Choi, Hojun, et al.
Pubblicazione: (2025)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
di: Lim, Youngsun, et al.
Pubblicazione: (2024)
di: Lim, Youngsun, et al.
Pubblicazione: (2024)
Weakly Supervised Semantic Segmentation for Driving Scenes
di: Kim, Dongseob, et al.
Pubblicazione: (2023)
di: Kim, Dongseob, et al.
Pubblicazione: (2023)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
di: Park, NaHyeon, et al.
Pubblicazione: (2025)
di: Park, NaHyeon, et al.
Pubblicazione: (2025)
Precision matters: Precision-aware ensemble for weakly supervised semantic segmentation
di: Park, Junsung, et al.
Pubblicazione: (2024)
di: Park, Junsung, et al.
Pubblicazione: (2024)
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval
di: Lim, Youngsun, et al.
Pubblicazione: (2024)
di: Lim, Youngsun, et al.
Pubblicazione: (2024)
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
di: Kim, Dongseob, et al.
Pubblicazione: (2025)
di: Kim, Dongseob, et al.
Pubblicazione: (2025)
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution
di: Kim, Jeongsoo, et al.
Pubblicazione: (2024)
di: Kim, Jeongsoo, et al.
Pubblicazione: (2024)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
di: Park, Seojeong, et al.
Pubblicazione: (2024)
di: Park, Seojeong, et al.
Pubblicazione: (2024)
Grounding Driving VLA via Inverse Kinematics
di: Park, Junsung, et al.
Pubblicazione: (2026)
di: Park, Junsung, et al.
Pubblicazione: (2026)
Label-Augmented Dataset Distillation
di: Kang, Seoungyoon, et al.
Pubblicazione: (2024)
di: Kang, Seoungyoon, et al.
Pubblicazione: (2024)
Memory-Efficient Fine-Tuning for Quantized Diffusion Model
di: Ryu, Hyogon, et al.
Pubblicazione: (2024)
di: Ryu, Hyogon, et al.
Pubblicazione: (2024)
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
di: Park, Junsung, et al.
Pubblicazione: (2024)
di: Park, Junsung, et al.
Pubblicazione: (2024)
Mitigating Cross-Image Information Leakage in LVLMs for Multi-Image Tasks
di: Park, Yeji, et al.
Pubblicazione: (2025)
di: Park, Yeji, et al.
Pubblicazione: (2025)
Understanding Multi-Granularity for Open-Vocabulary Part Segmentation
di: Choi, Jiho, et al.
Pubblicazione: (2024)
di: Choi, Jiho, et al.
Pubblicazione: (2024)
Open-Vocabulary Object Detection via Language Hierarchy
di: Huang, Jiaxing, et al.
Pubblicazione: (2024)
di: Huang, Jiaxing, et al.
Pubblicazione: (2024)
Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention
di: Juanico, Drandreb Earl O., et al.
Pubblicazione: (2025)
di: Juanico, Drandreb Earl O., et al.
Pubblicazione: (2025)
From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects
di: Li, Zizhao, et al.
Pubblicazione: (2024)
di: Li, Zizhao, et al.
Pubblicazione: (2024)
No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather
di: Park, Junsung, et al.
Pubblicazione: (2025)
di: Park, Junsung, et al.
Pubblicazione: (2025)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
di: An, Na Min, et al.
Pubblicazione: (2025)
di: An, Na Min, et al.
Pubblicazione: (2025)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
di: Lee, Seonho, et al.
Pubblicazione: (2025)
di: Lee, Seonho, et al.
Pubblicazione: (2025)
Mitigating Context Bias in Domain Adaptation for Object Detection using Mask Pooling
di: Son, Hojun, et al.
Pubblicazione: (2025)
di: Son, Hojun, et al.
Pubblicazione: (2025)
Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments
di: Etchegaray, Djamahl, et al.
Pubblicazione: (2024)
di: Etchegaray, Djamahl, et al.
Pubblicazione: (2024)
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
di: Choi, Jiho, et al.
Pubblicazione: (2025)
di: Choi, Jiho, et al.
Pubblicazione: (2025)
Quantifying Context Bias in Domain Adaptation for Object Detection
di: Son, Hojun, et al.
Pubblicazione: (2024)
di: Son, Hojun, et al.
Pubblicazione: (2024)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
di: Zhou, Yang, et al.
Pubblicazione: (2025)
di: Zhou, Yang, et al.
Pubblicazione: (2025)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
di: Park, Yeji, et al.
Pubblicazione: (2024)
di: Park, Yeji, et al.
Pubblicazione: (2024)
VOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking
di: Qian, Zekun, et al.
Pubblicazione: (2024)
di: Qian, Zekun, et al.
Pubblicazione: (2024)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
di: Yu, Seungjun, et al.
Pubblicazione: (2025)
di: Yu, Seungjun, et al.
Pubblicazione: (2025)
SFUOD: Source-Free Unknown Object Detection
di: Park, Keon-Hee, et al.
Pubblicazione: (2025)
di: Park, Keon-Hee, et al.
Pubblicazione: (2025)
Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
di: Yu, Seungjun, et al.
Pubblicazione: (2025)
di: Yu, Seungjun, et al.
Pubblicazione: (2025)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
di: Ahn, Geo, et al.
Pubblicazione: (2026)
di: Ahn, Geo, et al.
Pubblicazione: (2026)
OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation
di: Hwang, Dongjun, et al.
Pubblicazione: (2024)
di: Hwang, Dongjun, et al.
Pubblicazione: (2024)
DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation
di: Kim, Jiwook, et al.
Pubblicazione: (2024)
di: Kim, Jiwook, et al.
Pubblicazione: (2024)
Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy Prediction
di: Yan, Chi, et al.
Pubblicazione: (2025)
di: Yan, Chi, et al.
Pubblicazione: (2025)
I0T: Embedding Standardization Method Towards Zero Modality Gap
di: An, Na Min, et al.
Pubblicazione: (2024)
di: An, Na Min, et al.
Pubblicazione: (2024)
Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph
di: Linok, Sergey, et al.
Pubblicazione: (2024)
di: Linok, Sergey, et al.
Pubblicazione: (2024)
OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision
di: Wang, Junjie, et al.
Pubblicazione: (2024)
di: Wang, Junjie, et al.
Pubblicazione: (2024)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
di: Jiao, Pengkun, et al.
Pubblicazione: (2024)
di: Jiao, Pengkun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
di: Kang, Inha, et al.
Pubblicazione: (2025) -
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
di: Choi, Hojun, et al.
Pubblicazione: (2025) -
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
di: Lim, Youngsun, et al.
Pubblicazione: (2024) -
Weakly Supervised Semantic Segmentation for Driving Scenes
di: Kim, Dongseob, et al.
Pubblicazione: (2023) -
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
di: Park, NaHyeon, et al.
Pubblicazione: (2025)