What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Inha, Lim, Youngsun, Lee, Seonho, Choi, Jiho, Choe, Junsuk, Shim, Hyunjung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Sampling Bag of Views for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2024)
by: Choi, Hojun, et al.
Published: (2024)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Understanding Multi-Granularity for Open-Vocabulary Part Segmentation
by: Choi, Jiho, et al.
Published: (2024)
by: Choi, Jiho, et al.
Published: (2024)
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
by: Choi, Jiho, et al.
Published: (2025)
by: Choi, Jiho, et al.
Published: (2025)
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval
by: Lim, Youngsun, et al.
Published: (2024)
by: Lim, Youngsun, et al.
Published: (2024)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
by: Lim, Youngsun, et al.
Published: (2024)
by: Lim, Youngsun, et al.
Published: (2024)
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2025)
by: Choi, Hojun, et al.
Published: (2025)
Label-Augmented Dataset Distillation
by: Kang, Seoungyoon, et al.
Published: (2024)
by: Kang, Seoungyoon, et al.
Published: (2024)
Weakly Supervised Semantic Segmentation for Driving Scenes
by: Kim, Dongseob, et al.
Published: (2023)
by: Kim, Dongseob, et al.
Published: (2023)
DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation
by: Kim, Jiwook, et al.
Published: (2024)
by: Kim, Jiwook, et al.
Published: (2024)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather
by: Park, Junsung, et al.
Published: (2025)
by: Park, Junsung, et al.
Published: (2025)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
by: Park, Seojeong, et al.
Published: (2024)
by: Park, Seojeong, et al.
Published: (2024)
Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
Self-Supervised Vision Transformers Are Efficient Segmentation Learners for Imperfect Labels
by: Lee, Seungho, et al.
Published: (2024)
by: Lee, Seungho, et al.
Published: (2024)
Rethinking Direct Preference Optimization in Diffusion Models
by: Kang, Junyong, et al.
Published: (2025)
by: Kang, Junyong, et al.
Published: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
by: Lee, Minhyun, et al.
Published: (2023)
by: Lee, Minhyun, et al.
Published: (2023)
Learning from Spatio-temporal Correlation for Semi-Supervised LiDAR Semantic Segmentation
by: Lee, Seungho, et al.
Published: (2024)
by: Lee, Seungho, et al.
Published: (2024)
Enhancing Multi-Image Understanding through Delimiter Token Scaling
by: Lee, Minyoung, et al.
Published: (2026)
by: Lee, Minyoung, et al.
Published: (2026)
Memory-Efficient Fine-Tuning for Quantized Diffusion Model
by: Ryu, Hyogon, et al.
Published: (2024)
by: Ryu, Hyogon, et al.
Published: (2024)
Grounding Driving VLA via Inverse Kinematics
by: Park, Junsung, et al.
Published: (2026)
by: Park, Junsung, et al.
Published: (2026)
Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
by: Kang, Inha, et al.
Published: (2025)
by: Kang, Inha, et al.
Published: (2025)
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
by: Kim, Dongseob, et al.
Published: (2025)
by: Kim, Dongseob, et al.
Published: (2025)
Precision matters: Precision-aware ensemble for weakly supervised semantic segmentation
by: Park, Junsung, et al.
Published: (2024)
by: Park, Junsung, et al.
Published: (2024)
DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models
by: Ryu, Hyogon, et al.
Published: (2025)
by: Ryu, Hyogon, et al.
Published: (2025)
SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals
by: Yoon, Soyeon, et al.
Published: (2026)
by: Yoon, Soyeon, et al.
Published: (2026)
Mitigating Cross-Image Information Leakage in LVLMs for Multi-Image Tasks
by: Park, Yeji, et al.
Published: (2025)
by: Park, Yeji, et al.
Published: (2025)
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution
by: Kim, Jeongsoo, et al.
Published: (2024)
by: Kim, Jeongsoo, et al.
Published: (2024)
Lossless Token Merging Even Without Fine-Tuning in Vision Transformers
by: Lee, Jaeyeon, et al.
Published: (2025)
by: Lee, Jaeyeon, et al.
Published: (2025)
Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision
by: Jung, Raehyuk, et al.
Published: (2025)
by: Jung, Raehyuk, et al.
Published: (2025)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
by: Kim, Myunsoo, et al.
Published: (2025)
by: Kim, Myunsoo, et al.
Published: (2025)
OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation
by: Hwang, Dongjun, et al.
Published: (2024)
by: Hwang, Dongjun, et al.
Published: (2024)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
by: Lee, Chanhyuk, et al.
Published: (2025)
by: Lee, Chanhyuk, et al.
Published: (2025)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
by: Lee, Jaehoon, et al.
Published: (2026)
by: Lee, Jaehoon, et al.
Published: (2026)
ToSA: Token Merging with Spatial Awareness
by: Huang, Hsiang-Wei, et al.
Published: (2025)
by: Huang, Hsiang-Wei, et al.
Published: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)
by: Zhang, Luyuan, et al.
Published: (2026)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
by: Park, NaHyeon, et al.
Published: (2024)
by: Park, NaHyeon, et al.
Published: (2024)
Similar Items
-
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025) -
Sampling Bag of Views for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2024) -
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
by: Lee, Seonho, et al.
Published: (2024) -
Understanding Multi-Granularity for Open-Vocabulary Part Segmentation
by: Choi, Jiho, et al.
Published: (2024) -
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
by: Choi, Jiho, et al.
Published: (2025)