Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xiaoran, Yang, Jiangang, Chong, Wenyue, Shi, Wenhui, Sun, Shichu, Xing, Jing, Liu, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection
by: Xu, Xiaoran, et al.
Published: (2024)
by: Xu, Xiaoran, et al.
Published: (2024)
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
by: Xu, Xiaoran, et al.
Published: (2026)
by: Xu, Xiaoran, et al.
Published: (2026)
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
by: Luo, Luqing, et al.
Published: (2024)
by: Luo, Luqing, et al.
Published: (2024)
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
by: Xue, Weiying, et al.
Published: (2024)
by: Xue, Weiying, et al.
Published: (2024)
Towards Robust Semantic Correspondence: A Benchmark and Insights
by: Chong, Wenyue
Published: (2025)
by: Chong, Wenyue
Published: (2025)
Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection
by: Hu, Yupeng, et al.
Published: (2025)
by: Hu, Yupeng, et al.
Published: (2025)
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
by: Chen, Wenyue, et al.
Published: (2026)
by: Chen, Wenyue, et al.
Published: (2026)
Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs
by: Cai, Zhaolin, et al.
Published: (2025)
by: Cai, Zhaolin, et al.
Published: (2025)
Language Prompt vs. Image Enhancement: Boosting Object Detection With CLIP in Hazy Environments
by: Pang, Jian, et al.
Published: (2026)
by: Pang, Jian, et al.
Published: (2026)
Contextualized Representation Learning for Effective Human-Object Interaction Detection
by: Li, Zhehao, et al.
Published: (2025)
by: Li, Zhehao, et al.
Published: (2025)
Boosting Open-Vocabulary Object Detection by Handling Background Samples
by: Zeng, Ruizhe, et al.
Published: (2024)
by: Zeng, Ruizhe, et al.
Published: (2024)
Boosting 3D Object Detection with Semantic-Aware Multi-Branch Framework
by: Jing, Hao, et al.
Published: (2024)
by: Jing, Hao, et al.
Published: (2024)
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
by: Kogilathota, Sai Akhil, et al.
Published: (2026)
by: Kogilathota, Sai Akhil, et al.
Published: (2026)
Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
by: Zhou, Zirun, et al.
Published: (2025)
by: Zhou, Zirun, et al.
Published: (2025)
Spatial-Temporal Human-Object Interaction Detection
by: Sun, Xu, et al.
Published: (2025)
by: Sun, Xu, et al.
Published: (2025)
Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation
by: Li, Lin, et al.
Published: (2025)
by: Li, Lin, et al.
Published: (2025)
Boosting Salient Object Detection with Knowledge Distillated from Large Foundation Models
by: He, Miaoyang, et al.
Published: (2025)
by: He, Miaoyang, et al.
Published: (2025)
Adaptive Event Stream Slicing for Open-Vocabulary Event-Based Object Detection via Vision-Language Knowledge Distillation
by: Zhang, Jinchang, et al.
Published: (2025)
by: Zhang, Jinchang, et al.
Published: (2025)
Tiny Object Detection with Single Point Supervision
by: Zhu, Haoran, et al.
Published: (2024)
by: Zhu, Haoran, et al.
Published: (2024)
Frequency-Spatial Entanglement Learning for Camouflaged Object Detection
by: Sun, Yanguang, et al.
Published: (2024)
by: Sun, Yanguang, et al.
Published: (2024)
Boosting Object Detection with Zero-Shot Day-Night Domain Adaptation
by: Du, Zhipeng, et al.
Published: (2023)
by: Du, Zhipeng, et al.
Published: (2023)
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
by: Jiang, Zhangqi, et al.
Published: (2024)
by: Jiang, Zhangqi, et al.
Published: (2024)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
by: Seo, Soo Won, et al.
Published: (2026)
by: Seo, Soo Won, et al.
Published: (2026)
Language-Driven Dual Style Mixing for Single-Domain Generalized Object Detection
by: Qin, Hongda, et al.
Published: (2025)
by: Qin, Hongda, et al.
Published: (2025)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)
by: Bao, Chen, et al.
Published: (2024)
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
by: Zhang, Songyan, et al.
Published: (2024)
by: Zhang, Songyan, et al.
Published: (2024)
Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
by: Deng, Huilin, et al.
Published: (2025)
by: Deng, Huilin, et al.
Published: (2025)
SOOD++: Leveraging Unlabeled Data to Boost Oriented Object Detection
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
MOSE: Boosting Vision-based Roadside 3D Object Detection with Scene Cues
by: Chen, Xiahan, et al.
Published: (2024)
by: Chen, Xiahan, et al.
Published: (2024)
Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
by: Li, Lin, et al.
Published: (2025)
by: Li, Lin, et al.
Published: (2025)
Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection
by: Wang, Xuquan, et al.
Published: (2026)
by: Wang, Xuquan, et al.
Published: (2026)
RadarDistill: Boosting Radar-based Object Detection Performance via Knowledge Distillation from LiDAR Features
by: Bang, Geonho, et al.
Published: (2024)
by: Bang, Geonho, et al.
Published: (2024)
BoostRad: Enhancing Object Detection by Boosting Radar Reflections
by: Haitman, Yuval, et al.
Published: (2024)
by: Haitman, Yuval, et al.
Published: (2024)
Decomposing and Composing: Towards Efficient Vision-Language Continual Learning via Rank-1 Expert Pool in a Single LoRA
by: Fa, Zhan, et al.
Published: (2026)
by: Fa, Zhan, et al.
Published: (2026)
Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport
by: Zheng, Taoran, et al.
Published: (2025)
by: Zheng, Taoran, et al.
Published: (2025)
GEM: Boost Simple Network for Glass Surface Segmentation via Vision Foundation Models
by: Hao, Jing, et al.
Published: (2023)
by: Hao, Jing, et al.
Published: (2023)
RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation
by: Chu, Wenhui
Published: (2026)
by: Chu, Wenhui
Published: (2026)
Similar Items
-
PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection
by: Xu, Xiaoran, et al.
Published: (2024) -
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
by: Xu, Xiaoran, et al.
Published: (2026) -
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
by: Luo, Luqing, et al.
Published: (2024) -
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
by: Xue, Weiying, et al.
Published: (2024) -
Towards Robust Semantic Correspondence: A Benchmark and Insights
by: Chong, Wenyue
Published: (2025)