Structure-aware Prompt Adaptation from Seen to Unseen for Open-Vocabulary Compositional Zero-Shot Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Yihang, Wang, Jiong, Zeng, Pengpeng, Zhang, Ji, Zhao, Lei, Wang, Chong, Song, Jingkuan, Gao, Lianli |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024)
by: Zhang, Haonan, et al.
Published: (2024)
CFReID: Continual Few-shot Person Re-Identification
by: Ni, Hao, et al.
Published: (2025)
by: Ni, Hao, et al.
Published: (2025)
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
by: Cai, Xiao, et al.
Published: (2026)
by: Cai, Xiao, et al.
Published: (2026)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024)
by: Wu, Shihan, et al.
Published: (2024)
ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval
by: Fang, Kaipeng, et al.
Published: (2023)
by: Fang, Kaipeng, et al.
Published: (2023)
DePT: Decoupled Prompt Tuning
by: Zhang, Ji, et al.
Published: (2023)
by: Zhang, Ji, et al.
Published: (2023)
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
by: Cai, Xiao, et al.
Published: (2024)
by: Cai, Xiao, et al.
Published: (2024)
From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
by: Zhang, Ji, et al.
Published: (2023)
by: Zhang, Ji, et al.
Published: (2023)
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
by: Chen, Cheng, et al.
Published: (2026)
by: Chen, Cheng, et al.
Published: (2026)
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
by: Wang, Xuanhan, et al.
Published: (2025)
by: Wang, Xuanhan, et al.
Published: (2025)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Reversible Inversion for Training-Free Exemplar-guided Image Editing
by: Li, Yuke, et al.
Published: (2025)
by: Li, Yuke, et al.
Published: (2025)
Sim-and-Human Co-training for Data-Efficient and Generalizable Robotic Manipulation
by: Fang, Kaipeng, et al.
Published: (2026)
by: Fang, Kaipeng, et al.
Published: (2026)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
by: Guo, Jiaqi, et al.
Published: (2024)
by: Guo, Jiaqi, et al.
Published: (2024)
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
by: Hong, Yu, et al.
Published: (2025)
by: Hong, Yu, et al.
Published: (2025)
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
by: Su, Sitong, et al.
Published: (2023)
by: Su, Sitong, et al.
Published: (2023)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
by: Liu, Ke, et al.
Published: (2025)
by: Liu, Ke, et al.
Published: (2025)
Any Target Can be Offense: Adversarial Example Generation via Generalized Latent Infection
by: Sun, Youheng, et al.
Published: (2024)
by: Sun, Youheng, et al.
Published: (2024)
GT23D-Bench: A Comprehensive General Text-to-3D Generation Benchmark
by: Cai, Xiao, et al.
Published: (2024)
by: Cai, Xiao, et al.
Published: (2024)
Reliable Few-shot Learning under Dual Noises
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation
by: Zhu, Junhong, et al.
Published: (2026)
by: Zhu, Junhong, et al.
Published: (2026)
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
by: Luo, Xu, et al.
Published: (2026)
by: Luo, Xu, et al.
Published: (2026)
Debiased Orthogonal Boundary-Driven Efficient Noise Mitigation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages
by: Mullov, Carlos, et al.
Published: (2024)
by: Mullov, Carlos, et al.
Published: (2024)
Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
by: Fan, Xiaomeng, et al.
Published: (2025)
by: Fan, Xiaomeng, et al.
Published: (2025)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
by: Wang, Xuanhan, et al.
Published: (2025)
by: Wang, Xuanhan, et al.
Published: (2025)
Zero-Shot Adaptation of Behavioral Foundation Models to Unseen Dynamics
by: Bobrin, Maksim, et al.
Published: (2025)
by: Bobrin, Maksim, et al.
Published: (2025)
Perceptions of the Seen and the Unseen World
by: Michaela Schäuble
Published: (2020)
by: Michaela Schäuble
Published: (2020)
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
by: Xu, Wenhao, et al.
Published: (2023)
by: Xu, Wenhao, et al.
Published: (2023)
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
by: Yin, Xiaoran, et al.
Published: (2025)
by: Yin, Xiaoran, et al.
Published: (2025)
Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting
by: Jeon, Inseok, et al.
Published: (2026)
by: Jeon, Inseok, et al.
Published: (2026)
FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
by: Yuan, Shengming, et al.
Published: (2025)
by: Yuan, Shengming, et al.
Published: (2025)
Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
by: Lei, Ting, et al.
Published: (2025)
by: Lei, Ting, et al.
Published: (2025)
AICL: Action In-Context Learning for Video Diffusion Model
by: Liu, Jianzhi, et al.
Published: (2024)
by: Liu, Jianzhi, et al.
Published: (2024)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
by: Lyu, Xinyu, et al.
Published: (2024)
by: Lyu, Xinyu, et al.
Published: (2024)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
DOZE: A Dataset for Open-Vocabulary Zero-Shot Object Navigation in Dynamic Environments
by: Ma, Ji, et al.
Published: (2024)
by: Ma, Ji, et al.
Published: (2024)
Similar Items
-
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024) -
CFReID: Continual Few-shot Person Re-Identification
by: Ni, Hao, et al.
Published: (2025) -
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
by: Cai, Xiao, et al.
Published: (2026) -
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024) -
ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval
by: Fang, Kaipeng, et al.
Published: (2023)