PSVMA+: Exploring Multi-granularity Semantic-visual Adaption for Generalized Zero-shot Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Man, Bai, Huihui, Li, Feng, Zhang, Chunjie, Wei, Yunchao, Wang, Meng, Chua, Tat-Seng, Zhao, Yao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
von: Liu, Man, et al.
Veröffentlicht: (2024)
von: Liu, Man, et al.
Veröffentlicht: (2024)
Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
von: Li, Peihao, et al.
Veröffentlicht: (2025)
von: Li, Peihao, et al.
Veröffentlicht: (2025)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
Universal Scene Graph Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Region-Adaptive Transform with Segmentation Prior for Image Compression
von: Liu, Yuxi, et al.
Veröffentlicht: (2024)
von: Liu, Yuxi, et al.
Veröffentlicht: (2024)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
von: Zhang, An, et al.
Veröffentlicht: (2024)
von: Zhang, An, et al.
Veröffentlicht: (2024)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
von: Chu, Meng, et al.
Veröffentlicht: (2023)
von: Chu, Meng, et al.
Veröffentlicht: (2023)
Extending Visual Dynamics for Video-to-Music Generation
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
von: He, Jinghan, et al.
Veröffentlicht: (2026)
von: He, Jinghan, et al.
Veröffentlicht: (2026)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
Bridge the Points: Graph-based Few-shot Segment Anything Semantically
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
IPSeg: Image Posterior Mitigates Semantic Drift in Class-Incremental Segmentation
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
von: Ren, Zhongwei, et al.
Veröffentlicht: (2025)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2025)
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Enabling Generalized Zero-shot Learning Towards Unseen Domains by Intrinsic Learning from Redundant LLM Semantics
von: Yue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Yue, Jiaqi, et al.
Veröffentlicht: (2024)
Audio-visual Generalized Zero-shot Learning the Easy Way
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection
von: Wei, Yao, et al.
Veröffentlicht: (2026)
von: Wei, Yao, et al.
Veröffentlicht: (2026)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
Instilling Multi-round Thinking to Text-guided Image Generation
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
Extremely Simple Out-of-distribution Detection for Audio-visual Generalized Zero-shot Learning
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective
von: Zheng, Zhedong, et al.
Veröffentlicht: (2022)
von: Zheng, Zhedong, et al.
Veröffentlicht: (2022)
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
von: Zhu, Hongguang, et al.
Veröffentlicht: (2025)
von: Zhu, Hongguang, et al.
Veröffentlicht: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
von: Qi, Ji, et al.
Veröffentlicht: (2025)
von: Qi, Ji, et al.
Veröffentlicht: (2025)
Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning
von: Pu, Haojie, et al.
Veröffentlicht: (2026)
von: Pu, Haojie, et al.
Veröffentlicht: (2026)
Prioritized Semantic Learning for Zero-shot Instance Navigation
von: Sun, Xinyu, et al.
Veröffentlicht: (2024)
von: Sun, Xinyu, et al.
Veröffentlicht: (2024)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis
von: Lai, Haoran, et al.
Veröffentlicht: (2025)
von: Lai, Haoran, et al.
Veröffentlicht: (2025)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
von: Liu, Guimeng, et al.
Veröffentlicht: (2025)
von: Liu, Guimeng, et al.
Veröffentlicht: (2025)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
von: Liu, Han, et al.
Veröffentlicht: (2025)
von: Liu, Han, et al.
Veröffentlicht: (2025)
Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation
von: Zhang, Bingfeng, et al.
Veröffentlicht: (2024)
von: Zhang, Bingfeng, et al.
Veröffentlicht: (2024)
Towards Modality Generalization: A Benchmark and Prospective Analysis
von: Liu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2024)
Boosting Audio-visual Zero-shot Learning with Large Language Models
von: Chen, Haoxing, et al.
Veröffentlicht: (2023)
von: Chen, Haoxing, et al.
Veröffentlicht: (2023)
Generalizable Semantic Vision Query Generation for Zero-shot Panoptic and Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2024)
von: Chen, Jialei, et al.
Veröffentlicht: (2024)
InstantID: Zero-shot Identity-Preserving Generation in Seconds
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
von: Wang, Qixun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
von: Liu, Man, et al.
Veröffentlicht: (2024) -
Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
von: Li, Peihao, et al.
Veröffentlicht: (2025) -
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025) -
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
von: Li, Juncheng, et al.
Veröffentlicht: (2023) -
Universal Scene Graph Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)