AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yidan, Zhuang, Chenyi, Liu, Wutao, Gao, Pan, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025)
by: Zuo, Zhi, et al.
Published: (2025)
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023)
by: Xu, Xiaoxu, et al.
Published: (2023)
Loomis Painter: Reconstructing the Painting Process
by: Pobitzer, Markus, et al.
Published: (2025)
by: Pobitzer, Markus, et al.
Published: (2025)
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
by: Shu, Yan, et al.
Published: (2026)
by: Shu, Yan, et al.
Published: (2026)
PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models
by: Xie, Chang, et al.
Published: (2025)
by: Xie, Chang, et al.
Published: (2025)
Prototypical Hash Encoding for On-the-Fly Fine-Grained Category Discovery
by: Zheng, Haiyang, et al.
Published: (2024)
by: Zheng, Haiyang, et al.
Published: (2024)
Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts
by: Zheng, Haiyang, et al.
Published: (2025)
by: Zheng, Haiyang, et al.
Published: (2025)
Democratizing Fine-grained Visual Recognition with Large Language Models
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
by: Guo, Hao, et al.
Published: (2025)
by: Guo, Hao, et al.
Published: (2025)
Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery
by: Zheng, Haiyang, et al.
Published: (2024)
by: Zheng, Haiyang, et al.
Published: (2024)
Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category Discovery
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
by: Xu, Xiaoxu, et al.
Published: (2024)
by: Xu, Xiaoxu, et al.
Published: (2024)
Wasserstein-Aligned Hyperbolic Multi-View Clustering
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-labeling
by: Wang, Xu, et al.
Published: (2024)
by: Wang, Xu, et al.
Published: (2024)
3D Part Segmentation via Geometric Aggregation of 2D Visual Features
by: Garosi, Marco, et al.
Published: (2024)
by: Garosi, Marco, et al.
Published: (2024)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
by: Lobba, Davide, et al.
Published: (2025)
by: Lobba, Davide, et al.
Published: (2025)
First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
by: Liu, Wutao, et al.
Published: (2025)
by: Liu, Wutao, et al.
Published: (2025)
Training and Tuning Generative Neural Radiance Fields for Attribute-Conditional 3D-Aware Face Generation
by: Zhang, Jichao, et al.
Published: (2022)
by: Zhang, Jichao, et al.
Published: (2022)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
by: Zheng, Haiyang, et al.
Published: (2025)
by: Zheng, Haiyang, et al.
Published: (2025)
3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
by: Xu, Xiaoxu, et al.
Published: (2025)
by: Xu, Xiaoxu, et al.
Published: (2025)
Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models Function
by: Zhuang, Chenyi, et al.
Published: (2024)
by: Zhuang, Chenyi, et al.
Published: (2024)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
by: Huy, Ta Duc, et al.
Published: (2025)
by: Huy, Ta Duc, et al.
Published: (2025)
Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
by: Dong, Jiahua, et al.
Published: (2025)
by: Dong, Jiahua, et al.
Published: (2025)
Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
by: Gao, Yongbiao, et al.
Published: (2024)
by: Gao, Yongbiao, et al.
Published: (2024)
Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning
by: Pu, Haojie, et al.
Published: (2026)
by: Pu, Haojie, et al.
Published: (2026)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
by: Song, Yue, et al.
Published: (2023)
by: Song, Yue, et al.
Published: (2023)
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
by: Hu, Ying, et al.
Published: (2024)
by: Hu, Ying, et al.
Published: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
by: Wang, Weijie, et al.
Published: (2022)
by: Wang, Weijie, et al.
Published: (2022)
Reverse Personalization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
by: Chen, Bolei, et al.
Published: (2025)
by: Chen, Bolei, et al.
Published: (2025)
Prompt Categories Cluster for Weakly Supervised Semantic Segmentation
by: Wu, Wangyu, et al.
Published: (2024)
by: Wu, Wangyu, et al.
Published: (2024)
Diff9D: Diffusion-Based Domain-Generalized Category-Level 9-DoF Object Pose Estimation
by: Liu, Jian, et al.
Published: (2025)
by: Liu, Jian, et al.
Published: (2025)
Unified Unsupervised Salient Object Detection via Knowledge Transfer
by: Yuan, Yao, et al.
Published: (2024)
by: Yuan, Yao, et al.
Published: (2024)
Similar Items
-
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025) -
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
by: Li, Xiaoqi, et al.
Published: (2025) -
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023) -
Loomis Painter: Reconstructing the Painting Process
by: Pobitzer, Markus, et al.
Published: (2025) -
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
by: Shu, Yan, et al.
Published: (2026)