Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
Fuente:
arXiv
Salvato in:
| Autori principali: | Pei, Gensheng, Chen, Tao, Wang, Yujia, Cai, Xinhao, Shu, Xiangbo, Zhou, Tianfei, Yao, Yazhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficiency Follows Global-Local Decoupling
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation
di: Pei, Gensheng, et al.
Pubblicazione: (2026)
di: Pei, Gensheng, et al.
Pubblicazione: (2026)
PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation
di: Yin, Jianjian, et al.
Pubblicazione: (2026)
di: Yin, Jianjian, et al.
Pubblicazione: (2026)
Towards Remote Sensing Change Detection with Neural Memory
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
Beyond Quadratic: Linear-Time Change Detection with RWKV
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)
Taming SAM3 in the Wild: A Concept Bank for Open-Vocabulary Segmentation
di: Pei, Gensheng, et al.
Pubblicazione: (2026)
di: Pei, Gensheng, et al.
Pubblicazione: (2026)
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection
di: Cai, Xinhao, et al.
Pubblicazione: (2026)
di: Cai, Xinhao, et al.
Pubblicazione: (2026)
Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation
di: Cai, Xinhao, et al.
Pubblicazione: (2026)
di: Cai, Xinhao, et al.
Pubblicazione: (2026)
VideoMAC: Video Masked Autoencoders Meet ConvNets
di: Pei, Gensheng, et al.
Pubblicazione: (2024)
di: Pei, Gensheng, et al.
Pubblicazione: (2024)
Knowledge Transfer with Simulated Inter-Image Erasing for Weakly Supervised Semantic Segmentation
di: Chen, Tao, et al.
Pubblicazione: (2024)
di: Chen, Tao, et al.
Pubblicazione: (2024)
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
di: Yin, Jianjian, et al.
Pubblicazione: (2025)
di: Yin, Jianjian, et al.
Pubblicazione: (2025)
FTMoMamba: Motion Generation with Frequency and Text State Space Models
di: Li, Chengjian, et al.
Pubblicazione: (2024)
di: Li, Chengjian, et al.
Pubblicazione: (2024)
Relating CNN-Transformer Fusion Network for Change Detection
di: Gao, Yuhao, et al.
Pubblicazione: (2024)
di: Gao, Yuhao, et al.
Pubblicazione: (2024)
Dynamic in Static: Hybrid Visual Correspondence for Self-Supervised Video Object Segmentation
di: Pei, Gensheng, et al.
Pubblicazione: (2024)
di: Pei, Gensheng, et al.
Pubblicazione: (2024)
Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
di: Zhou, Bo, et al.
Pubblicazione: (2026)
di: Zhou, Bo, et al.
Pubblicazione: (2026)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
di: Zhang, Wang, et al.
Pubblicazione: (2024)
di: Zhang, Wang, et al.
Pubblicazione: (2024)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
di: Chang, Boyu, et al.
Pubblicazione: (2026)
di: Chang, Boyu, et al.
Pubblicazione: (2026)
OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
di: Qu, Hongyu, et al.
Pubblicazione: (2025)
di: Qu, Hongyu, et al.
Pubblicazione: (2025)
Poly Kernel Inception Network for Remote Sensing Detection
di: Cai, Xinhao, et al.
Pubblicazione: (2024)
di: Cai, Xinhao, et al.
Pubblicazione: (2024)
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
di: Ni, Ziqi, et al.
Pubblicazione: (2025)
di: Ni, Ziqi, et al.
Pubblicazione: (2025)
AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity Recognition
di: Cao, Meiqi, et al.
Pubblicazione: (2024)
di: Cao, Meiqi, et al.
Pubblicazione: (2024)
All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning
di: Yang, Zheng, et al.
Pubblicazione: (2025)
di: Yang, Zheng, et al.
Pubblicazione: (2025)
On-Road Object Importance Estimation: A New Dataset and A Model with Multi-Fold Top-Down Guidance
di: Nan, Zhixiong, et al.
Pubblicazione: (2024)
di: Nan, Zhixiong, et al.
Pubblicazione: (2024)
Diffusion Feedback Helps CLIP See Better
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
di: Fang, Hao, et al.
Pubblicazione: (2024)
di: Fang, Hao, et al.
Pubblicazione: (2024)
What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality
di: Wang, Zihan, et al.
Pubblicazione: (2024)
di: Wang, Zihan, et al.
Pubblicazione: (2024)
Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning
di: Sun, Hao, et al.
Pubblicazione: (2026)
di: Sun, Hao, et al.
Pubblicazione: (2026)
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
di: Zhang, Dengke, et al.
Pubblicazione: (2024)
di: Zhang, Dengke, et al.
Pubblicazione: (2024)
Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset Distillation
di: Zhong, Xinhao, et al.
Pubblicazione: (2024)
di: Zhong, Xinhao, et al.
Pubblicazione: (2024)
When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP
di: Chu, Beilin, et al.
Pubblicazione: (2025)
di: Chu, Beilin, et al.
Pubblicazione: (2025)
FrameOracle: Learning What to See and How Much to See in Videos
di: Li, Chaoyu, et al.
Pubblicazione: (2025)
di: Li, Chaoyu, et al.
Pubblicazione: (2025)
Kernel-Aware Graph Prompt Learning for Few-Shot Anomaly Detection
di: Tao, Fenfang, et al.
Pubblicazione: (2024)
di: Tao, Fenfang, et al.
Pubblicazione: (2024)
C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake Detection
di: Tan, Chuangchuang, et al.
Pubblicazione: (2024)
di: Tan, Chuangchuang, et al.
Pubblicazione: (2024)
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches
di: Wu, Cheng-En, et al.
Pubblicazione: (2024)
di: Wu, Cheng-En, et al.
Pubblicazione: (2024)
UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
di: Dai, Guangzhao, et al.
Pubblicazione: (2024)
di: Dai, Guangzhao, et al.
Pubblicazione: (2024)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
di: Bai, Jiawang, et al.
Pubblicazione: (2023)
di: Bai, Jiawang, et al.
Pubblicazione: (2023)
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
di: Wang, Zhichuan, et al.
Pubblicazione: (2025)
di: Wang, Zhichuan, et al.
Pubblicazione: (2025)
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
di: Corvi, Riccardo, et al.
Pubblicazione: (2025)
di: Corvi, Riccardo, et al.
Pubblicazione: (2025)
MedSAD-CLIP: Supervised CLIP with Token-Patch Cross-Attention for Medical Anomaly Detection and Segmentation
di: Tran, Thuy Truong, et al.
Pubblicazione: (2026)
di: Tran, Thuy Truong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Efficiency Follows Global-Local Decoupling
di: Yang, Zhenyu, et al.
Pubblicazione: (2026) -
PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation
di: Pei, Gensheng, et al.
Pubblicazione: (2026) -
PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation
di: Yin, Jianjian, et al.
Pubblicazione: (2026) -
Towards Remote Sensing Change Detection with Neural Memory
di: Yang, Zhenyu, et al.
Pubblicazione: (2026) -
Beyond Quadratic: Linear-Time Change Detection with RWKV
di: Yang, Zhenyu, et al.
Pubblicazione: (2026)