Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Zijie, Hu, Zhenzhen, Hong, Richang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
von: Song, Zijie, et al.
Veröffentlicht: (2025)
von: Song, Zijie, et al.
Veröffentlicht: (2025)
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022)
von: Song, Zijie, et al.
Veröffentlicht: (2022)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
von: Xu, Yunbo, et al.
Veröffentlicht: (2025)
von: Xu, Yunbo, et al.
Veröffentlicht: (2025)
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
Iterative Adversarial Attack on Image-guided Story Ending Generation
von: Wang, Youze, et al.
Veröffentlicht: (2023)
von: Wang, Youze, et al.
Veröffentlicht: (2023)
Jigsaw Regularization in Whole-Slide Image Classification
von: Jeong, So Won, et al.
Veröffentlicht: (2026)
von: Jeong, So Won, et al.
Veröffentlicht: (2026)
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
von: Chen, Yin, et al.
Veröffentlicht: (2024)
von: Chen, Yin, et al.
Veröffentlicht: (2024)
MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
von: Fang, Zhenying, et al.
Veröffentlicht: (2025)
von: Fang, Zhenying, et al.
Veröffentlicht: (2025)
EditCLIP: Representation Learning for Image Editing
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
GT-Mean Loss: A Simple Yet Effective Solution for Brightness Mismatch in Low-Light Image Enhancement
von: Liao, Jingxi, et al.
Veröffentlicht: (2025)
von: Liao, Jingxi, et al.
Veröffentlicht: (2025)
Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
von: Ye, Weixin, et al.
Veröffentlicht: (2026)
von: Ye, Weixin, et al.
Veröffentlicht: (2026)
Doubly Abductive Counterfactual Inference for Text-based Image Editing
von: Song, Xue, et al.
Veröffentlicht: (2024)
von: Song, Xue, et al.
Veröffentlicht: (2024)
Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
von: Yu, Yangche, et al.
Veröffentlicht: (2025)
von: Yu, Yangche, et al.
Veröffentlicht: (2025)
Visual Jigsaw Post-Training Improves MLLMs
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
Solving Convex Partition Visual Jigsaw Puzzles
von: Ohayon, Yaniv, et al.
Veröffentlicht: (2025)
von: Ohayon, Yaniv, et al.
Veröffentlicht: (2025)
Attention Head Purification: A New Perspective to Harness CLIP for Domain Generalization
von: Wang, Yingfan, et al.
Veröffentlicht: (2024)
von: Wang, Yingfan, et al.
Veröffentlicht: (2024)
Boundary Discretization and Reliable Classification Network for Temporal Action Detection
von: Fang, Zhenying, et al.
Veröffentlicht: (2023)
von: Fang, Zhenying, et al.
Veröffentlicht: (2023)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
von: Wang, Juan, et al.
Veröffentlicht: (2026)
von: Wang, Juan, et al.
Veröffentlicht: (2026)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
von: Agarwal, Aishwarya, et al.
Veröffentlicht: (2025)
von: Agarwal, Aishwarya, et al.
Veröffentlicht: (2025)
From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in Videos
von: Chen, Yin, et al.
Veröffentlicht: (2023)
von: Chen, Yin, et al.
Veröffentlicht: (2023)
Jigsaw++: Imagining Complete Shape Priors for Object Reassembly
von: Lu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Lu, Jiaxin, et al.
Veröffentlicht: (2024)
Solving Masked Jigsaw Puzzles with Diffusion Vision Transformers
von: Liu, Jinyang, et al.
Veröffentlicht: (2024)
von: Liu, Jinyang, et al.
Veröffentlicht: (2024)
Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
von: Wang, Zifu, et al.
Veröffentlicht: (2025)
von: Wang, Zifu, et al.
Veröffentlicht: (2025)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIP
von: Song, Chenyue, et al.
Veröffentlicht: (2025)
von: Song, Chenyue, et al.
Veröffentlicht: (2025)
CLIP-IT: CLIP-based Pairing for Histology Images Classification
von: Karimian, Banafsheh, et al.
Veröffentlicht: (2025)
von: Karimian, Banafsheh, et al.
Veröffentlicht: (2025)
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
von: Shi, Liangtao, et al.
Veröffentlicht: (2025)
von: Shi, Liangtao, et al.
Veröffentlicht: (2025)
Unwarping Screen Content Images via Structure-texture Enhancement Network and Transformation Self-estimation
von: Xiao, Zhenzhen, et al.
Veröffentlicht: (2025)
von: Xiao, Zhenzhen, et al.
Veröffentlicht: (2025)
Multi-Perspective Subimage CLIP with Keyword Guidance for Remote Sensing Image-Text Retrieval
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIP
von: Zeng, Yirui, et al.
Veröffentlicht: (2025)
von: Zeng, Yirui, et al.
Veröffentlicht: (2025)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
von: Ali, Muhammad, et al.
Veröffentlicht: (2024)
von: Ali, Muhammad, et al.
Veröffentlicht: (2024)
CLIP in Medical Imaging: A Survey
von: Zhao, Zihao, et al.
Veröffentlicht: (2023)
von: Zhao, Zihao, et al.
Veröffentlicht: (2023)
CLIP-SR: Collaborative Linguistic and Image Processing for Super-Resolution
von: Hu, Bingwen, et al.
Veröffentlicht: (2024)
von: Hu, Bingwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023) -
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
von: Song, Zijie, et al.
Veröffentlicht: (2025) -
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022) -
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2025) -
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)