Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Tianzhe, Tong, Shengbang, Ding, Tianjiao, Dai, Xili, Haeffele, Benjamin David, Vidal, René, Ma, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023)
by: Yu, Yaodong, et al.
Published: (2023)
Ctrl123: Consistent Novel View Synthesis via Closed-Loop Transcription
by: Zhao, Hongxiang, et al.
Published: (2024)
by: Zhao, Hongxiang, et al.
Published: (2024)
Hierarchical Concept Embedding & Pursuit for Interpretable Image Classification
by: Nguyen, Nghia, et al.
Published: (2026)
by: Nguyen, Nghia, et al.
Published: (2026)
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
by: Huang, Yaxuan, et al.
Published: (2025)
by: Huang, Yaxuan, et al.
Published: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
Voyaging into Perpetual Dynamic Scenes from a Single View
by: Tian, Fengrui, et al.
Published: (2025)
by: Tian, Fengrui, et al.
Published: (2025)
Temporal Rate Reduction Clustering for Human Motion Segmentation
by: Meng, Xianghan, et al.
Published: (2025)
by: Meng, Xianghan, et al.
Published: (2025)
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)
by: Zheng, Boyang, et al.
Published: (2025)
Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection
by: Qi, Dacheng, et al.
Published: (2026)
by: Qi, Dacheng, et al.
Published: (2026)
Asymmetric Idiosyncrasies in Multimodal Models
by: Tao, Muzi, et al.
Published: (2026)
by: Tao, Muzi, et al.
Published: (2026)
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025)
by: Luo, Jinqi, et al.
Published: (2025)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Anti-Collapse Loss for Deep Metric Learning Based on Coding Rate Metric
by: Jiang, Xiruo, et al.
Published: (2024)
by: Jiang, Xiruo, et al.
Published: (2024)
Graph Cut-guided Maximal Coding Rate Reduction for Learning Image Embedding and Clustering
by: He, W., et al.
Published: (2024)
by: He, W., et al.
Published: (2024)
BAFNet: Bilateral Attention Fusion Network for Lightweight Semantic Segmentation of Urban Remote Sensing Images
by: Wang, Wentao, et al.
Published: (2024)
by: Wang, Wentao, et al.
Published: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
by: Wu, Ziyang, et al.
Published: (2024)
by: Wu, Ziyang, et al.
Published: (2024)
AmorLIP: Efficient Language-Image Pretraining via Amortization
by: Sun, Haotian, et al.
Published: (2025)
by: Sun, Haotian, et al.
Published: (2025)
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
by: Fang, Irving, et al.
Published: (2025)
by: Fang, Irving, et al.
Published: (2025)
USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
by: Chu, Xiangxiang, et al.
Published: (2025)
by: Chu, Xiangxiang, et al.
Published: (2025)
Adaptive Stain Normalization for Cross-Domain Medical Histology
by: Xu, Tianyue, et al.
Published: (2025)
by: Xu, Tianyue, et al.
Published: (2025)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Action Detection via an Image Diffusion Process
by: Foo, Lin Geng, et al.
Published: (2024)
by: Foo, Lin Geng, et al.
Published: (2024)
Clustering Propagation for Universal Medical Image Segmentation
by: Ding, Yuhang, et al.
Published: (2024)
by: Ding, Yuhang, et al.
Published: (2024)
ConsistEdit: Highly Consistent and Precise Training-free Visual Editing
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
Taming Transformer Without Using Learning Rate Warmup
by: Qi, Xianbiao, et al.
Published: (2025)
by: Qi, Xianbiao, et al.
Published: (2025)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
BrainAnytime: Anatomy-Aware Cross-Modal Pretraining for Brain Image Analysis with Arbitrary Modality Availability
by: Yang, Guangqian, et al.
Published: (2026)
by: Yang, Guangqian, et al.
Published: (2026)
[CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation
by: Wang, Akang, et al.
Published: (2026)
by: Wang, Akang, et al.
Published: (2026)
BIFRÖST: 3D-Aware Image compositing with Language Instructions
by: Li, Lingxiao, et al.
Published: (2024)
by: Li, Lingxiao, et al.
Published: (2024)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
EPIC: Explanation of Pretrained Image Classification Networks via Prototype
by: Borycki, Piotr, et al.
Published: (2025)
by: Borycki, Piotr, et al.
Published: (2025)
STF: Spatial Temporal Fusion for Trajectory Prediction
by: Han, Pengqian, et al.
Published: (2023)
by: Han, Pengqian, et al.
Published: (2023)
LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
MedTri: A Platform for Structured Medical Report Normalization to Enhance Vision-Language Pretraining
by: Chu, Yuetan, et al.
Published: (2026)
by: Chu, Yuetan, et al.
Published: (2026)
Task Specific Pretraining with Noisy Labels for Remote Sensing Image Segmentation
by: Liu, Chenying, et al.
Published: (2024)
by: Liu, Chenying, et al.
Published: (2024)
Similar Items
-
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023) -
Ctrl123: Consistent Novel View Synthesis via Closed-Loop Transcription
by: Zhao, Hongxiang, et al.
Published: (2024) -
Hierarchical Concept Embedding & Pursuit for Interpretable Image Classification
by: Nguyen, Nghia, et al.
Published: (2026) -
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
by: Huang, Yaxuan, et al.
Published: (2025) -
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)