Communication-Inspired Tokenization for Structured Image Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Davtyan, Aram, Sahin, Yusuf, Haghighi, Yasaman, Stapf, Sebastian, Acuaviva, Pablo, Alahi, Alexandre, Favaro, Paolo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video Generation
by: Davtyan, Aram, et al.
Published: (2023)
by: Davtyan, Aram, et al.
Published: (2023)
Composition of Memory Experts for Diffusion World Models
by: Stapf, Sebastian, et al.
Published: (2026)
by: Stapf, Sebastian, et al.
Published: (2026)
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
by: Haghighi, Yasaman, et al.
Published: (2026)
by: Haghighi, Yasaman, et al.
Published: (2026)
CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation
by: Davtyan, Aram, et al.
Published: (2024)
by: Davtyan, Aram, et al.
Published: (2024)
Faster Inference of Flow-Based Generative Models via Improved Data-Noise Coupling
by: Davtyan, Aram, et al.
Published: (2026)
by: Davtyan, Aram, et al.
Published: (2026)
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
by: Hassan, Mariam, et al.
Published: (2024)
by: Hassan, Mariam, et al.
Published: (2024)
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
LayerSync: Self-aligning Intermediate Layers
by: Haghighi, Yasaman, et al.
Published: (2025)
by: Haghighi, Yasaman, et al.
Published: (2025)
CODE: Confident Ordinary Differential Editing
by: van Delft, Bastien, et al.
Published: (2024)
by: van Delft, Bastien, et al.
Published: (2024)
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
by: Corbière, Charles, et al.
Published: (2025)
by: Corbière, Charles, et al.
Published: (2025)
A Multi-Loss Strategy for Vehicle Trajectory Prediction: Combining Off-Road, Diversity, and Directional Consistency Losses
by: Rahimi, Ahmad, et al.
Published: (2024)
by: Rahimi, Ahmad, et al.
Published: (2024)
KOALA: A Kalman Optimization Algorithm with Loss Adaptivity
by: Davtyan, Aram, et al.
Published: (2021)
by: Davtyan, Aram, et al.
Published: (2021)
Causal Disentanglement-Inspired Degradation Representation Learning for Full-Reference Image Quality Assessment
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Forecast-PEFT: Parameter-Efficient Fine-Tuning for Pre-trained Motion Forecasting Models
by: Wang, Jifeng, et al.
Published: (2024)
by: Wang, Jifeng, et al.
Published: (2024)
Helvipad: A Real-World Dataset for Omnidirectional Stereo Depth Estimation
by: Zayene, Mehdi, et al.
Published: (2024)
by: Zayene, Mehdi, et al.
Published: (2024)
EgoSim: An Egocentric Multi-view Simulator and Real Dataset for Body-worn Cameras during Motion and Activity
by: Hollidt, Dominik, et al.
Published: (2025)
by: Hollidt, Dominik, et al.
Published: (2025)
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
by: Li, Wuyang, et al.
Published: (2026)
by: Li, Wuyang, et al.
Published: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
Sim-to-Real Causal Transfer: A Metric Learning Approach to Causally-Aware Interaction Representations
by: Rahimi, Ahmad, et al.
Published: (2023)
by: Rahimi, Ahmad, et al.
Published: (2023)
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
by: Dong, Jiajun, et al.
Published: (2025)
by: Dong, Jiajun, et al.
Published: (2025)
Geometrically Constrained and Token-Based Probabilistic Spatial Transformers
by: Schmidt, Johann, et al.
Published: (2025)
by: Schmidt, Johann, et al.
Published: (2025)
On the Adversarial Robustness of Discrete Image Tokenizers
by: Bhagwatkar, Rishika, et al.
Published: (2026)
by: Bhagwatkar, Rishika, et al.
Published: (2026)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
by: Shao, Run, et al.
Published: (2024)
by: Shao, Run, et al.
Published: (2024)
Evolve to Inspire: Novelty Search for Diverse Image Generation
by: Inch, Alex, et al.
Published: (2025)
by: Inch, Alex, et al.
Published: (2025)
Cognitively-Inspired Tokens Overcome Egocentric Bias in Multimodal Models
by: Leonard, Bridget, et al.
Published: (2026)
by: Leonard, Bridget, et al.
Published: (2026)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
by: Zhu, Xuanyu, et al.
Published: (2026)
by: Zhu, Xuanyu, et al.
Published: (2026)
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
Scaling Image Tokenizers with Grouped Spherical Quantization
by: Wang, Jiangtao, et al.
Published: (2024)
by: Wang, Jiangtao, et al.
Published: (2024)
Scalable Image Tokenization with Index Backpropagation Quantization
by: Shi, Fengyuan, et al.
Published: (2024)
by: Shi, Fengyuan, et al.
Published: (2024)
Memory-Inspired Temporal Prompt Interaction for Text-Image Classification
by: Yu, Xinyao, et al.
Published: (2024)
by: Yu, Xinyao, et al.
Published: (2024)
Discriminative Class Tokens for Text-to-Image Diffusion Models
by: Schwartz, Idan, et al.
Published: (2023)
by: Schwartz, Idan, et al.
Published: (2023)
xT: Nested Tokenization for Larger Context in Large Images
by: Gupta, Ritwik, et al.
Published: (2024)
by: Gupta, Ritwik, et al.
Published: (2024)
TMCIR: Token Merge Benefits Composed Image Retrieval
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
MacTok: Robust Continuous Tokenization for Image Generation
by: Zeng, Hengyu, et al.
Published: (2026)
by: Zeng, Hengyu, et al.
Published: (2026)
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
by: Zhang, Xinliang, et al.
Published: (2025)
by: Zhang, Xinliang, et al.
Published: (2025)
Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single-Image Denoising
by: Li, Huaqiu, et al.
Published: (2025)
by: Li, Huaqiu, et al.
Published: (2025)
DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
by: Singh, Azad, et al.
Published: (2025)
by: Singh, Azad, et al.
Published: (2025)
Similar Items
-
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025) -
Rethinking Visual Intelligence: Insights from Video Pretraining
by: Acuaviva, Pablo, et al.
Published: (2025) -
Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video Generation
by: Davtyan, Aram, et al.
Published: (2023) -
Composition of Memory Experts for Diffusion World Models
by: Stapf, Sebastian, et al.
Published: (2026) -
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
by: Haghighi, Yasaman, et al.
Published: (2026)