The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Weichen, Diao, Haiwen, Wang, Quan, Lin, Dahua, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
by: Diao, Haiwen, et al.
Published: (2025)
by: Diao, Haiwen, et al.
Published: (2025)
From Pixels to Words -- Towards Native One-Vision Models at Scale
by: Diao, Haiwen, et al.
Published: (2026)
by: Diao, Haiwen, et al.
Published: (2026)
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025)
by: Si, Chenyang, et al.
Published: (2025)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
by: Xiang, Wenzhao, et al.
Published: (2026)
by: Xiang, Wenzhao, et al.
Published: (2026)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
by: Qiu, Xuerui, et al.
Published: (2026)
by: Qiu, Xuerui, et al.
Published: (2026)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models
by: Fan, Weichen, et al.
Published: (2025)
by: Fan, Weichen, et al.
Published: (2025)
ConsistCompose: Unified Multimodal Layout Control for Image Composition
by: Shi, Xuanke, et al.
Published: (2025)
by: Shi, Xuanke, et al.
Published: (2025)
FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain
by: Wang, YuAn, et al.
Published: (2025)
by: Wang, YuAn, et al.
Published: (2025)
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
by: Gao, Sensen, et al.
Published: (2026)
by: Gao, Sensen, et al.
Published: (2026)
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
by: Zhao, Yaqi, et al.
Published: (2026)
by: Zhao, Yaqi, et al.
Published: (2026)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
by: Xie, Haozhe, et al.
Published: (2026)
by: Xie, Haozhe, et al.
Published: (2026)
Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart Parsing
by: Li, Jinsong, et al.
Published: (2026)
by: Li, Jinsong, et al.
Published: (2026)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
by: Li, Rongjie, et al.
Published: (2024)
by: Li, Rongjie, et al.
Published: (2024)
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
by: Karypidis, Efstathios, et al.
Published: (2026)
by: Karypidis, Efstathios, et al.
Published: (2026)
Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
by: Yun, Jooyeol, et al.
Published: (2025)
by: Yun, Jooyeol, et al.
Published: (2025)
Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
by: Qian, Rui, et al.
Published: (2023)
by: Qian, Rui, et al.
Published: (2023)
Unified Multi-Site Multi-Sequence Brain MRI Harmonization Enriched by Biomedical Semantic Style
by: Wu, Mengqi, et al.
Published: (2026)
by: Wu, Mengqi, et al.
Published: (2026)
Unified Human-Scene Interaction via Prompted Chain-of-Contacts
by: Xiao, Zeqi, et al.
Published: (2023)
by: Xiao, Zeqi, et al.
Published: (2023)
Exploiting Semantic and Pixel Representations for Ultra-Low Bitrate Image Compression
by: Wei, Hao, et al.
Published: (2026)
by: Wei, Hao, et al.
Published: (2026)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation
by: Ming, Yuhang, et al.
Published: (2025)
by: Ming, Yuhang, et al.
Published: (2025)
RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations
by: Ge, Yanhao, et al.
Published: (2026)
by: Ge, Yanhao, et al.
Published: (2026)
Decoupling Semantics and Fingerprints: A Universal Representation for AI-Generated Image Detection
by: Wang, Zhiyuan, et al.
Published: (2026)
by: Wang, Zhiyuan, et al.
Published: (2026)
SepRep-Net: Multi-source Free Domain Adaptation via Model Separation And Reparameterization
by: Jin, Ying, et al.
Published: (2024)
by: Jin, Ying, et al.
Published: (2024)
PixelBytes: Catching Unified Representation for Multimodal Generation
by: Furfaro, Fabien
Published: (2024)
by: Furfaro, Fabien
Published: (2024)
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
by: Zhang, Yuhan, et al.
Published: (2025)
by: Zhang, Yuhan, et al.
Published: (2025)
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
The Indra Representation Hypothesis for Multimodal Alignment
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
by: Diao, Haiwen, et al.
Published: (2026)
by: Diao, Haiwen, et al.
Published: (2026)
Pixel Sentence Representation Learning
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
PixelHacker: Image Inpainting with Structural and Semantic Consistency
by: Xu, Ziyang, et al.
Published: (2025)
by: Xu, Ziyang, et al.
Published: (2025)
From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
by: He, Feng, et al.
Published: (2025)
by: He, Feng, et al.
Published: (2025)
Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
by: Pan, Yueming, et al.
Published: (2025)
by: Pan, Yueming, et al.
Published: (2025)
Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
by: Su, Zhuo, et al.
Published: (2024)
by: Su, Zhuo, et al.
Published: (2024)
Similar Items
-
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
by: Diao, Haiwen, et al.
Published: (2025) -
From Pixels to Words -- Towards Native One-Vision Models at Scale
by: Diao, Haiwen, et al.
Published: (2026) -
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025) -
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024) -
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)