What matters for Representation Alignment: Global Information or Spatial Structure?
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Jaskirat, Leng, Xingjian, Wu, Zongze, Zheng, Liang, Zhang, Richard, Shechtman, Eli, Xie, Saining |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improved Baselines with Representation Autoencoders
by: Singh, Jaskirat, et al.
Published: (2026)
by: Singh, Jaskirat, et al.
Published: (2026)
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
by: Gandikota, Rohit, et al.
Published: (2025)
by: Gandikota, Rohit, et al.
Published: (2025)
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026)
by: Duggal, Shivam, et al.
Published: (2026)
Lazy Diffusion Transformer for Interactive Image Editing
by: Nitzan, Yotam, et al.
Published: (2024)
by: Nitzan, Yotam, et al.
Published: (2024)
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
by: Leng, Xingjian, et al.
Published: (2025)
by: Leng, Xingjian, et al.
Published: (2025)
Structure-Guided Image Completion with Image-level and Object-level Semantic Discriminators
by: Zheng, Haitian, et al.
Published: (2022)
by: Zheng, Haitian, et al.
Published: (2022)
VLM-Guided Adaptive Negative Prompting for Creative Generation
by: Golan, Shelly, et al.
Published: (2025)
by: Golan, Shelly, et al.
Published: (2025)
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
by: Chen, Jiacheng, et al.
Published: (2025)
by: Chen, Jiacheng, et al.
Published: (2025)
Distilling Diffusion Models into Conditional GANs
by: Kang, Minguk, et al.
Published: (2024)
by: Kang, Minguk, et al.
Published: (2024)
TexSpot: 3D Texture Enhancement with Spatially-uniform Point Latent Representation
by: Lu, Ziteng, et al.
Published: (2026)
by: Lu, Ziteng, et al.
Published: (2026)
SplatFont3D: Structure-Aware Text-to-3D Artistic Font Generation with Part-Level Style Control
by: Gan, Ji, et al.
Published: (2025)
by: Gan, Ji, et al.
Published: (2025)
GSDiff: Synthesizing Vector Floorplans via Geometry-enhanced Structural Graph Generation
by: Hu, Sizhe, et al.
Published: (2024)
by: Hu, Sizhe, et al.
Published: (2024)
StructInbet: Integrating Explicit Structural Guidance into Inbetween Frame Generation
by: Pan, Zhenglin, et al.
Published: (2025)
by: Pan, Zhenglin, et al.
Published: (2025)
AdaContour: Adaptive Contour Descriptor with Hierarchical Representation
by: Ding, Tianyu, et al.
Published: (2024)
by: Ding, Tianyu, et al.
Published: (2024)
DISK: Differentiable Sparse Kernel Complex for Efficient Spatially-Variant Convolution
by: Wu, Zhizhen, et al.
Published: (2025)
by: Wu, Zhizhen, et al.
Published: (2025)
GeoFusion-CAD: Structure-Aware Diffusion with Geometric State Space for Parametric 3D Design
by: Zhou, Xiaolei, et al.
Published: (2026)
by: Zhou, Xiaolei, et al.
Published: (2026)
Workflow-Aware Structured Layer Decomposition for Illustration Production
by: Zhang, Tianyu, et al.
Published: (2026)
by: Zhang, Tianyu, et al.
Published: (2026)
3D Gaussian Inverse Rendering with Approximated Global Illumination
by: Wu, Zirui, et al.
Published: (2025)
by: Wu, Zirui, et al.
Published: (2025)
Negative Token Merging: Image-based Adversarial Feature Guidance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
LGTM: Local-to-Global Text-Driven Human Motion Diffusion Model
by: Sun, Haowen, et al.
Published: (2024)
by: Sun, Haowen, et al.
Published: (2024)
CompGS++: Compressed Gaussian Splatting for Static and Dynamic Scene Representation
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
CompGS: Efficient 3D Scene Representation via Compressed Gaussian Splatting
by: Liu, Xiangrui, et al.
Published: (2024)
by: Liu, Xiangrui, et al.
Published: (2024)
DaReNeRF: Direction-aware Representation for Dynamic Scenes
by: Lou, Ange, et al.
Published: (2024)
by: Lou, Ange, et al.
Published: (2024)
A Generalizable Light Transport 3D Embedding for Global Illumination
by: Xu, Bing, et al.
Published: (2025)
by: Xu, Bing, et al.
Published: (2025)
Towards Geometric-Photometric Joint Alignment for Facial Mesh Registration
by: Wang, Xizhi, et al.
Published: (2024)
by: Wang, Xizhi, et al.
Published: (2024)
VLMaterial: Procedural Material Generation with Large Vision-Language Models
by: Li, Beichen, et al.
Published: (2025)
by: Li, Beichen, et al.
Published: (2025)
LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation
by: Lu, Xin, et al.
Published: (2025)
by: Lu, Xin, et al.
Published: (2025)
Causality in Video Diffusers is Separable from Denoising
by: Bai, Xingjian, et al.
Published: (2026)
by: Bai, Xingjian, et al.
Published: (2026)
WalkTheDog: Cross-Morphology Motion Alignment via Phase Manifolds
by: Li, Peizhuo, et al.
Published: (2024)
by: Li, Peizhuo, et al.
Published: (2024)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
by: Zhao, Junchuan, et al.
Published: (2026)
by: Zhao, Junchuan, et al.
Published: (2026)
TurboEdit: Instant text-based image editing
by: Wu, Zongze, et al.
Published: (2024)
by: Wu, Zongze, et al.
Published: (2024)
DMesh: A Differentiable Mesh Representation
by: Son, Sanghyun, et al.
Published: (2024)
by: Son, Sanghyun, et al.
Published: (2024)
SD-GS: Structured Deformable 3D Gaussians for Efficient Dynamic Scene Reconstruction
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
Fine-Grained Spatially Varying Material Selection in Images
by: Guerrero-Viu, Julia, et al.
Published: (2025)
by: Guerrero-Viu, Julia, et al.
Published: (2025)
Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface
by: Luo, Yihao, et al.
Published: (2025)
by: Luo, Yihao, et al.
Published: (2025)
Limitations of (Procrustes) Alignment in Assessing Multi-Person Human Pose and Shape Estimation
by: Martin, Drazic, et al.
Published: (2024)
by: Martin, Drazic, et al.
Published: (2024)
Text-to-Vector Generation with Neural Path Representation
by: Zhang, Peiying, et al.
Published: (2024)
by: Zhang, Peiying, et al.
Published: (2024)
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
by: Liu, Haofeng, et al.
Published: (2026)
by: Liu, Haofeng, et al.
Published: (2026)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
Similar Items
-
Improved Baselines with Representation Autoencoders
by: Singh, Jaskirat, et al.
Published: (2026) -
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
by: Gandikota, Rohit, et al.
Published: (2025) -
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026) -
Lazy Diffusion Transformer for Interactive Image Editing
by: Nitzan, Yotam, et al.
Published: (2024) -
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
by: Leng, Xingjian, et al.
Published: (2025)