Improved Baselines with Representation Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Jaskirat, Zheng, Boyang, Wu, Zongze, Zhang, Richard, Shechtman, Eli, Xie, Saining |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What matters for Representation Alignment: Global Information or Spatial Structure?
von: Singh, Jaskirat, et al.
Veröffentlicht: (2025)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2025)
End-to-End Training for Unified Tokenization and Latent Denoising
von: Duggal, Shivam, et al.
Veröffentlicht: (2026)
von: Duggal, Shivam, et al.
Veröffentlicht: (2026)
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
von: Gandikota, Rohit, et al.
Veröffentlicht: (2025)
von: Gandikota, Rohit, et al.
Veröffentlicht: (2025)
Lazy Diffusion Transformer for Interactive Image Editing
von: Nitzan, Yotam, et al.
Veröffentlicht: (2024)
von: Nitzan, Yotam, et al.
Veröffentlicht: (2024)
Negative Token Merging: Image-based Adversarial Feature Guidance
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
Improving Video Generation with Human Feedback
von: Liu, Jie, et al.
Veröffentlicht: (2025)
von: Liu, Jie, et al.
Veröffentlicht: (2025)
Diffusion Transformers with Representation Autoencoders
von: Zheng, Boyang, et al.
Veröffentlicht: (2025)
von: Zheng, Boyang, et al.
Veröffentlicht: (2025)
Causality in Video Diffusers is Separable from Denoising
von: Bai, Xingjian, et al.
Veröffentlicht: (2026)
von: Bai, Xingjian, et al.
Veröffentlicht: (2026)
Multi-Spectral Gaussian Splatting with Neural Color Representation
von: Meyer, Lukas, et al.
Veröffentlicht: (2025)
von: Meyer, Lukas, et al.
Veröffentlicht: (2025)
Distilling Diffusion Models into Conditional GANs
von: Kang, Minguk, et al.
Veröffentlicht: (2024)
von: Kang, Minguk, et al.
Veröffentlicht: (2024)
Uncertainty-Informed Volume Visualization using Implicit Neural Representation
von: Saklani, Shanu, et al.
Veröffentlicht: (2024)
von: Saklani, Shanu, et al.
Veröffentlicht: (2024)
Few-Shot Unsupervised Implicit Neural Shape Representation Learning with Spatial Adversaries
von: Ouasfi, Amine, et al.
Veröffentlicht: (2024)
von: Ouasfi, Amine, et al.
Veröffentlicht: (2024)
STEP-Parts: Geometric Partitioning of Boundary Representations for Large-Scale CAD Processing
von: Fan, Shen, et al.
Veröffentlicht: (2026)
von: Fan, Shen, et al.
Veröffentlicht: (2026)
BulletGen: Improving 4D Reconstruction with Bullet-Time Generation
von: Rozumny, Denis, et al.
Veröffentlicht: (2025)
von: Rozumny, Denis, et al.
Veröffentlicht: (2025)
Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
Edge-preserving noise for diffusion models
von: Vandersanden, Jente, et al.
Veröffentlicht: (2024)
von: Vandersanden, Jente, et al.
Veröffentlicht: (2024)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
Diffusion Self-Distillation for Zero-Shot Customized Image Generation
von: Cai, Shengqu, et al.
Veröffentlicht: (2024)
von: Cai, Shengqu, et al.
Veröffentlicht: (2024)
PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
von: Xie, Tianyi, et al.
Veröffentlicht: (2023)
von: Xie, Tianyi, et al.
Veröffentlicht: (2023)
ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning
von: Zhang, David Junhao, et al.
Veröffentlicht: (2024)
von: Zhang, David Junhao, et al.
Veröffentlicht: (2024)
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
von: Lin, Chieh Hubert, et al.
Veröffentlicht: (2025)
von: Lin, Chieh Hubert, et al.
Veröffentlicht: (2025)
RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation
von: Tan, Xianfeng, et al.
Veröffentlicht: (2024)
von: Tan, Xianfeng, et al.
Veröffentlicht: (2024)
ArtiFixer: Enhancing and Extending 3D Reconstruction with Auto-Regressive Diffusion Models
von: de Lutio, Riccardo, et al.
Veröffentlicht: (2026)
von: de Lutio, Riccardo, et al.
Veröffentlicht: (2026)
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
von: Seitzer, Maximilian, et al.
Veröffentlicht: (2023)
von: Seitzer, Maximilian, et al.
Veröffentlicht: (2023)
An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion
von: Ali, Sharib, et al.
Veröffentlicht: (2024)
von: Ali, Sharib, et al.
Veröffentlicht: (2024)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
PAPR in Motion: Seamless Point-level 3D Scene Interpolation
von: Peng, Shichong, et al.
Veröffentlicht: (2024)
von: Peng, Shichong, et al.
Veröffentlicht: (2024)
Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation
von: Cao, Tianshi, et al.
Veröffentlicht: (2026)
von: Cao, Tianshi, et al.
Veröffentlicht: (2026)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
von: Hong, Susung, et al.
Veröffentlicht: (2025)
von: Hong, Susung, et al.
Veröffentlicht: (2025)
Learning to Edit Visual Programs with Self-Supervision
von: Jones, R. Kenny, et al.
Veröffentlicht: (2024)
von: Jones, R. Kenny, et al.
Veröffentlicht: (2024)
Gaussian Splashing: Unified Particles for Versatile Motion Synthesis and Rendering
von: Feng, Yutao, et al.
Veröffentlicht: (2024)
von: Feng, Yutao, et al.
Veröffentlicht: (2024)
NeuSDFusion: A Spatial-Aware Generative Model for 3D Shape Completion, Reconstruction, and Generation
von: Cui, Ruikai, et al.
Veröffentlicht: (2024)
von: Cui, Ruikai, et al.
Veröffentlicht: (2024)
SimVS: Simulating World Inconsistencies for Robust View Synthesis
von: Trevithick, Alex, et al.
Veröffentlicht: (2024)
von: Trevithick, Alex, et al.
Veröffentlicht: (2024)
MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting
von: Chang, Hanzhi, et al.
Veröffentlicht: (2025)
von: Chang, Hanzhi, et al.
Veröffentlicht: (2025)
Controlling Text-to-Image Diffusion by Orthogonal Finetuning
von: Qiu, Zeju, et al.
Veröffentlicht: (2023)
von: Qiu, Zeju, et al.
Veröffentlicht: (2023)
BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation
von: Hou, Xiaolu, et al.
Veröffentlicht: (2025)
von: Hou, Xiaolu, et al.
Veröffentlicht: (2025)
LRM: Large Reconstruction Model for Single Image to 3D
von: Hong, Yicong, et al.
Veröffentlicht: (2023)
von: Hong, Yicong, et al.
Veröffentlicht: (2023)
Improving Physics-Augmented Continuum Neural Radiance Field-Based Geometry-Agnostic System Identification with Lagrangian Particle Optimization
von: Kaneko, Takuhiro
Veröffentlicht: (2024)
von: Kaneko, Takuhiro
Veröffentlicht: (2024)
Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
von: Zheng, Shuhong, et al.
Veröffentlicht: (2025)
von: Zheng, Shuhong, et al.
Veröffentlicht: (2025)
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
von: Sani, Samin Mahdizadeh, et al.
Veröffentlicht: (2026)
von: Sani, Samin Mahdizadeh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
What matters for Representation Alignment: Global Information or Spatial Structure?
von: Singh, Jaskirat, et al.
Veröffentlicht: (2025) -
End-to-End Training for Unified Tokenization and Latent Denoising
von: Duggal, Shivam, et al.
Veröffentlicht: (2026) -
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
von: Gandikota, Rohit, et al.
Veröffentlicht: (2025) -
Lazy Diffusion Transformer for Interactive Image Editing
von: Nitzan, Yotam, et al.
Veröffentlicht: (2024) -
Negative Token Merging: Image-based Adversarial Feature Guidance
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)