Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
Fuente:
arXiv
Guardado en:
| Autores principales: | Lin, Zinan, Liu, Enshu, Ning, Xuefei, Zhu, Junyi, Wang, Wenyu, Yekhanin, Sergey |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better
por: Liu, Enshu, et al.
Publicado: (2024)
por: Liu, Enshu, et al.
Publicado: (2024)
Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching
por: Liu, Enshu, et al.
Publicado: (2024)
por: Liu, Enshu, et al.
Publicado: (2024)
Differentially Private Synthetic Data via APIs 3: Using Simulators Instead of Foundation Model
por: Lin, Zinan, et al.
Publicado: (2025)
por: Lin, Zinan, et al.
Publicado: (2025)
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
por: Zhao, Tianchen, et al.
Publicado: (2024)
por: Zhao, Tianchen, et al.
Publicado: (2024)
Differentially Private Synthetic Data via Foundation Model APIs 1: Images
por: Lin, Zinan, et al.
Publicado: (2023)
por: Lin, Zinan, et al.
Publicado: (2023)
FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation
por: Huang, Kaiyi, et al.
Publicado: (2025)
por: Huang, Kaiyi, et al.
Publicado: (2025)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
por: Huang, Kaiyi, et al.
Publicado: (2024)
por: Huang, Kaiyi, et al.
Publicado: (2024)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
por: Zhao, Tianchen, et al.
Publicado: (2024)
por: Zhao, Tianchen, et al.
Publicado: (2024)
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
por: Zhao, Lin, et al.
Publicado: (2024)
por: Zhao, Lin, et al.
Publicado: (2024)
Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
por: Ma, Qianli, et al.
Publicado: (2024)
por: Ma, Qianli, et al.
Publicado: (2024)
Latent Action Control for Reasoning-Guided Unified Image Generation
por: Zhai, Fuxiang, et al.
Publicado: (2026)
por: Zhai, Fuxiang, et al.
Publicado: (2026)
DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
por: Liu, Dongxu, et al.
Publicado: (2025)
por: Liu, Dongxu, et al.
Publicado: (2025)
Multi-scale Unified Network for Image Classification
por: Liu, Wenzhuo, et al.
Publicado: (2024)
por: Liu, Wenzhuo, et al.
Publicado: (2024)
CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
por: Huang, Kaiyi, et al.
Publicado: (2026)
por: Huang, Kaiyi, et al.
Publicado: (2026)
Robust Latent Representation Tuning for Image-text Classification
por: Sun, Hao, et al.
Publicado: (2024)
por: Sun, Hao, et al.
Publicado: (2024)
UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation
por: Wu, Guanjun, et al.
Publicado: (2025)
por: Wu, Guanjun, et al.
Publicado: (2025)
Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective
por: Zhu, Yongxin, et al.
Publicado: (2024)
por: Zhu, Yongxin, et al.
Publicado: (2024)
Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
por: Wu, Junyi, et al.
Publicado: (2026)
por: Wu, Junyi, et al.
Publicado: (2026)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
por: Huang, Hai, et al.
Publicado: (2024)
por: Huang, Hai, et al.
Publicado: (2024)
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
por: Zhu, Lianghui, et al.
Publicado: (2025)
por: Zhu, Lianghui, et al.
Publicado: (2025)
AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding
por: Jabin, Md Asaduzzaman, et al.
Publicado: (2025)
por: Jabin, Md Asaduzzaman, et al.
Publicado: (2025)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
por: Lin, Yijing, et al.
Publicado: (2025)
por: Lin, Yijing, et al.
Publicado: (2025)
PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation
por: Fan, Zehua, et al.
Publicado: (2026)
por: Fan, Zehua, et al.
Publicado: (2026)
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
por: Zheng, Haoyu, et al.
Publicado: (2026)
por: Zheng, Haoyu, et al.
Publicado: (2026)
Towards Latent Masked Image Modeling for Self-Supervised Visual Representation Learning
por: Wei, Yibing, et al.
Publicado: (2024)
por: Wei, Yibing, et al.
Publicado: (2024)
LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
por: Debnath, Soumyaratna, et al.
Publicado: (2026)
por: Debnath, Soumyaratna, et al.
Publicado: (2026)
Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
por: Caetano, Francisco, et al.
Publicado: (2025)
por: Caetano, Francisco, et al.
Publicado: (2025)
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
por: Xing, Jiazheng, et al.
Publicado: (2026)
por: Xing, Jiazheng, et al.
Publicado: (2026)
PixelBytes: Catching Unified Representation for Multimodal Generation
por: Furfaro, Fabien
Publicado: (2024)
por: Furfaro, Fabien
Publicado: (2024)
Synthesize Privacy-Preserving High-Resolution Images via Private Textual Intermediaries
por: Wang, Haoxiang, et al.
Publicado: (2025)
por: Wang, Haoxiang, et al.
Publicado: (2025)
A World Model of Radiologist Reading for Medical Image Representation Learning
por: Li, Yiwei, et al.
Publicado: (2026)
por: Li, Yiwei, et al.
Publicado: (2026)
Multi-modal Relation Distillation for Unified 3D Representation Learning
por: Wang, Huiqun, et al.
Publicado: (2024)
por: Wang, Huiqun, et al.
Publicado: (2024)
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
por: Hao, Shaozhe, et al.
Publicado: (2024)
por: Hao, Shaozhe, et al.
Publicado: (2024)
Improving the Training of Rectified Flows
por: Lee, Sangyun, et al.
Publicado: (2024)
por: Lee, Sangyun, et al.
Publicado: (2024)
AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
por: Wang, Zile, et al.
Publicado: (2025)
por: Wang, Zile, et al.
Publicado: (2025)
Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space
por: Zhu, Jian, et al.
Publicado: (2025)
por: Zhu, Jian, et al.
Publicado: (2025)
Unified Continuous Generative Models
por: Sun, Peng, et al.
Publicado: (2025)
por: Sun, Peng, et al.
Publicado: (2025)
Adaptive Disentangled Representation Learning for Incomplete Multi-View Multi-Label Classification
por: Li, Quanjiang, et al.
Publicado: (2026)
por: Li, Quanjiang, et al.
Publicado: (2026)
TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
por: Li, Zongming, et al.
Publicado: (2025)
por: Li, Zongming, et al.
Publicado: (2025)
Implicit Neural Representations for Robust Joint Sparse-View CT Reconstruction
por: Shi, Jiayang, et al.
Publicado: (2024)
por: Shi, Jiayang, et al.
Publicado: (2024)
Ejemplares similares
-
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better
por: Liu, Enshu, et al.
Publicado: (2024) -
Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching
por: Liu, Enshu, et al.
Publicado: (2024) -
Differentially Private Synthetic Data via APIs 3: Using Simulators Instead of Foundation Model
por: Lin, Zinan, et al.
Publicado: (2025) -
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
por: Zhao, Tianchen, et al.
Publicado: (2024) -
Differentially Private Synthetic Data via Foundation Model APIs 1: Images
por: Lin, Zinan, et al.
Publicado: (2023)