Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Kai, Li, Xiang, Kuen, Jason, Chen, Hao, Xu, Xiaohao, Gu, Jiuxiang, Luo, Yinyi, Raj, Bhiksha, Lin, Zhe, Savvides, Marios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image Tokenizer Needs Post-Training
by: Qiu, Kai, et al.
Published: (2025)
by: Qiu, Kai, et al.
Published: (2025)
ImageFolder: Autoregressive Image Generation with Folded Tokens
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Self-Corrected Image Generation with Explainable Latent Rewards
by: Luo, Yinyi, et al.
Published: (2026)
by: Luo, Yinyi, et al.
Published: (2026)
Efficient Autoregressive Audio Modeling via Next-Scale Prediction
by: Qiu, Kai, et al.
Published: (2024)
by: Qiu, Kai, et al.
Published: (2024)
LatentUMM: Dual Latent Alignment for Unified Multimodal Models
by: Luo, Yinyi, et al.
Published: (2026)
by: Luo, Yinyi, et al.
Published: (2026)
ControlVAR: Exploring Controllable Visual Autoregressive Modeling
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
by: Luo, Yinyi, et al.
Published: (2025)
by: Luo, Yinyi, et al.
Published: (2025)
SNCE: Geometry-Aware Supervision for Scalable Discrete Image Generation
by: Li, Shufan, et al.
Published: (2026)
by: Li, Shufan, et al.
Published: (2026)
$\text{R}^2$-Bench: Benchmarking the Robustness of Referring Perception Models under Perturbations
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
An Embarrassingly Simple Baseline for Imbalanced Semi-Supervised Learning
by: Chen, Hao, et al.
Published: (2022)
by: Chen, Hao, et al.
Published: (2022)
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
Customizable Perturbation Synthesis for Robust SLAM Benchmarking
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
SOLAR: Scalable Optimization of Large-scale Architecture for Reasoning
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
by: Luo, Yinyi, et al.
Published: (2026)
by: Luo, Yinyi, et al.
Published: (2026)
Revisiting Acoustic Features for Robust ASR
by: Shah, Muhammad A., et al.
Published: (2024)
by: Shah, Muhammad A., et al.
Published: (2024)
LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models
by: Li, Shufan, et al.
Published: (2026)
by: Li, Shufan, et al.
Published: (2026)
DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation
by: Li, Zhuoling, et al.
Published: (2026)
by: Li, Zhuoling, et al.
Published: (2026)
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
by: Baali, Massa, et al.
Published: (2025)
by: Baali, Massa, et al.
Published: (2025)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
by: Chen, Fangyi, et al.
Published: (2024)
by: Chen, Fangyi, et al.
Published: (2024)
Hierarchical Knowledge Graph Construction from Images for Scalable E-Commerce
by: Yang, Zhantao, et al.
Published: (2024)
by: Yang, Zhantao, et al.
Published: (2024)
Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration
by: Li, Zhuoling, et al.
Published: (2026)
by: Li, Zhuoling, et al.
Published: (2026)
From Perfect to Noisy World Simulation: Customizable Embodied Multi-modal Perturbations for SLAM Robustness Benchmarking
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning
by: Yao, Changwei, et al.
Published: (2025)
by: Yao, Changwei, et al.
Published: (2025)
On the Robust Approximation of ASR Metrics
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
Human Voice is Unique
by: Singh, Rita, et al.
Published: (2025)
by: Singh, Rita, et al.
Published: (2025)
SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
by: Ye, Hanrong, et al.
Published: (2023)
by: Ye, Hanrong, et al.
Published: (2023)
Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video
by: Xu, Xiaohao, et al.
Published: (2025)
by: Xu, Xiaohao, et al.
Published: (2025)
On Catastrophic Inheritance of Large Foundation Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024)
by: Sharma, Roshan, et al.
Published: (2024)
Perturbation Ontology based Graph Attention Networks
by: Wang, Yichen, et al.
Published: (2024)
by: Wang, Yichen, et al.
Published: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
by: Krichli, Tomer, et al.
Published: (2025)
by: Krichli, Tomer, et al.
Published: (2025)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
by: Dixit, Satvik, et al.
Published: (2024)
by: Dixit, Satvik, et al.
Published: (2024)
AURA Score: A Metric For Holistic Audio Question Answering Evaluation
by: Dixit, Satvik, et al.
Published: (2025)
by: Dixit, Satvik, et al.
Published: (2025)
Domain Adaptation for Contrastive Audio-Language Models
by: Deshmukh, Soham, et al.
Published: (2024)
by: Deshmukh, Soham, et al.
Published: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
by: Cohen, Eyal, et al.
Published: (2025)
by: Cohen, Eyal, et al.
Published: (2025)
On Fairness of Unified Multimodal Large Language Model for Image Generation
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
SOHES: Self-supervised Open-world Hierarchical Entity Segmentation
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
Similar Items
-
Image Tokenizer Needs Post-Training
by: Qiu, Kai, et al.
Published: (2025) -
ImageFolder: Autoregressive Image Generation with Folded Tokens
by: Li, Xiang, et al.
Published: (2024) -
XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
by: Li, Xiang, et al.
Published: (2024) -
Self-Corrected Image Generation with Explainable Latent Rewards
by: Luo, Yinyi, et al.
Published: (2026) -
Efficient Autoregressive Audio Modeling via Next-Scale Prediction
by: Qiu, Kai, et al.
Published: (2024)