Guess the Unified Model: How Much Can We Recover from Generated Images?
Fuente:
arXiv
Saved in:
| Main Authors: | Cekinmez, Jasin, Mitsuhashi, Ryo, Wu, Addison J., Yin, Yida |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
by: Mitsuhashi, Ryo, et al.
Published: (2026)
by: Mitsuhashi, Ryo, et al.
Published: (2026)
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
by: Cekinmez, Jasin, et al.
Published: (2025)
by: Cekinmez, Jasin, et al.
Published: (2025)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
by: Chen, Haoyu, et al.
Published: (2026)
by: Chen, Haoyu, et al.
Published: (2026)
How Much 3D Do Video Foundation Models Encode?
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
by: Li, Ruihang, et al.
Published: (2026)
by: Li, Ruihang, et al.
Published: (2026)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
How Much You Ate? Food Portion Estimation on Spoons
by: Sharma, Aaryam, et al.
Published: (2024)
by: Sharma, Aaryam, et al.
Published: (2024)
Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion Models
by: Lopez, Eleonora, et al.
Published: (2024)
by: Lopez, Eleonora, et al.
Published: (2024)
MRecover: A Conditional Generative Model for Recovering Motion-Corrupted MR images Using AI Generated Contrast
by: Li, Jinghang, et al.
Published: (2026)
by: Li, Jinghang, et al.
Published: (2026)
PICABench: How Far Are We from Physically Realistic Image Editing?
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
From Sora What We Can See: A Survey of Text-to-Video Generation
by: Sun, Rui, et al.
Published: (2024)
by: Sun, Rui, et al.
Published: (2024)
How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video Generation
by: Davtyan, Aram, et al.
Published: (2023)
by: Davtyan, Aram, et al.
Published: (2023)
Can We Change the Stroke Size for Easier Diffusion?
by: Bai, Yunwei, et al.
Published: (2026)
by: Bai, Yunwei, et al.
Published: (2026)
Large Language Models Can Understanding Depth from Monocular Images
by: Xia, Zhongyi, et al.
Published: (2024)
by: Xia, Zhongyi, et al.
Published: (2024)
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict
by: Wu, Chaochen, et al.
Published: (2025)
by: Wu, Chaochen, et al.
Published: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
by: Zhuang, Shaobin, et al.
Published: (2026)
by: Zhuang, Shaobin, et al.
Published: (2026)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New Perspective
by: Yin, Zeyuan, et al.
Published: (2023)
by: Yin, Zeyuan, et al.
Published: (2023)
Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
by: Kumar, Ashwin, et al.
Published: (2026)
by: Kumar, Ashwin, et al.
Published: (2026)
Recovering Diagnostic Value: Super-Resolution-Aided Echocardiographic Classification in Resource-Constrained Imaging
by: Babu, Krishan Agyakari Raja, et al.
Published: (2025)
by: Babu, Krishan Agyakari Raja, et al.
Published: (2025)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Unified and Generalized Approach
by: Yan, Jiebin, et al.
Published: (2026)
by: Yan, Jiebin, et al.
Published: (2026)
Unified Thinker: A General Reasoning Modular Core for Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Bridging Synthetic and Real-World Domains: A Human-in-the-Loop Weakly-Supervised Framework for Industrial Toxic Emission Segmentation
by: Tao, Yida, et al.
Published: (2025)
by: Tao, Yida, et al.
Published: (2025)
Latent Action Control for Reasoning-Guided Unified Image Generation
by: Zhai, Fuxiang, et al.
Published: (2026)
by: Zhai, Fuxiang, et al.
Published: (2026)
Unified Text-Image Generation with Weakness-Targeted Post-Training
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
U-Mamba2: Scaling State Space Models for Dental Anatomy Segmentation in CBCT
by: Tan, Zhi Qin, et al.
Published: (2025)
by: Tan, Zhi Qin, et al.
Published: (2025)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
by: Bahram, Yara, et al.
Published: (2025)
by: Bahram, Yara, et al.
Published: (2025)
V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation
by: Zhang, Guiwei, et al.
Published: (2025)
by: Zhang, Guiwei, et al.
Published: (2025)
DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing
by: Wang, Dianyi, et al.
Published: (2026)
by: Wang, Dianyi, et al.
Published: (2026)
DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
by: Li, Xiao-Hui, et al.
Published: (2025)
by: Li, Xiao-Hui, et al.
Published: (2025)
Are General-Purpose Vision Models All We Need for 2D Medical Image Segmentation? A Cross-Dataset Empirical Study
by: Borst, Vanessa, et al.
Published: (2026)
by: Borst, Vanessa, et al.
Published: (2026)
Similar Items
-
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
by: Mitsuhashi, Ryo, et al.
Published: (2026) -
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
by: Cekinmez, Jasin, et al.
Published: (2025) -
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
by: Chen, Haoyu, et al.
Published: (2026) -
How Much 3D Do Video Foundation Models Encode?
by: Huang, Zixuan, et al.
Published: (2025) -
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
by: Li, Ruihang, et al.
Published: (2026)