Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Dong, Wei, Fangyun, Wan, Ziyu, Chen, Dongdong, Zhang, Jiawei, Zhao, Jinjing, Zhang, Sirui, Yue, Yang, Liang, Zhiyang, Guo, Baining, Luo, Chong, Bao, Jianmin, Li, Ji, Shi, Lei, Yang, Qinhong, Wu, Xiuyu, Feng, Xuelu, Lu, Yan, Dong, Yanchen, Wang, Yitong, Chen, Yunuo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast Autoregressive Models for Continuous Latent Generation
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
von: Hang, Tiankai, et al.
Veröffentlicht: (2025)
Diffusion Models without Classifier-free Guidance
von: Tang, Zhicong, et al.
Veröffentlicht: (2025)
von: Tang, Zhicong, et al.
Veröffentlicht: (2025)
MageBench: Bridging Large Multimodal Models to Agents
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
VolumeDiffusion: Flexible Text-to-3D Generation with Efficient Volumetric Encoder
von: Tang, Zhicong, et al.
Veröffentlicht: (2023)
von: Tang, Zhicong, et al.
Veröffentlicht: (2023)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
Rethinking Generative Large Language Model Evaluation for Semantic Comprehension
von: Wei, Fangyun, et al.
Veröffentlicht: (2024)
von: Wei, Fangyun, et al.
Veröffentlicht: (2024)
SynChart: Synthesizing Charts from Language Models
von: Liu, Mengchen, et al.
Veröffentlicht: (2024)
von: Liu, Mengchen, et al.
Veröffentlicht: (2024)
LACON: Training Text-to-Image Model from Uncurated Data
von: Liang, Zhiyang, et al.
Veröffentlicht: (2026)
von: Liang, Zhiyang, et al.
Veröffentlicht: (2026)
Efficient Diffusion Training via Min-SNR Weighting Strategy
von: Hang, Tiankai, et al.
Veröffentlicht: (2023)
von: Hang, Tiankai, et al.
Veröffentlicht: (2023)
From Virtual Games to Real-World Play
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
Semantic Image Synthesis via Diffusion Models
von: Zhou, Wengang, et al.
Veröffentlicht: (2022)
von: Zhou, Wengang, et al.
Veröffentlicht: (2022)
CCEdit: Creative and Controllable Video Editing via Diffusion Models
von: Feng, Ruoyu, et al.
Veröffentlicht: (2023)
von: Feng, Ruoyu, et al.
Veröffentlicht: (2023)
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians
von: Zhang, Xiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiyu, et al.
Veröffentlicht: (2025)
Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
von: Yue, Yang, et al.
Veröffentlicht: (2026)
von: Yue, Yang, et al.
Veröffentlicht: (2026)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
SmartEraser: Remove Anything from Images using Masked-Region Guidance
von: Jiang, Longtao, et al.
Veröffentlicht: (2025)
von: Jiang, Longtao, et al.
Veröffentlicht: (2025)
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
von: Wang, Wenbo, et al.
Veröffentlicht: (2026)
von: Wang, Wenbo, et al.
Veröffentlicht: (2026)
Spatia: Video Generation with Updatable Spatial Memory
von: Zhao, Jinjing, et al.
Veröffentlicht: (2025)
von: Zhao, Jinjing, et al.
Veröffentlicht: (2025)
Full‐Dimensional Penetration Strategy with Degradable PEAI Enables 8.21% Efficiency in Bulk Heterojunction Sb2S3 Solar Cells
von: Yang Wang, et al.
Veröffentlicht: (2025)
von: Yang Wang, et al.
Veröffentlicht: (2025)
Laparoscopic Cervical Cerclage Offers Greatest Benefit for Patients With Multiple Conizations or Significantly Shortened Cervix: A Retrospective Study Based on Causal Inference Modeling
von: Xia Li, et al.
Veröffentlicht: (2026)
von: Xia Li, et al.
Veröffentlicht: (2026)
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Animate Any Character in Any World
von: Wang, Yitong, et al.
Veröffentlicht: (2025)
von: Wang, Yitong, et al.
Veröffentlicht: (2025)
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
von: Chen, Jierun, et al.
Veröffentlicht: (2024)
von: Chen, Jierun, et al.
Veröffentlicht: (2024)
Rethinking Quantum Noise in Quantum Machine Learning: When Noise Improves Learning
von: Zhu, Linghua, et al.
Veröffentlicht: (2026)
von: Zhu, Linghua, et al.
Veröffentlicht: (2026)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformer
von: Dong, Qingji, et al.
Veröffentlicht: (2026)
von: Dong, Qingji, et al.
Veröffentlicht: (2026)
Antibacterial Activity of MC‐170, a 2,2‐Disubstituted Indole‐3‐one Derivative, Against Staphylococcus aureus via Phosphatidylglycerol
von: Yumiao Zhao, et al.
Veröffentlicht: (2025)
von: Yumiao Zhao, et al.
Veröffentlicht: (2025)
Rethinking the Vulnerability of Concept Erasure and a New Method
von: Richardson, Alex D., et al.
Veröffentlicht: (2025)
von: Richardson, Alex D., et al.
Veröffentlicht: (2025)
Simplified Diffusion Schrödinger Bridge
von: Tang, Zhicong, et al.
Veröffentlicht: (2024)
von: Tang, Zhicong, et al.
Veröffentlicht: (2024)
CCA: Collaborative Competitive Agents for Image Editing
von: Hang, Tiankai, et al.
Veröffentlicht: (2024)
von: Hang, Tiankai, et al.
Veröffentlicht: (2024)
Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
von: Tang, Zhicong, et al.
Veröffentlicht: (2025)
von: Tang, Zhicong, et al.
Veröffentlicht: (2025)
Rethinking Model Ensemble in Transfer-based Adversarial Attacks
von: Chen, Huanran, et al.
Veröffentlicht: (2023)
von: Chen, Huanran, et al.
Veröffentlicht: (2023)
A High‐Efficiency Codesign Method for Bandgap Circuit by Submodule Optimization
von: Yunqi Yang, et al.
Veröffentlicht: (2025)
von: Yunqi Yang, et al.
Veröffentlicht: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
New record-breaking binary linear codes constructed from group codes
von: Yu, Cong, et al.
Veröffentlicht: (2024)
von: Yu, Cong, et al.
Veröffentlicht: (2024)
Inevitable Encounters: Backdoor Attacks Involving Lossy Compression
von: Li, Qian, et al.
Veröffentlicht: (2026)
von: Li, Qian, et al.
Veröffentlicht: (2026)
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Fast Autoregressive Models for Continuous Latent Generation
von: Hang, Tiankai, et al.
Veröffentlicht: (2025) -
Diffusion Models without Classifier-free Guidance
von: Tang, Zhicong, et al.
Veröffentlicht: (2025) -
MageBench: Bridging Large Multimodal Models to Agents
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024) -
VolumeDiffusion: Flexible Text-to-3D Generation with Efficient Volumetric Encoder
von: Tang, Zhicong, et al.
Veröffentlicht: (2023) -
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024)