iFSQ: Improving FSQ for Image Generation with 1 Line of Code
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Bin, Li, Zongjian, Niu, Yuwei, Gong, Kaixiong, Ge, Yunyang, Lin, Yunlong, Zheng, Mingzhe, Zhang, JianWei, Yang, Miles, Zhong, Zhao, Bo, Liefeng, Yuan, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
by: Meng, Zhaoyang, et al.
Published: (2026)
by: Meng, Zhaoyang, et al.
Published: (2026)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025)
by: Lin, Bin, et al.
Published: (2025)
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
by: Zheng, Mingzhe, et al.
Published: (2026)
by: Zheng, Mingzhe, et al.
Published: (2026)
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
by: Niu, Yuwei, et al.
Published: (2025)
by: Niu, Yuwei, et al.
Published: (2025)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)
by: Chen, Liuhan, et al.
Published: (2024)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning
by: Ge, Yunyang, et al.
Published: (2026)
by: Ge, Yunyang, et al.
Published: (2026)
FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
by: Ge, Yunyang, et al.
Published: (2025)
by: Ge, Yunyang, et al.
Published: (2025)
Improving Search Agent with One Line of Code
by: Li, Jian, et al.
Published: (2026)
by: Li, Jian, et al.
Published: (2026)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
by: Li, Zongjian, et al.
Published: (2025)
by: Li, Zongjian, et al.
Published: (2025)
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
by: Zhao, Yaqi, et al.
Published: (2026)
by: Zhao, Yaqi, et al.
Published: (2026)
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
by: Liu, Xiangyue, et al.
Published: (2026)
by: Liu, Xiangyue, et al.
Published: (2026)
ImgEdit: A Unified Image Editing Dataset and Benchmark
by: Ye, Yang, et al.
Published: (2025)
by: Ye, Yang, et al.
Published: (2025)
MOCHA: Discovering Multi-Order Dynamic Causality in Temporal Point Processes
by: Cao, Yunyang, et al.
Published: (2025)
by: Cao, Yunyang, et al.
Published: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
DanceMeld: Unraveling Dance Phrases with Hierarchical Latent Codes for Music-to-Dance Synthesis
by: Gao, Xin, et al.
Published: (2023)
by: Gao, Xin, et al.
Published: (2023)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
by: Zhao, Chengshu, et al.
Published: (2025)
by: Zhao, Chengshu, et al.
Published: (2025)
Improved Lower Bounds for Approximating Parameterized Nearest Codeword and Related Problems under ETH
by: Li, Shuangle, et al.
Published: (2024)
by: Li, Shuangle, et al.
Published: (2024)
Interpretable Hybrid-Rule Temporal Point Processes
by: Cao, Yunyang, et al.
Published: (2025)
by: Cao, Yunyang, et al.
Published: (2025)
Rota-Baxter groups with weight zero and integration on topological groups
by: Gao, Xing, et al.
Published: (2024)
by: Gao, Xing, et al.
Published: (2024)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Cautious Optimizers: Improving Training with One Line of Code
by: Liang, Kaizhao, et al.
Published: (2024)
by: Liang, Kaizhao, et al.
Published: (2024)
An Improved Bouc–Wen Model for High‐Damping Rubber Bearings Incorporating Large‐Strain Behavior: Development, Validation, and Comparative Analysis
by: Peng Chen, et al.
Published: (2025)
by: Peng Chen, et al.
Published: (2025)
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker
by: Wu, Zongjian, et al.
Published: (2026)
by: Wu, Zongjian, et al.
Published: (2026)
Helios: Real Real-Time Long Video Generation Model
by: Yuan, Shenghai, et al.
Published: (2026)
by: Yuan, Shenghai, et al.
Published: (2026)
Towards Fine-grained Interactive Segmentation in Images and Videos
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling
by: Men, Yifang, et al.
Published: (2024)
by: Men, Yifang, et al.
Published: (2024)
MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection
by: Yang, Qiushi, et al.
Published: (2025)
by: Yang, Qiushi, et al.
Published: (2025)
DiffuEraser: A Diffusion Model for Video Inpainting
by: Li, Xiaowen, et al.
Published: (2025)
by: Li, Xiaowen, et al.
Published: (2025)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
BIFRÖST: 3D-Aware Image compositing with Language Instructions
by: Li, Lingxiao, et al.
Published: (2024)
by: Li, Lingxiao, et al.
Published: (2024)
The Extension Arm Design Method Based on a Two‐Bar Tension Stretchable Mechanism
by: Song Gao, et al.
Published: (2025)
by: Song Gao, et al.
Published: (2025)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
AnyText2: Visual Text Generation and Editing With Customizable Attributes
by: Tuo, Yuxiang, et al.
Published: (2024)
by: Tuo, Yuxiang, et al.
Published: (2024)
UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization
by: He, Junjie, et al.
Published: (2024)
by: He, Junjie, et al.
Published: (2024)
Line-level Semantic Structure Learning for Code Vulnerability Detection
by: Wang, Ziliang, et al.
Published: (2024)
by: Wang, Ziliang, et al.
Published: (2024)
StableShard: Stable and Scalable Blockchain Sharding with High Concurrency via Collaborative Committees
by: Li, Mingzhe, et al.
Published: (2024)
by: Li, Mingzhe, et al.
Published: (2024)
SpiralShard: Highly Concurrent and Secure Blockchain Sharding via Linked Cross-shard Endorsement
by: Lin, You, et al.
Published: (2024)
by: Lin, You, et al.
Published: (2024)
Similar Items
-
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
by: Meng, Zhaoyang, et al.
Published: (2026) -
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025) -
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
by: Zheng, Mingzhe, et al.
Published: (2026) -
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
by: Niu, Yuwei, et al.
Published: (2025) -
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)