Gespeichert in:
| Hauptverfasser: | Pan, Kaihang, Lin, Wang, Yue, Zhongqi, Ao, Tenglong, Jia, Liyu, Zhao, Wei, Li, Juncheng, Tang, Siliang, Zhang, Hanwang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2504.14666 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
von: Lin, Wang, et al.
Veröffentlicht: (2025)
von: Lin, Wang, et al.
Veröffentlicht: (2025)
Auto-Encoding Morph-Tokens for Multimodal LLM
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
von: Wang, Bohan, et al.
Veröffentlicht: (2025)
von: Wang, Bohan, et al.
Veröffentlicht: (2025)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
von: Qiu, Haiyi, et al.
Veröffentlicht: (2026)
von: Qiu, Haiyi, et al.
Veröffentlicht: (2026)
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
Few-shot Learner Parameterization by Diffusion Time-steps
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
Body of Her: A Preliminary Study on End-to-End Humanoid Agent
von: Ao, Tenglong
Veröffentlicht: (2024)
von: Ao, Tenglong
Veröffentlicht: (2024)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
von: Fan, Zhaoyu, et al.
Veröffentlicht: (2025)
von: Fan, Zhaoyu, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
von: Gao, Minghe, et al.
Veröffentlicht: (2025)
von: Gao, Minghe, et al.
Veröffentlicht: (2025)
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
On Path to Multimodal Generalist: General-Level and General-Bench
von: Fei, Hao, et al.
Veröffentlicht: (2025)
von: Fei, Hao, et al.
Veröffentlicht: (2025)
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
Exploring Diffusion Time-steps for Unsupervised Representation Learning
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought
von: Zhao, Kesen, et al.
Veröffentlicht: (2026)
von: Zhao, Kesen, et al.
Veröffentlicht: (2026)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
von: Qiu, Haiyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haiyi, et al.
Veröffentlicht: (2024)
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
von: Bu, Wendong, et al.
Veröffentlicht: (2025)
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
von: Liu, Yaofang, et al.
Veröffentlicht: (2025)
von: Liu, Yaofang, et al.
Veröffentlicht: (2025)
TASR: Timestep-Aware Diffusion Model for Image Super-Resolution
von: Lin, Qinwei, et al.
Veröffentlicht: (2024)
von: Lin, Qinwei, et al.
Veröffentlicht: (2024)
Timestep-Aware Correction for Quantized Diffusion Models
von: Yao, Yuzhe, et al.
Veröffentlicht: (2024)
von: Yao, Yuzhe, et al.
Veröffentlicht: (2024)
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
von: Wang, Yuqing, et al.
Veröffentlicht: (2026)
von: Wang, Yuqing, et al.
Veröffentlicht: (2026)
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
von: Liu, Feng, et al.
Veröffentlicht: (2024)
von: Liu, Feng, et al.
Veröffentlicht: (2024)
OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
von: Pan, Kaihang, et al.
Veröffentlicht: (2026)
von: Pan, Kaihang, et al.
Veröffentlicht: (2026)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
ERTACache: Error Rectification and Timesteps Adjustment for Efficient Diffusion
von: Peng, Xurui, et al.
Veröffentlicht: (2025)
von: Peng, Xurui, et al.
Veröffentlicht: (2025)
Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
von: Zheng, Zhixin, et al.
Veröffentlicht: (2025)
von: Zheng, Zhixin, et al.
Veröffentlicht: (2025)
TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping
von: Ding, Ning, et al.
Veröffentlicht: (2025)
von: Ding, Ning, et al.
Veröffentlicht: (2025)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
von: Lin, Wang, et al.
Veröffentlicht: (2025) -
Auto-Encoding Morph-Tokens for Multimodal LLM
von: Pan, Kaihang, et al.
Veröffentlicht: (2024) -
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
von: Wang, Bohan, et al.
Veröffentlicht: (2025) -
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
von: Yu, Qifan, et al.
Veröffentlicht: (2024) -
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)