Saved in:
| Main Authors: | Wang, Bohan, Yue, Zhongqi, Zhang, Fengda, Chen, Shuo, Bi, Li'an, Zhang, Junzhe, Song, Xue, Chan, Kennard Yanting, Pan, Jiachun, Wu, Weijia, Zhou, Mingze, Lin, Wang, Pan, Kaihang, Zhang, Saining, Jia, Liyu, Hu, Wentao, Zhao, Wei, Zhang, Hanwang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.07538 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
by: Lin, Wang, et al.
Published: (2025)
by: Lin, Wang, et al.
Published: (2025)
MIND: Benchmarking Memory Consistency and Action Control in World Models
by: Ye, Yixuan, et al.
Published: (2026)
by: Ye, Yixuan, et al.
Published: (2026)
Few-shot Learner Parameterization by Diffusion Time-steps
by: Yue, Zhongqi, et al.
Published: (2024)
by: Yue, Zhongqi, et al.
Published: (2024)
Auto-Encoding Morph-Tokens for Multimodal LLM
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
by: Yu, Qifan, et al.
Published: (2024)
by: Yu, Qifan, et al.
Published: (2024)
IntegratedPIFu: Integrated Pixel Aligned Implicit Function for Single-view Human Reconstruction
by: Chan, Kennard Yanting, et al.
Published: (2022)
by: Chan, Kennard Yanting, et al.
Published: (2022)
Distributionally Generative Augmentation for Fair Facial Attribute Classification
by: Zhang, Fengda, et al.
Published: (2024)
by: Zhang, Fengda, et al.
Published: (2024)
From Values to Tokens: An LLM-Driven Framework for Context-aware Time Series Forecasting via Symbolic Discretization
by: Tao, Xiaoyu, et al.
Published: (2025)
by: Tao, Xiaoyu, et al.
Published: (2025)
ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models
by: Gao, Kaifeng, et al.
Published: (2024)
by: Gao, Kaifeng, et al.
Published: (2024)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
Fine Structure-Aware Sampling: A New Sampling Training Scheme for Pixel-Aligned Implicit Models in Single-View Human Reconstruction
by: Chan, Kennard Yanting, et al.
Published: (2024)
by: Chan, Kennard Yanting, et al.
Published: (2024)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
MTMD: Multi-Scale Temporal Memory Learning and Efficient Debiasing Framework for Stock Trend Forecasting
by: Wang, Mingjie, et al.
Published: (2022)
by: Wang, Mingjie, et al.
Published: (2022)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
Exploring Diffusion Time-steps for Unsupervised Representation Learning
by: Yue, Zhongqi, et al.
Published: (2024)
by: Yue, Zhongqi, et al.
Published: (2024)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
by: Zheng, Guangting, et al.
Published: (2025)
by: Zheng, Guangting, et al.
Published: (2025)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought
by: Zhao, Kesen, et al.
Published: (2026)
by: Zhao, Kesen, et al.
Published: (2026)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
by: Gao, Kaifeng, et al.
Published: (2024)
by: Gao, Kaifeng, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Enhancing CLIP Robustness via Cross-Modality Alignment
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Recent Advances in Discrete Speech Tokens: A Review
by: Guo, Yiwei, et al.
Published: (2025)
by: Guo, Yiwei, et al.
Published: (2025)
Discrete forecast reconciliation
by: Zhang, Bohan, et al.
Published: (2023)
by: Zhang, Bohan, et al.
Published: (2023)
Accelerating Inference of Discrete Autoregressive Normalizing Flows by Selective Jacobi Decoding
by: Zhang, Jiaru, et al.
Published: (2025)
by: Zhang, Jiaru, et al.
Published: (2025)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Hierarchical‐Split Multi‐Scale Convolution Network With Multi‐Task Learning for Human Activity Recognition in Wearable Devices
by: Pengwei Zhang, et al.
Published: (2026)
by: Pengwei Zhang, et al.
Published: (2026)
Adaptive Tokenization: On the Hop-Overpriority Problem in Tokenized Graph Learning Models
by: Wang, Zhibiao, et al.
Published: (2025)
by: Wang, Zhibiao, et al.
Published: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior
by: Wu, Zike, et al.
Published: (2024)
by: Wu, Zike, et al.
Published: (2024)
LottieGPT: Tokenizing Vector Animation for Autoregressive Generation
by: Chen, Junhao, et al.
Published: (2026)
by: Chen, Junhao, et al.
Published: (2026)
Efficient Matrix Implementation for Rotary Position Embedding
by: Minqi, Chen, et al.
Published: (2026)
by: Minqi, Chen, et al.
Published: (2026)
Efficient Optimization of Variational Autoregressive Networks with Natural Gradient
by: Liu, Jing, et al.
Published: (2024)
by: Liu, Jing, et al.
Published: (2024)
Entropy-Guided Token Dropout: Training Autoregressive Language Models with Limited Domain Data
by: Wang, Jiapeng, et al.
Published: (2025)
by: Wang, Jiapeng, et al.
Published: (2025)
Similar Items
-
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
by: Pan, Kaihang, et al.
Published: (2025) -
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
by: Chow, Wei, et al.
Published: (2025) -
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
by: Lin, Wang, et al.
Published: (2025) -
MIND: Benchmarking Memory Consistency and Action Control in World Models
by: Ye, Yixuan, et al.
Published: (2026) -
Few-shot Learner Parameterization by Diffusion Time-steps
by: Yue, Zhongqi, et al.
Published: (2024)