Mirai: Autoregressive Visual Generation Needs Foresight
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Yonghao, Huang, Lang, Wang, Zerun, Li, Runyi, Yamasaki, Toshihiko |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCOMatch: Alleviating Overtrusting in Open-set Semi-supervised Learning
by: Wang, Zerun, et al.
Published: (2024)
by: Wang, Zerun, et al.
Published: (2024)
Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data
by: Wang, Zerun, et al.
Published: (2024)
by: Wang, Zerun, et al.
Published: (2024)
From Obstacles to Resources: Semi-supervised Learning Faces Synthetic Data Contamination
by: Wang, Zerun, et al.
Published: (2024)
by: Wang, Zerun, et al.
Published: (2024)
Online Open-set Semi-supervised Object Detection with Dual Competing Head
by: Wang, Zerun, et al.
Published: (2023)
by: Wang, Zerun, et al.
Published: (2023)
Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up
by: Huang, Lang, et al.
Published: (2025)
by: Huang, Lang, et al.
Published: (2025)
Unified Vector Floorplan Generation via Markup Representation
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
by: Kuroki, Michihiro, et al.
Published: (2025)
by: Kuroki, Michihiro, et al.
Published: (2025)
BSED: Baseline Shapley-Based Explainable Detector
by: Kuroki, Michihiro, et al.
Published: (2023)
by: Kuroki, Michihiro, et al.
Published: (2023)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
by: Zhao, Zaiying, et al.
Published: (2025)
by: Zhao, Zaiying, et al.
Published: (2025)
Face2Diffusion for Fast and Editable Face Personalization
by: Shiohara, Kaede, et al.
Published: (2024)
by: Shiohara, Kaede, et al.
Published: (2024)
ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing Points
by: Okumura, Ryota, et al.
Published: (2025)
by: Okumura, Ryota, et al.
Published: (2025)
Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
by: Tanji, Naoto, et al.
Published: (2025)
by: Tanji, Naoto, et al.
Published: (2025)
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
by: Xiao, Ling, et al.
Published: (2026)
by: Xiao, Ling, et al.
Published: (2026)
ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
Reward Incremental Learning in Text-to-Image Generation
by: Wang, Maorong, et al.
Published: (2024)
by: Wang, Maorong, et al.
Published: (2024)
Language-guided Detection and Mitigation of Unknown Dataset Bias
by: Zhao, Zaiying, et al.
Published: (2024)
by: Zhao, Zaiying, et al.
Published: (2024)
Randomized Autoregressive Visual Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
Parallelized Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
CAR: Controllable Autoregressive Modeling for Visual Generation
by: Yao, Ziyu, et al.
Published: (2024)
by: Yao, Ziyu, et al.
Published: (2024)
Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
by: Sugihara, Tomoya, et al.
Published: (2024)
by: Sugihara, Tomoya, et al.
Published: (2024)
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
by: Sbrolli, Cristian, et al.
Published: (2026)
by: Sbrolli, Cristian, et al.
Published: (2026)
Adversarial Training from Mean Field Perspective
by: Kumano, Soichiro, et al.
Published: (2025)
by: Kumano, Soichiro, et al.
Published: (2025)
Theoretical Understanding of Learning from Adversarial Perturbations
by: Kumano, Soichiro, et al.
Published: (2024)
by: Kumano, Soichiro, et al.
Published: (2024)
Adversarially Pretrained Transformers May Be Universally Robust In-Context Learners
by: Kumano, Soichiro, et al.
Published: (2025)
by: Kumano, Soichiro, et al.
Published: (2025)
Wide Two-Layer Networks can Learn from Adversarial Perturbations
by: Kumano, Soichiro, et al.
Published: (2024)
by: Kumano, Soichiro, et al.
Published: (2024)
Visual Implicit Autoregressive Modeling
by: Jiang, Pengfei, et al.
Published: (2026)
by: Jiang, Pengfei, et al.
Published: (2026)
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Spectral Probing of Feature Upsamplers in 2D-to-3D Scene Reconstruction
by: Xiao, Ling, et al.
Published: (2026)
by: Xiao, Ling, et al.
Published: (2026)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
by: Wang, Junke, et al.
Published: (2025)
by: Wang, Junke, et al.
Published: (2025)
From Sequential to Spatial: Reordering Autoregression for Efficient Visual Generation
by: Wang, Siyang, et al.
Published: (2025)
by: Wang, Siyang, et al.
Published: (2025)
Efficient Conditional Generation on Scale-based Visual Autoregressive Models
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Continual Distillation of Teachers from Different Domains
by: Michel, Nicolas, et al.
Published: (2026)
by: Michel, Nicolas, et al.
Published: (2026)
Dealing with Synthetic Data Contamination in Online Continual Learning
by: Wang, Maorong, et al.
Published: (2024)
by: Wang, Maorong, et al.
Published: (2024)
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
by: Xiao, Ling, et al.
Published: (2022)
by: Xiao, Ling, et al.
Published: (2022)
Next Patch Prediction for Autoregressive Visual Generation
by: Pang, Yatian, et al.
Published: (2024)
by: Pang, Yatian, et al.
Published: (2024)
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
by: Xu, Boxun, et al.
Published: (2025)
by: Xu, Boxun, et al.
Published: (2025)
Merlin:Empowering Multimodal LLMs with Foresight Minds
by: Yu, En, et al.
Published: (2023)
by: Yu, En, et al.
Published: (2023)
Similar Items
-
SCOMatch: Alleviating Overtrusting in Open-set Semi-supervised Learning
by: Wang, Zerun, et al.
Published: (2024) -
Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data
by: Wang, Zerun, et al.
Published: (2024) -
From Obstacles to Resources: Semi-supervised Learning Faces Synthetic Data Contamination
by: Wang, Zerun, et al.
Published: (2024) -
Online Open-set Semi-supervised Object Detection with Dual Competing Head
by: Wang, Zerun, et al.
Published: (2023) -
Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up
by: Huang, Lang, et al.
Published: (2025)