Spanning Tree Autoregressive Visual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Sangkyu, Lee, Changho, Han, Janghoon, Song, Hosung, You, Tackgeun, Lim, Hwasup, Choi, Stanley Jungkyu, Lee, Honglak, Yu, Youngjae |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
MASS: Overcoming Language Bias in Image-Text Matching
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning
by: Han, Janghoon, et al.
Published: (2024)
by: Han, Janghoon, et al.
Published: (2024)
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
by: Han, Janghoon, et al.
Published: (2025)
by: Han, Janghoon, et al.
Published: (2025)
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
by: Lee, Changho, et al.
Published: (2024)
by: Lee, Changho, et al.
Published: (2024)
HFI: A unified framework for training-free detection and implicit watermarking of latent diffusion model generated images
by: Choi, Sungik, et al.
Published: (2024)
by: Choi, Sungik, et al.
Published: (2024)
Bidirectional Temporal Diffusion Model for Temporally Consistent Human Animation
by: Adiya, Tserendorj, et al.
Published: (2023)
by: Adiya, Tserendorj, et al.
Published: (2023)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
by: Choi, Changho, et al.
Published: (2025)
by: Choi, Changho, et al.
Published: (2025)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025)
by: Cho, Janghoon, et al.
Published: (2025)
View Selection for 3D Captioning via Diffusion Ranking
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
Improving Real-Time Omnidirectional 3D Multi-Person Human Pose Estimation with People Matching and Unsupervised 2D-3D Lifting
by: Knap, Pawel, et al.
Published: (2024)
by: Knap, Pawel, et al.
Published: (2024)
SemCity: Semantic Scene Generation with Triplane Diffusion
by: Lee, Jumin, et al.
Published: (2024)
by: Lee, Jumin, et al.
Published: (2024)
LatentSwap: An Efficient Latent Code Mapping Framework for Face Swapping
by: Choi, Changho, et al.
Published: (2024)
by: Choi, Changho, et al.
Published: (2024)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Fine-Tuning Visual Autoregressive Models for Subject-Driven Generation
by: Chung, Jiwoo, et al.
Published: (2025)
by: Chung, Jiwoo, et al.
Published: (2025)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
by: Chu, Sanghyeok, et al.
Published: (2026)
by: Chu, Sanghyeok, et al.
Published: (2026)
Generalized Contrastive Learning for Universal Multimodal Retrieval
by: Lee, Jungsoo, et al.
Published: (2025)
by: Lee, Jungsoo, et al.
Published: (2025)
MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
by: Choi, Changho, et al.
Published: (2025)
by: Choi, Changho, et al.
Published: (2025)
MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents
by: Kwon, Minkyung, et al.
Published: (2026)
by: Kwon, Minkyung, et al.
Published: (2026)
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing
by: Lionar, Stefan, et al.
Published: (2025)
by: Lionar, Stefan, et al.
Published: (2025)
Parallelized Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
Randomized Autoregressive Visual Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
by: Kim, Jeongin, et al.
Published: (2025)
by: Kim, Jeongin, et al.
Published: (2025)
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
by: Baek, Junyeob, et al.
Published: (2025)
by: Baek, Junyeob, et al.
Published: (2025)
Visually Dehallucinative Instruction Generation
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
by: Zhang, Jinhua, et al.
Published: (2025)
by: Zhang, Jinhua, et al.
Published: (2025)
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
by: Park, Mingue, et al.
Published: (2025)
by: Park, Mingue, et al.
Published: (2025)
Next Patch Prediction for Autoregressive Visual Generation
by: Pang, Yatian, et al.
Published: (2024)
by: Pang, Yatian, et al.
Published: (2024)
It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models
by: Choi, Jaeha, et al.
Published: (2026)
by: Choi, Jaeha, et al.
Published: (2026)
VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Mirai: Autoregressive Visual Generation Needs Foresight
by: Yu, Yonghao, et al.
Published: (2026)
by: Yu, Yonghao, et al.
Published: (2026)
Similar Items
-
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025) -
MASS: Overcoming Language Bias in Image-Text Matching
by: Chung, Jiwan, et al.
Published: (2025) -
Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning
by: Han, Janghoon, et al.
Published: (2024) -
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
by: Han, Janghoon, et al.
Published: (2025) -
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
by: Lee, Changho, et al.
Published: (2024)