An Empirical Study of Autoregressive Pre-training from Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Rajasegaran, Jathushan, Radosavovic, Ilija, Ravishankar, Rahul, Gandelsman, Yossi, Feichtenhofer, Christoph, Malik, Jitendra |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Properties of Diffusion Models for Perceptual Tasks
by: Ravishankar, Rahul, et al.
Published: (2024)
by: Ravishankar, Rahul, et al.
Published: (2024)
Synthesizing Moving People with 3D Control
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Gaussian Masked Autoencoders
by: Rajasegaran, Jathushan, et al.
Published: (2025)
by: Rajasegaran, Jathushan, et al.
Published: (2025)
Poly-Autoregressive Prediction for Modeling Interactions
by: Thakkar, Neerja, et al.
Published: (2025)
by: Thakkar, Neerja, et al.
Published: (2025)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
by: Bu, Edmund, et al.
Published: (2025)
by: Bu, Edmund, et al.
Published: (2025)
Tracking by Predicting 3-D Gaussians Over Time
by: Baranwal, Tanish, et al.
Published: (2025)
by: Baranwal, Tanish, et al.
Published: (2025)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
Synergy and Synchrony in Couple Dances
by: Maluleke, Vongani, et al.
Published: (2024)
by: Maluleke, Vongani, et al.
Published: (2024)
Generative Pre-trained Autoregressive Diffusion Transformer
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
Learning Video Representations without Natural Videos
by: Yu, Xueyang, et al.
Published: (2024)
by: Yu, Xueyang, et al.
Published: (2024)
Jailbreaking Vision-Language Models Through the Visual Modality
by: Azulay, Aharon, et al.
Published: (2026)
by: Azulay, Aharon, et al.
Published: (2026)
FewShotNeRF: Meta-Learning-based Novel View Synthesis for Rapid Scene-Specific Adaptation
by: Sivakumar, Piraveen, et al.
Published: (2024)
by: Sivakumar, Piraveen, et al.
Published: (2024)
The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
by: Ekin, Yigit, et al.
Published: (2026)
by: Ekin, Yigit, et al.
Published: (2026)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
VideoQA in the Era of LLMs: An Empirical Study
by: Xiao, Junbin, et al.
Published: (2024)
by: Xiao, Junbin, et al.
Published: (2024)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Speculative Decoding for Autoregressive Video Generation
by: Hu, Yuezhou, et al.
Published: (2026)
by: Hu, Yuezhou, et al.
Published: (2026)
Comment-aided Video-Language Alignment via Contrastive Pre-training for Short-form Video Humor Detection
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
MAGI-1: Autoregressive Video Generation at Scale
by: ai, Sand., et al.
Published: (2025)
by: ai, Sand., et al.
Published: (2025)
Fast Autoregressive Video Generation with Diagonal Decoding
by: Ye, Yang, et al.
Published: (2025)
by: Ye, Yang, et al.
Published: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Adapting VACE for Real-Time Autoregressive Video Diffusion
by: Fosdick, Ryan
Published: (2026)
by: Fosdick, Ryan
Published: (2026)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
xT: Nested Tokenization for Larger Context in Large Images
by: Gupta, Ritwik, et al.
Published: (2024)
by: Gupta, Ritwik, et al.
Published: (2024)
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
by: Li, Lingen, et al.
Published: (2026)
by: Li, Lingen, et al.
Published: (2026)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024)
by: Wang, Hanyu, et al.
Published: (2024)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
by: Tian, Ran, et al.
Published: (2023)
by: Tian, Ran, et al.
Published: (2023)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
by: Yesiltepe, Hidir, et al.
Published: (2026)
by: Yesiltepe, Hidir, et al.
Published: (2026)
Efficient Pre-training for Localized Instruction Generation of Videos
by: Batra, Anil, et al.
Published: (2023)
by: Batra, Anil, et al.
Published: (2023)
Interpreting the Second-Order Effects of Neurons in CLIP
by: Gandelsman, Yossi, et al.
Published: (2024)
by: Gandelsman, Yossi, et al.
Published: (2024)
Lightweight, Pre-trained Transformers for Remote Sensing Timeseries
by: Tseng, Gabriel, et al.
Published: (2023)
by: Tseng, Gabriel, et al.
Published: (2023)
One-Forcing: Towards Stable One-Step Autoregressive Video Generation
by: Feng, Jiaqi, et al.
Published: (2026)
by: Feng, Jiaqi, et al.
Published: (2026)
Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity
by: Tian, Jiahao, et al.
Published: (2026)
by: Tian, Jiahao, et al.
Published: (2026)
A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency
by: Long, Do Xuan, et al.
Published: (2026)
by: Long, Do Xuan, et al.
Published: (2026)
Similar Items
-
Scaling Properties of Diffusion Models for Perceptual Tasks
by: Ravishankar, Rahul, et al.
Published: (2024) -
Synthesizing Moving People with 3D Control
by: Li, Boyi, et al.
Published: (2024) -
Gaussian Masked Autoencoders
by: Rajasegaran, Jathushan, et al.
Published: (2025) -
Poly-Autoregressive Prediction for Modeling Interactions
by: Thakkar, Neerja, et al.
Published: (2025) -
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)