From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wallingford, Matthew, Bhattad, Anand, Kusupati, Aditya, Ramanujan, Vivek, Deitke, Matt, Kakade, Sham, Kembhavi, Aniruddha, Mottaghi, Roozbeh, Ma, Wei-Chiu, Farhadi, Ali |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Matryoshka Representation Learning
par: Kusupati, Aditya, et autres
Publié: (2022)
par: Kusupati, Aditya, et autres
Publié: (2022)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
par: Luo, Rundong, et autres
Publié: (2025)
par: Luo, Rundong, et autres
Publié: (2025)
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
par: Shen, Ethan, et autres
Publié: (2024)
par: Shen, Ethan, et autres
Publié: (2024)
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
par: Soraki, Rustin, et autres
Publié: (2026)
par: Soraki, Rustin, et autres
Publié: (2026)
Posterior Augmented Flow Matching
par: Stoica, George, et autres
Publié: (2026)
par: Stoica, George, et autres
Publié: (2026)
MatFormer: Nested Transformer for Elastic Inference
par: Devvrit, et autres
Publié: (2023)
par: Devvrit, et autres
Publié: (2023)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
par: Gao, Ziqi, et autres
Publié: (2024)
par: Gao, Ziqi, et autres
Publié: (2024)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
par: Bharadhwaj, Homanga, et autres
Publié: (2024)
par: Bharadhwaj, Homanga, et autres
Publié: (2024)
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
par: Ramanujan, Vivek, et autres
Publié: (2024)
par: Ramanujan, Vivek, et autres
Publié: (2024)
Polaris: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair
par: Kakade, Aditya, et autres
Publié: (2026)
par: Kakade, Aditya, et autres
Publié: (2026)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
par: Fan, Xiang, et autres
Publié: (2024)
par: Fan, Xiang, et autres
Publié: (2024)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
par: Yang, Yue, et autres
Publié: (2025)
par: Yang, Yue, et autres
Publié: (2025)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
par: Gupta, Tanmay, et autres
Publié: (2024)
par: Gupta, Tanmay, et autres
Publié: (2024)
The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
par: Geng, Scott, et autres
Publié: (2024)
par: Geng, Scott, et autres
Publié: (2024)
MIMIC: Masked Image Modeling with Image Correspondences
par: Marathe, Kalyani, et autres
Publié: (2023)
par: Marathe, Kalyani, et autres
Publié: (2023)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
par: Salehi, Mohammadreza, et autres
Publié: (2024)
par: Salehi, Mohammadreza, et autres
Publié: (2024)
Imagine360: Immersive 360 Video Generation from Perspective Anchor
par: Tan, Jing, et autres
Publié: (2024)
par: Tan, Jing, et autres
Publié: (2024)
Contrastive Flow Matching
par: Stoica, George, et autres
Publié: (2025)
par: Stoica, George, et autres
Publié: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
par: Brandfonbrener, David, et autres
Publié: (2024)
par: Brandfonbrener, David, et autres
Publié: (2024)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
par: Su, Huangyuan, et autres
Publié: (2025)
par: Su, Huangyuan, et autres
Publié: (2025)
Seeing Fast and Slow: Learning the Flow of Time in Videos
par: Wu, Yen-Siang, et autres
Publié: (2026)
par: Wu, Yen-Siang, et autres
Publié: (2026)
Task Me Anything
par: Zhang, Jieyu, et autres
Publié: (2024)
par: Zhang, Jieyu, et autres
Publié: (2024)
LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting
par: Xing, Xiaoyan, et autres
Publié: (2024)
par: Xing, Xiaoyan, et autres
Publié: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
par: Zheng, Chenhao, et autres
Publié: (2024)
par: Zheng, Chenhao, et autres
Publié: (2024)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
par: Duggal, Shivam, et autres
Publié: (2025)
par: Duggal, Shivam, et autres
Publié: (2025)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
par: Yadav, Tanush, et autres
Publié: (2026)
par: Yadav, Tanush, et autres
Publié: (2026)
Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning
par: Ramrakhya, Ram, et autres
Publié: (2025)
par: Ramrakhya, Ram, et autres
Publié: (2025)
UrbanIR: Large-Scale Urban Scene Inverse Rendering from a Single Video
par: Lin, Chih-Hao, et autres
Publié: (2023)
par: Lin, Chih-Hao, et autres
Publié: (2023)
Preserving Identity with Variational Score for General-purpose 3D Editing
par: Le, Duong H., et autres
Publié: (2024)
par: Le, Duong H., et autres
Publié: (2024)
Generative Blocks World: Moving Things Around in Pictures
par: Vavilala, Vaibhav, et autres
Publié: (2025)
par: Vavilala, Vaibhav, et autres
Publié: (2025)
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
par: Hu, Jiaheng, et autres
Publié: (2024)
par: Hu, Jiaheng, et autres
Publié: (2024)
Algebraic generators of the skein algebra of a surface
par: Santharoubane, Ramanujan
Publié: (2018)
par: Santharoubane, Ramanujan
Publié: (2018)
LOTION: Smoothing the Optimization Landscape for Quantized Training
par: Kwun, Mujin, et autres
Publié: (2025)
par: Kwun, Mujin, et autres
Publié: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
par: Morwani, Depen, et autres
Publié: (2025)
par: Morwani, Depen, et autres
Publié: (2025)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
par: Zhang, Hanlin, et autres
Publié: (2026)
par: Zhang, Hanlin, et autres
Publié: (2026)
GQ-VAE: A gated quantized VAE for learning variable length tokens
par: Datta, Theo, et autres
Publié: (2025)
par: Datta, Theo, et autres
Publié: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
par: Abreu, Natalie, et autres
Publié: (2025)
par: Abreu, Natalie, et autres
Publié: (2025)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
par: Jin, Jikai, et autres
Publié: (2025)
par: Jin, Jikai, et autres
Publié: (2025)
Neuro-Parametric Spectral Classification of Black Hole and Neutron Star X-ray Binary Systems
par: Garg, Akash, et autres
Publié: (2026)
par: Garg, Akash, et autres
Publié: (2026)
Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now
par: Thozhiyoor, Varun Varma, et autres
Publié: (2025)
par: Thozhiyoor, Varun Varma, et autres
Publié: (2025)
Documents similaires
-
Matryoshka Representation Learning
par: Kusupati, Aditya, et autres
Publié: (2022) -
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
par: Luo, Rundong, et autres
Publié: (2025) -
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
par: Shen, Ethan, et autres
Publié: (2024) -
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
par: Soraki, Rustin, et autres
Publié: (2026) -
Posterior Augmented Flow Matching
par: Stoica, George, et autres
Publié: (2026)