VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bartoccioni, Florent, Ramzi, Elias, Besnier, Victor, Venkataramanan, Shashanka, Vu, Tuan-Hung, Xu, Yihong, Chambon, Loick, Gidaris, Spyros, Odabas, Serkan, Hurych, David, Marlet, Renaud, Boulch, Alexandre, Chen, Mickael, Zablocki, Éloi, Bursuc, Andrei, Valle, Eduardo, Cord, Matthieu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GaussRender: Learning 3D Occupancy with Gaussian Rendering
von: Chambon, Loïck, et al.
Veröffentlicht: (2025)
von: Chambon, Loïck, et al.
Veröffentlicht: (2025)
PointBeV: A Sparse Approach to BeV Predictions
von: Chambon, Loick, et al.
Veröffentlicht: (2023)
von: Chambon, Loick, et al.
Veröffentlicht: (2023)
Valeo4Cast: A Modular Approach to End-to-End Forecasting
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
Driving on Registers
von: Kirby, Ellington, et al.
Veröffentlicht: (2026)
von: Kirby, Ellington, et al.
Veröffentlicht: (2026)
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
von: Chambon, Loick, et al.
Veröffentlicht: (2025)
von: Chambon, Loick, et al.
Veröffentlicht: (2025)
Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?
von: Xu, Yihong, et al.
Veröffentlicht: (2023)
von: Xu, Yihong, et al.
Veröffentlicht: (2023)
OccFeat: Self-supervised Occupancy Feature Prediction for Pretraining BEV Segmentation Networks
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2024)
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2024)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
Three Pillars improving Vision Foundation Model Distillation for Lidar
von: Puy, Gilles, et al.
Veröffentlicht: (2023)
von: Puy, Gilles, et al.
Veröffentlicht: (2023)
Annealed Winner-Takes-All for Motion Forecasting
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
von: Puy, Gilles, et al.
Veröffentlicht: (2026)
von: Puy, Gilles, et al.
Veröffentlicht: (2026)
Halton Scheduler For Masked Generative Image Transformer
von: Besnier, Victor, et al.
Veröffentlicht: (2025)
von: Besnier, Victor, et al.
Veröffentlicht: (2025)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
BIGFix: Bidirectional Image Generation with Token Fixing
von: Besnier, Victor, et al.
Veröffentlicht: (2025)
von: Besnier, Victor, et al.
Veröffentlicht: (2025)
PPT: Pretraining with Pseudo-Labeled Trajectories for Motion Forecasting
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
IPA: An Information-Reconstructive Input Projection Framework for Efficient Foundation Model Adaptation
von: Yin, Yuan, et al.
Veröffentlicht: (2025)
von: Yin, Yuan, et al.
Veröffentlicht: (2025)
ReGentS: Real-World Safety-Critical Driving Scenario Generation Made Stable
von: Yin, Yuan, et al.
Veröffentlicht: (2024)
von: Yin, Yuan, et al.
Veröffentlicht: (2024)
LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension
von: Cardiel, Amaia, et al.
Veröffentlicht: (2024)
von: Cardiel, Amaia, et al.
Veröffentlicht: (2024)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
von: Vobecky, Antonin, et al.
Veröffentlicht: (2022)
von: Vobecky, Antonin, et al.
Veröffentlicht: (2022)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
von: Vobecky, Antonin, et al.
Veröffentlicht: (2024)
von: Vobecky, Antonin, et al.
Veröffentlicht: (2024)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
von: Gidaris, Spyros, et al.
Veröffentlicht: (2023)
von: Gidaris, Spyros, et al.
Veröffentlicht: (2023)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
To VaR, or Not to VaR, That is the Question
von: Olkhov, Victor
Veröffentlicht: (2021)
von: Olkhov, Victor
Veröffentlicht: (2021)
DRIV-EX: Counterfactual Explanations for Driving LLMs
von: Cardiel, Amaia, et al.
Veröffentlicht: (2026)
von: Cardiel, Amaia, et al.
Veröffentlicht: (2026)
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2025)
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2025)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2026)
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2026)
Mojtababzp/VaPOrS: VaPOrS v1.0.1
von: Mojtaba Bezaatpour
Veröffentlicht: (2025)
von: Mojtaba Bezaatpour
Veröffentlicht: (2025)
MAD: Motion Appearance Decoupling for efficient Driving World Models
von: Rahimi, Ahmad, et al.
Veröffentlicht: (2026)
von: Rahimi, Ahmad, et al.
Veröffentlicht: (2026)
LOGen: Toward Lidar Object Generation by Point Diffusion
von: Kirby, Ellington, et al.
Veröffentlicht: (2024)
von: Kirby, Ellington, et al.
Veröffentlicht: (2024)
Is clustering enough for LiDAR instance segmentation? A state-of-the-art training-free baseline
von: Sautier, Corentin, et al.
Veröffentlicht: (2025)
von: Sautier, Corentin, et al.
Veröffentlicht: (2025)
UNIT: Unsupervised Online Instance Segmentation through Time
von: Sautier, Corentin, et al.
Veröffentlicht: (2024)
von: Sautier, Corentin, et al.
Veröffentlicht: (2024)
VaEIN3.1‐VaERF057‐VaFBA1 Module Positively Regulates Cold Tolerance by Accumulating Soluble Sugar in Grapevine
von: Huimin Zhou, et al.
Veröffentlicht: (2025)
von: Huimin Zhou, et al.
Veröffentlicht: (2025)
Supervised Anomaly Detection for Complex Industrial Images
von: Baitieva, Aimira, et al.
Veröffentlicht: (2024)
von: Baitieva, Aimira, et al.
Veröffentlicht: (2024)
Hawaiki Nui Va'a
Veröffentlicht: (1995)
Veröffentlicht: (1995)
Value at Risk (VaR)
von: Juan Gaytán Cortés
Veröffentlicht: (2022)
von: Juan Gaytán Cortés
Veröffentlicht: (2022)
LiDPM: Rethinking Point Diffusion for Lidar Scene Completion
von: Martyniuk, Tetiana, et al.
Veröffentlicht: (2025)
von: Martyniuk, Tetiana, et al.
Veröffentlicht: (2025)
JAFAR: Jack up Any Feature at Any Resolution
von: Couairon, Paul, et al.
Veröffentlicht: (2025)
von: Couairon, Paul, et al.
Veröffentlicht: (2025)
GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers
von: Zablocki, Éloi, et al.
Veröffentlicht: (2024)
von: Zablocki, Éloi, et al.
Veröffentlicht: (2024)
Plans for the 1942 season in Norfolk, Va.
von: NA
Veröffentlicht: (1942)
von: NA
Veröffentlicht: (1942)
Coevolving Representations in Joint Image-Feature Diffusion
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2026)
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GaussRender: Learning 3D Occupancy with Gaussian Rendering
von: Chambon, Loïck, et al.
Veröffentlicht: (2025) -
PointBeV: A Sparse Approach to BeV Predictions
von: Chambon, Loick, et al.
Veröffentlicht: (2023) -
Valeo4Cast: A Modular Approach to End-to-End Forecasting
von: Xu, Yihong, et al.
Veröffentlicht: (2024) -
Driving on Registers
von: Kirby, Ellington, et al.
Veröffentlicht: (2026) -
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
von: Chambon, Loick, et al.
Veröffentlicht: (2025)