Wolf: Dense Video Captioning with a World Summarization Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Boyi, Zhu, Ligeng, Tian, Ran, Tan, Shuhan, Chen, Yuxiao, Lu, Yao, Cui, Yin, Veer, Sushant, Ehrlich, Max, Philion, Jonah, Weng, Xinshuo, Xue, Fuzhao, Fan, Linxi, Zhu, Yuke, Kautz, Jan, Tao, Andrew, Liu, Ming-Yu, Fidler, Sanja, Ivanovic, Boris, Darrell, Trevor, Malik, Jitendra, Han, Song, Pavone, Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trajeglish: Traffic Modeling as Next-Token Prediction
von: Philion, Jonah, et al.
Veröffentlicht: (2023)
von: Philion, Jonah, et al.
Veröffentlicht: (2023)
Promptable Closed-loop Traffic Simulation
von: Tan, Shuhan, et al.
Veröffentlicht: (2024)
von: Tan, Shuhan, et al.
Veröffentlicht: (2024)
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
von: Gan, Yulu, et al.
Veröffentlicht: (2025)
von: Gan, Yulu, et al.
Veröffentlicht: (2025)
Driving Everywhere with Large Language Model Policy Adaptation
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
von: Tian, Ran, et al.
Veröffentlicht: (2024)
von: Tian, Ran, et al.
Veröffentlicht: (2024)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
von: Zhao, Yue, et al.
Veröffentlicht: (2025)
von: Zhao, Yue, et al.
Veröffentlicht: (2025)
The Case for Negative Data: From Crash Reports to Counterfactuals for Reasonable Driving
von: Patrikar, Jay, et al.
Veröffentlicht: (2025)
von: Patrikar, Jay, et al.
Veröffentlicht: (2025)
RealDrive: Retrieval-Augmented Driving with Diffusion Models
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
System-Level Safety Monitoring and Recovery for Perception Failures in Autonomous Vehicles
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2024)
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2024)
Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
von: Lin, Toru, et al.
Veröffentlicht: (2025)
von: Lin, Toru, et al.
Veröffentlicht: (2025)
Gen-Drive: Enhancing Diffusion Generative Driving Policies with Reward Modeling and Reinforcement Learning Fine-tuning
von: Huang, Zhiyu, et al.
Veröffentlicht: (2024)
von: Huang, Zhiyu, et al.
Veröffentlicht: (2024)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents
von: Grigsby, Jake, et al.
Veröffentlicht: (2023)
von: Grigsby, Jake, et al.
Veröffentlicht: (2023)
Surprise Potential as a Measure of Interactivity in Driving Scenarios
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
RuleFuser: An Evidential Bayes Approach for Rule Injection in Imitation Learned Planners and Predictors for Robustness under Distribution Shifts
von: Patrikar, Jay, et al.
Veröffentlicht: (2024)
von: Patrikar, Jay, et al.
Veröffentlicht: (2024)
LoRD: Adapting Differentiable Driving Policies to Distribution Shifts
von: Diehl, Christopher, et al.
Veröffentlicht: (2024)
von: Diehl, Christopher, et al.
Veröffentlicht: (2024)
Online Aggregation of Trajectory Predictors
von: Tong, Alex, et al.
Veröffentlicht: (2025)
von: Tong, Alex, et al.
Veröffentlicht: (2025)
Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2025)
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2025)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Scaling Vision Pre-Training to 4K Resolution
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
Language-Image Models with 3D Understanding
von: Cho, Jang Hyun, et al.
Veröffentlicht: (2024)
von: Cho, Jang Hyun, et al.
Veröffentlicht: (2024)
Learning Humanoid Locomotion over Challenging Terrain
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
xT: Nested Tokenization for Larger Context in Large Images
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
DTPP: Differentiable Joint Conditional Prediction and Cost Evaluation for Tree Policy Planning in Autonomous Driving
von: Huang, Zhiyu, et al.
Veröffentlicht: (2023)
von: Huang, Zhiyu, et al.
Veröffentlicht: (2023)
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
von: Hirota, Yusuke, et al.
Veröffentlicht: (2025)
von: Hirota, Yusuke, et al.
Veröffentlicht: (2025)
DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features
von: Wang, Letian, et al.
Veröffentlicht: (2024)
von: Wang, Letian, et al.
Veröffentlicht: (2024)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
von: Yang, Jiawei, et al.
Veröffentlicht: (2025)
von: Yang, Jiawei, et al.
Veröffentlicht: (2025)
DreamDrive: Generative 4D Scene Modeling from Street View Images
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
von: Lu, Ziqi, et al.
Veröffentlicht: (2024)
von: Lu, Ziqi, et al.
Veröffentlicht: (2024)
ViR: Towards Efficient Vision Retention Backbones
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
Sim2Val: Leveraging Correlation Across Test Platforms for Variance-Reduced Metric Estimation
von: Luo, Rachel, et al.
Veröffentlicht: (2025)
von: Luo, Rachel, et al.
Veröffentlicht: (2025)
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
von: Sabour, Amirmojtaba, et al.
Veröffentlicht: (2025)
von: Sabour, Amirmojtaba, et al.
Veröffentlicht: (2025)
Align Your Steps: Optimizing Sampling Schedules in Diffusion Models
von: Sabour, Amirmojtaba, et al.
Veröffentlicht: (2024)
von: Sabour, Amirmojtaba, et al.
Veröffentlicht: (2024)
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Accelerating Structured Chain-of-Thought in Autonomous Vehicles
von: Gu, Yi, et al.
Veröffentlicht: (2026)
von: Gu, Yi, et al.
Veröffentlicht: (2026)
Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models
von: Zhang, Zhejun, et al.
Veröffentlicht: (2024)
von: Zhang, Zhejun, et al.
Veröffentlicht: (2024)
CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Latent Chain-of-Thought World Modeling for End-to-End Driving
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
VILA$^2$: VILA Augmented VILA
von: Fang, Yunhao, et al.
Veröffentlicht: (2024)
von: Fang, Yunhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Trajeglish: Traffic Modeling as Next-Token Prediction
von: Philion, Jonah, et al.
Veröffentlicht: (2023) -
Promptable Closed-loop Traffic Simulation
von: Tan, Shuhan, et al.
Veröffentlicht: (2024) -
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
von: Gan, Yulu, et al.
Veröffentlicht: (2025) -
Driving Everywhere with Large Language Model Policy Adaptation
von: Li, Boyi, et al.
Veröffentlicht: (2024) -
Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
von: Tian, Ran, et al.
Veröffentlicht: (2024)