Saved in:
| Main Authors: | Inbasekar, Karthik, Rom, Guy, Shlomovits, Omer |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.03475 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An End-to-End Framework for Video Multi-Person Pose Estimation
by: Wei, Zhihong
Published: (2025)
by: Wei, Zhihong
Published: (2025)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
NuiWorld: Exploring a Scalable Framework for End-to-End Controllable World Generation
by: Lee, Han-Hung, et al.
Published: (2026)
by: Lee, Han-Hung, et al.
Published: (2026)
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
by: Zhang, Jinqing, et al.
Published: (2026)
by: Zhang, Jinqing, et al.
Published: (2026)
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)
by: Li, Yingyan, et al.
Published: (2024)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation
by: Ma, Enhui, et al.
Published: (2024)
by: Ma, Enhui, et al.
Published: (2024)
BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving
by: Mohan, Karthik, et al.
Published: (2025)
by: Mohan, Karthik, et al.
Published: (2025)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
Divide and Merge: Motion and Semantic Learning in End-to-End Autonomous Driving
by: Shen, Yinzhe, et al.
Published: (2025)
by: Shen, Yinzhe, et al.
Published: (2025)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025)
by: Yu, Yonghui, et al.
Published: (2025)
Latent Chain-of-Thought World Modeling for End-to-End Driving
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
by: Zheng, Chaoda, et al.
Published: (2026)
by: Zheng, Chaoda, et al.
Published: (2026)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
STORM: End-to-End Referring Multi-Object Tracking in Videos
by: Lu, Zijia, et al.
Published: (2026)
by: Lu, Zijia, et al.
Published: (2026)
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
by: Sheng, Zihao, et al.
Published: (2026)
by: Sheng, Zihao, et al.
Published: (2026)
End-to-End Driving with Online Trajectory Evaluation via BEV World Model
by: Li, Yingyan, et al.
Published: (2025)
by: Li, Yingyan, et al.
Published: (2025)
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
by: Wang, Ruoyu, et al.
Published: (2026)
by: Wang, Ruoyu, et al.
Published: (2026)
World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
by: Guan, Yanchen, et al.
Published: (2025)
by: Guan, Yanchen, et al.
Published: (2025)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB Videos
by: Zhong, Pengzhi, et al.
Published: (2025)
by: Zhong, Pengzhi, et al.
Published: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
by: Zheng, Yupeng, et al.
Published: (2025)
by: Zheng, Yupeng, et al.
Published: (2025)
LMVC: An End-to-End Learned Multiview Video Coding Framework
by: Sheng, Xihua, et al.
Published: (2025)
by: Sheng, Xihua, et al.
Published: (2025)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
DualAD: Disentangling the Dynamic and Static World for End-to-End Driving
by: Doll, Simon, et al.
Published: (2024)
by: Doll, Simon, et al.
Published: (2024)
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
by: Shao, Hao, et al.
Published: (2026)
by: Shao, Hao, et al.
Published: (2026)
From Category to Scenery: An End-to-End Framework for Multi-Person Human-Object Interaction Recognition in Videos
by: Qiao, Tanqiu, et al.
Published: (2024)
by: Qiao, Tanqiu, et al.
Published: (2024)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
by: Hu, Yihan, et al.
Published: (2025)
by: Hu, Yihan, et al.
Published: (2025)
HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models
by: Cho, Hoonhee, et al.
Published: (2026)
by: Cho, Hoonhee, et al.
Published: (2026)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing
by: Li, Shuo, et al.
Published: (2026)
by: Li, Shuo, et al.
Published: (2026)
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
by: Lin, Hongbin, et al.
Published: (2025)
by: Lin, Hongbin, et al.
Published: (2025)
Similar Items
-
An End-to-End Framework for Video Multi-Person Pose Estimation
by: Wei, Zhihong
Published: (2025) -
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025) -
NuiWorld: Exploring a Scalable Framework for End-to-End Controllable World Generation
by: Lee, Han-Hung, et al.
Published: (2026) -
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
by: Zhang, Jinqing, et al.
Published: (2026) -
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)