VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Eskandar, George, Shen, Fengyi, Altillawi, Mohammad, Chen, Dong, Bai, Yang, Yang, Liudi, Liu, Ziyuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
by: Yang, Liudi, et al.
Published: (2026)
by: Yang, Liudi, et al.
Published: (2026)
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
CE-NPBG: Connectivity Enhanced Neural Point-Based Graphics for Novel View Synthesis in Autonomous Driving Scenes
by: Altillawi, Mohammad, et al.
Published: (2025)
by: Altillawi, Mohammad, et al.
Published: (2025)
ConsistentDreamer: View-Consistent Meshes Through Balanced Multi-View Gaussian Optimization
by: Şahin, Onat, et al.
Published: (2025)
by: Şahin, Onat, et al.
Published: (2025)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026)
by: Durante, Zane, et al.
Published: (2026)
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
by: Liu, Penghui, et al.
Published: (2025)
by: Liu, Penghui, et al.
Published: (2025)
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025)
by: Shi, Jingwei, et al.
Published: (2025)
VCA: Video Curious Agent for Long Video Understanding
by: Yang, Zeyuan, et al.
Published: (2024)
by: Yang, Zeyuan, et al.
Published: (2024)
View while Moving: Efficient Video Recognition in Long-untrimmed Videos
by: Tian, Ye, et al.
Published: (2023)
by: Tian, Ye, et al.
Published: (2023)
MultiWorld: Scalable Multi-Agent Multi-View Video World Models
by: Wu, Haoyu, et al.
Published: (2026)
by: Wu, Haoyu, et al.
Published: (2026)
Lifelong 3D Mapping Framework for Hand-held & Robot-mounted LiDAR Mapping Systems
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
by: Jiang, Songtao, et al.
Published: (2026)
by: Jiang, Songtao, et al.
Published: (2026)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
Agent-based Video Trimming
by: Yang, Lingfeng, et al.
Published: (2024)
by: Yang, Lingfeng, et al.
Published: (2024)
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
by: Zhang, Peixuan, et al.
Published: (2026)
by: Zhang, Peixuan, et al.
Published: (2026)
MV2MAE: Multi-View Video Masked Autoencoders
by: Shah, Ketul, et al.
Published: (2024)
by: Shah, Ketul, et al.
Published: (2024)
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
by: Yin, Yufei, et al.
Published: (2026)
by: Yin, Yufei, et al.
Published: (2026)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
by: Yang, Yuhang, et al.
Published: (2025)
by: Yang, Yuhang, et al.
Published: (2025)
YingVideo-MV: Music-Driven Multi-Stage Video Generation
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
by: Li, Peiyan, et al.
Published: (2026)
by: Li, Peiyan, et al.
Published: (2026)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
by: Bai, Purui, et al.
Published: (2026)
by: Bai, Purui, et al.
Published: (2026)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
by: Liang, Feng, et al.
Published: (2025)
by: Liang, Feng, et al.
Published: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
by: Yang, Xiangpeng, et al.
Published: (2025)
by: Yang, Xiangpeng, et al.
Published: (2025)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
by: Mei, Jianbiao, et al.
Published: (2024)
by: Mei, Jianbiao, et al.
Published: (2024)
HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion
by: Chang, Di, et al.
Published: (2026)
by: Chang, Di, et al.
Published: (2026)
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025)
by: Zhang, Yehang, et al.
Published: (2025)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
DreamJourney: Perpetual View Generation with Video Diffusion Models
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
by: Zuo, Qi, et al.
Published: (2024)
by: Zuo, Qi, et al.
Published: (2024)
Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
by: Bai, Haoran, et al.
Published: (2025)
by: Bai, Haoran, et al.
Published: (2025)
Similar Items
-
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
by: Yang, Liudi, et al.
Published: (2025) -
DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
by: Bai, Yang, et al.
Published: (2025) -
RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping
by: Bai, Yang, et al.
Published: (2025) -
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
by: Yang, Liudi, et al.
Published: (2026) -
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
by: Yang, Liudi, et al.
Published: (2025)