Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Maluleke, Vongani H., Horiuchi, Kie, Wilken, Lea, Ng, Evonne, Malik, Jitendra, Kanazawa, Angjoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synergy and Synchrony in Couple Dances
by: Maluleke, Vongani, et al.
Published: (2024)
by: Maluleke, Vongani, et al.
Published: (2024)
Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos
by: Yang, Gengshan, et al.
Published: (2024)
by: Yang, Gengshan, et al.
Published: (2024)
Human-level 3D shape perception emerges from multi-view learning
by: Bonnen, Tyler, et al.
Published: (2026)
by: Bonnen, Tyler, et al.
Published: (2026)
Reconstructing People, Places, and Cameras
by: Müller, Lea, et al.
Published: (2024)
by: Müller, Lea, et al.
Published: (2024)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
by: Ng, Evonne, et al.
Published: (2024)
by: Ng, Evonne, et al.
Published: (2024)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
by: Kerr, Justin, et al.
Published: (2024)
by: Kerr, Justin, et al.
Published: (2024)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025)
by: Kerr, Justin, et al.
Published: (2025)
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
FIMP: Future Interaction Modeling for Multi-Agent Motion Prediction
by: Woo, Sungmin, et al.
Published: (2024)
by: Woo, Sungmin, et al.
Published: (2024)
Viser: Imperative, Web-based 3D Visualization in Python
by: Yi, Brent, et al.
Published: (2025)
by: Yi, Brent, et al.
Published: (2025)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
by: Huang, Huang, et al.
Published: (2025)
by: Huang, Huang, et al.
Published: (2025)
Estimating Body and Hand Motion in an Ego-sensed World
by: Yi, Brent, et al.
Published: (2024)
by: Yi, Brent, et al.
Published: (2024)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
Splatfacto-W: A Nerfstudio Implementation of Gaussian Splatting for Unconstrained Photo Collections
by: Xu, Congrong, et al.
Published: (2024)
by: Xu, Congrong, et al.
Published: (2024)
TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
by: Ze, Yanjie, et al.
Published: (2025)
by: Ze, Yanjie, et al.
Published: (2025)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
by: Tian, Ran, et al.
Published: (2023)
by: Tian, Ran, et al.
Published: (2023)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Decentralized Diffusion Models
by: McAllister, David, et al.
Published: (2025)
by: McAllister, David, et al.
Published: (2025)
SOAR: Self-Occluded Avatar Recovery from a Single Video In the Wild
by: Pan, Zhuoyang, et al.
Published: (2024)
by: Pan, Zhuoyang, et al.
Published: (2024)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
by: Zhao, Ruiteng, et al.
Published: (2026)
by: Zhao, Ruiteng, et al.
Published: (2026)
Image-to-Force Estimation for Soft Tissue Interaction in Robotic-Assisted Surgery Using Structured Light
by: Wang, Jiayin, et al.
Published: (2025)
by: Wang, Jiayin, et al.
Published: (2025)
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
by: Yu, Jiawen, et al.
Published: (2025)
by: Yu, Jiawen, et al.
Published: (2025)
DRoPE: Directional Rotary Position Embedding for Efficient Agent Interaction Modeling
by: Zhao, Jianbo, et al.
Published: (2025)
by: Zhao, Jianbo, et al.
Published: (2025)
SPIDER: Scalable Physics-Informed Dexterous Retargeting
by: Pan, Chaoyi, et al.
Published: (2025)
by: Pan, Chaoyi, et al.
Published: (2025)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025)
by: Wang, Renhao, et al.
Published: (2025)
Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation
by: Konstantinidis, Fabian, et al.
Published: (2025)
by: Konstantinidis, Fabian, et al.
Published: (2025)
World Model for Robot Learning: A Comprehensive Survey
by: Hou, Bohan, et al.
Published: (2026)
by: Hou, Bohan, et al.
Published: (2026)
Large Video Planner Enables Generalizable Robot Control
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024)
by: Subramanian, Sanjay, et al.
Published: (2024)
AgentAlign: Misalignment-Adapted Multi-Agent Perception for Resilient Inter-Agent Sensor Correlations
by: Meng, Zonglin, et al.
Published: (2024)
by: Meng, Zonglin, et al.
Published: (2024)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
by: Cui, Jieming, et al.
Published: (2024)
by: Cui, Jieming, et al.
Published: (2024)
Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos
by: Li, Can, et al.
Published: (2026)
by: Li, Can, et al.
Published: (2026)
FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
A Model-based Visual Contact Localization and Force Sensing System for Compliant Robotic Grippers
by: Zuo, Kaiwen, et al.
Published: (2026)
by: Zuo, Kaiwen, et al.
Published: (2026)
Similar Items
-
Synergy and Synchrony in Couple Dances
by: Maluleke, Vongani, et al.
Published: (2024) -
Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos
by: Yang, Gengshan, et al.
Published: (2024) -
Human-level 3D shape perception emerges from multi-view learning
by: Bonnen, Tyler, et al.
Published: (2026) -
Reconstructing People, Places, and Cameras
by: Müller, Lea, et al.
Published: (2024) -
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)