FIction: 4D Future Interaction Prediction from Video
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ashutosh, Kumar, Pavlakos, Georgios, Grauman, Kristen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ExpertAF: Expert Actionable Feedback from Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
Learning Skill-Attributes for Transferable Assessment in Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
Learning Object State Changes in Videos: An Open-World Perspective
von: Xue, Zihui, et al.
Veröffentlicht: (2023)
von: Xue, Zihui, et al.
Veröffentlicht: (2023)
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
SportSkills: Physical Skill Learning from Sports Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
Detours for Navigating Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
SkillSight: Efficient First-Person Skill Assessment with Gaze
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
von: An, Joungbin, et al.
Veröffentlicht: (2025)
von: An, Joungbin, et al.
Veröffentlicht: (2025)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2026)
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2026)
Reconstructing Hand-Held Objects in 3D from Images and Videos
von: Wu, Jane, et al.
Veröffentlicht: (2024)
von: Wu, Jane, et al.
Veröffentlicht: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
von: An, Joungbin, et al.
Veröffentlicht: (2026)
von: An, Joungbin, et al.
Veröffentlicht: (2026)
Vid2Coach: Transforming How-To Videos into Task Assistants
von: Huh, Mina, et al.
Veröffentlicht: (2025)
von: Huh, Mina, et al.
Veröffentlicht: (2025)
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
von: Chen, Changan, et al.
Veröffentlicht: (2024)
von: Chen, Changan, et al.
Veröffentlicht: (2024)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
Real3D: Scaling Up Large Reconstruction Models with Real-World Images
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
von: Baid, Ami, et al.
Veröffentlicht: (2026)
von: Baid, Ami, et al.
Veröffentlicht: (2026)
Progress-Aware Video Frame Captioning
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
EgoExo-WM: Unlocking Exo Video for Ego World Models
von: Tran, Danny, et al.
Veröffentlicht: (2026)
von: Tran, Danny, et al.
Veröffentlicht: (2026)
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
von: Luo, Mi, et al.
Veröffentlicht: (2024)
von: Luo, Mi, et al.
Veröffentlicht: (2024)
Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
von: Chatzichristodoulou, Georgios, et al.
Veröffentlicht: (2025)
von: Chatzichristodoulou, Georgios, et al.
Veröffentlicht: (2025)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2026)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2026)
Human detectors are surprisingly powerful reward models
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
SPOC: Spatially-Progressing Object State Change Segmentation in Video
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
Seeing the Arrow of Time in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos
von: Wei, Dingkun, et al.
Veröffentlicht: (2026)
von: Wei, Dingkun, et al.
Veröffentlicht: (2026)
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
von: Luo, Mi, et al.
Veröffentlicht: (2025)
von: Luo, Mi, et al.
Veröffentlicht: (2025)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
Expressive Gaussian Human Avatars from Monocular RGB Video
von: Hu, Hezhen, et al.
Veröffentlicht: (2024)
von: Hu, Hezhen, et al.
Veröffentlicht: (2024)
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
von: Karvounas, Giorgos, et al.
Veröffentlicht: (2025)
von: Karvounas, Giorgos, et al.
Veröffentlicht: (2025)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
Atlas Gaussians Diffusion for 3D Generation
von: Yang, Haitao, et al.
Veröffentlicht: (2024)
von: Yang, Haitao, et al.
Veröffentlicht: (2024)
Few-View Object Reconstruction with Unknown Categories and Camera Poses
von: Jiang, Hanwen, et al.
Veröffentlicht: (2022)
von: Jiang, Hanwen, et al.
Veröffentlicht: (2022)
CoFie: Learning Compact Neural Surface Representations with Coordinate Fields
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
Reconstructing Humans with a Biomechanically Accurate Skeleton
von: Xia, Yan, et al.
Veröffentlicht: (2025)
von: Xia, Yan, et al.
Veröffentlicht: (2025)
Seeing without Pixels: Perception from Camera Trajectories
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Personal Visual Context Learning in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2025)
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ExpertAF: Expert Actionable Feedback from Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024) -
Learning Skill-Attributes for Transferable Assessment in Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025) -
Learning Object State Changes in Videos: An Open-World Perspective
von: Xue, Zihui, et al.
Veröffentlicht: (2023) -
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025) -
SportSkills: Physical Skill Learning from Sports Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)