GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Souček, Tomáš, Damen, Dima, Wray, Michael, Laptev, Ivan, Sivic, Josef |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
by: Souček, Tomáš, et al.
Published: (2024)
by: Souček, Tomáš, et al.
Published: (2024)
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025)
by: Flanagan, Kevin, et al.
Published: (2025)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)
by: Bansal, Siddhant, et al.
Published: (2024)
Video Editing for Video Retrieval
by: Zhu, Bin, et al.
Published: (2024)
by: Zhu, Bin, et al.
Published: (2024)
Beyond Caption-Based Queries for Video Moment Retrieval
by: Pujol-Perich, David, et al.
Published: (2026)
by: Pujol-Perich, David, et al.
Published: (2026)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
by: Zhu, Zhifan, et al.
Published: (2023)
by: Zhu, Zhifan, et al.
Published: (2023)
6D Object Pose Tracking in Internet Videos for Robotic Manipulation
by: Ponimatkin, Georgy, et al.
Published: (2025)
by: Ponimatkin, Georgy, et al.
Published: (2025)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
by: Sinha, Saptarshi, et al.
Published: (2024)
by: Sinha, Saptarshi, et al.
Published: (2024)
EgoPoints: Advancing Point Tracking for Egocentric Videos
by: Darkhalil, Ahmad, et al.
Published: (2024)
by: Darkhalil, Ahmad, et al.
Published: (2024)
Video, How Do Your Tokens Merge?
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
by: Chalk, Jacob, et al.
Published: (2024)
by: Chalk, Jacob, et al.
Published: (2024)
PointSt3R: Point Tracking through 3D Grounded Correspondence
by: Guerrier, Rhodri, et al.
Published: (2025)
by: Guerrier, Rhodri, et al.
Published: (2025)
InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos
by: Zhang, Yangsong, et al.
Published: (2025)
by: Zhang, Yangsong, et al.
Published: (2025)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024)
by: Plizzari, Chiara, et al.
Published: (2024)
The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation
by: Hatano, Masashi, et al.
Published: (2025)
by: Hatano, Masashi, et al.
Published: (2025)
AMEGO: Active Memory from long EGOcentric videos
by: Goletto, Gabriele, et al.
Published: (2024)
by: Goletto, Gabriele, et al.
Published: (2024)
Learning from Streaming Video with Orthogonal Gradients
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
by: Bardhan, Jai, et al.
Published: (2026)
by: Bardhan, Jai, et al.
Published: (2026)
Segmenting Collision Sound Sources in Egocentric Videos
by: Parida, Kranti Kumar, et al.
Published: (2025)
by: Parida, Kranti Kumar, et al.
Published: (2025)
EditDuet: A Multi-Agent System for Video Non-Linear Editing
by: Sandoval-Castaneda, Marcelo, et al.
Published: (2025)
by: Sandoval-Castaneda, Marcelo, et al.
Published: (2025)
A Video Is Not Worth a Thousand Words
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
EPIC Fields: Marrying 3D Geometry and Video Understanding
by: Tschernezki, Vadim, et al.
Published: (2023)
by: Tschernezki, Vadim, et al.
Published: (2023)
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023)
by: Chen, Shoufa, et al.
Published: (2023)
GenCompositor: Generative Video Compositing with Diffusion Transformer
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
Unique Lives, Shared World: Learning from Single-Life Videos
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
by: Mikeštíková, Anna Šárová, et al.
Published: (2025)
by: Mikeštíková, Anna Šárová, et al.
Published: (2025)
HD-EPIC: A Highly-Detailed Egocentric Video Dataset
by: Perrett, Toby, et al.
Published: (2025)
by: Perrett, Toby, et al.
Published: (2025)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
by: Romero, David, et al.
Published: (2025)
by: Romero, David, et al.
Published: (2025)
Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
by: Hatano, Masashi, et al.
Published: (2025)
by: Hatano, Masashi, et al.
Published: (2025)
Similar Items
-
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
by: Souček, Tomáš, et al.
Published: (2024) -
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
by: Fragomeni, Adriano, et al.
Published: (2025) -
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025) -
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025) -
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)