Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiahao, Cherian, Anoop, Liu, Yanbin, Ben-Shabat, Yizhak, Rodriguez, Cristian, Gould, Stephen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporally Grounding Instructional Diagrams in Unconstrained Videos
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
3DInAction: Understanding Human Actions in 3D Point Clouds
by: Ben-Shabat, Yizhak, et al.
Published: (2023)
by: Ben-Shabat, Yizhak, et al.
Published: (2023)
Neural Experts: Mixture of Experts for Implicit Neural Representations
by: Ben-Shabat, Yizhak, et al.
Published: (2024)
by: Ben-Shabat, Yizhak, et al.
Published: (2024)
VI3NR: Variance Informed Initialization for Implicit Neural Representations
by: Koneputugodage, Chamin Hewa, et al.
Published: (2025)
by: Koneputugodage, Chamin Hewa, et al.
Published: (2025)
GraVoS: Voxel Selection for 3D Point-Cloud Detection
by: Shrout, Oren, et al.
Published: (2022)
by: Shrout, Oren, et al.
Published: (2022)
PatchContrast: Self-Supervised Pre-training for 3D Object Detection
by: Shrout, Oren, et al.
Published: (2023)
by: Shrout, Oren, et al.
Published: (2023)
Step Differences in Instructional Video
by: Nagarajan, Tushar, et al.
Published: (2024)
by: Nagarajan, Tushar, et al.
Published: (2024)
RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
ComplexVAD: Detecting Interaction Anomalies in Video
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models
by: Mumcu, Furkan, et al.
Published: (2026)
by: Mumcu, Furkan, et al.
Published: (2026)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
by: Souček, Tomáš, et al.
Published: (2024)
by: Souček, Tomáš, et al.
Published: (2024)
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
by: Zhao, Qinyu, et al.
Published: (2025)
by: Zhao, Qinyu, et al.
Published: (2025)
Less is More: Improving Motion Diffusion Models with Sparse Keyframes
by: Bae, Jinseok, et al.
Published: (2025)
by: Bae, Jinseok, et al.
Published: (2025)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
StepAL: Step-aware Active Learning for Cataract Surgical Videos
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
by: Liang, Zhanhao, et al.
Published: (2026)
by: Liang, Zhanhao, et al.
Published: (2026)
MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
by: Kogashi, Kaen, et al.
Published: (2025)
by: Kogashi, Kaen, et al.
Published: (2025)
Step by Step Network
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
by: Ge, Mingji, et al.
Published: (2026)
by: Ge, Mingji, et al.
Published: (2026)
LLM-Guided Agentic Object Detection for Open-World Understanding
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
A Step to Decouple Optimization in 3DGS
by: Ding, Renjie, et al.
Published: (2026)
by: Ding, Renjie, et al.
Published: (2026)
Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
by: Liu, Yong, et al.
Published: (2025)
by: Liu, Yong, et al.
Published: (2025)
Training-free Online Video Step Grounding
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation
by: Chu, Ernie, et al.
Published: (2026)
by: Chu, Ernie, et al.
Published: (2026)
Asymmetric VAE for One-Step Video Super-Resolution Acceleration
by: Li, Jianze, et al.
Published: (2025)
by: Li, Jianze, et al.
Published: (2025)
DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization
by: Ding, Zihan, et al.
Published: (2024)
by: Ding, Zihan, et al.
Published: (2024)
Watch Your Steps: Local Image and Scene Editing by Text Instructions
by: Mirzaei, Ashkan, et al.
Published: (2023)
by: Mirzaei, Ashkan, et al.
Published: (2023)
Phased One-Step Adversarial Equilibrium for Video Diffusion Models
by: Cheng, Jiaxiang, et al.
Published: (2025)
by: Cheng, Jiaxiang, et al.
Published: (2025)
Snakes and Ladders: Two Steps Up for VideoMamba
by: Lu, Hui, et al.
Published: (2024)
by: Lu, Hui, et al.
Published: (2024)
Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
by: Zhang, Ziyi, et al.
Published: (2024)
by: Zhang, Ziyi, et al.
Published: (2024)
STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft
by: Zhao, Zhonghan, et al.
Published: (2024)
by: Zhao, Zhonghan, et al.
Published: (2024)
StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
by: Jiao, Guanlong, et al.
Published: (2026)
by: Jiao, Guanlong, et al.
Published: (2026)
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
by: Li, Danrui, et al.
Published: (2026)
by: Li, Danrui, et al.
Published: (2026)
DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding
by: Wu, Yuchuan, et al.
Published: (2026)
by: Wu, Yuchuan, et al.
Published: (2026)
Similar Items
-
Temporally Grounding Instructional Diagrams in Unconstrained Videos
by: Zhang, Jiahao, et al.
Published: (2024) -
Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
by: Zhang, Jiahao, et al.
Published: (2024) -
3DInAction: Understanding Human Actions in 3D Point Clouds
by: Ben-Shabat, Yizhak, et al.
Published: (2023) -
Neural Experts: Mixture of Experts for Implicit Neural Representations
by: Ben-Shabat, Yizhak, et al.
Published: (2024) -
VI3NR: Variance Informed Initialization for Implicit Neural Representations
by: Koneputugodage, Chamin Hewa, et al.
Published: (2025)