Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Jinghan, Huang, Yifei, Lu, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
by: Samel, Karan, et al.
Published: (2025)
by: Samel, Karan, et al.
Published: (2025)
RECIPE: Procedural Planning via Grounding in Instructional Video
by: Seminara, Luigi, et al.
Published: (2026)
by: Seminara, Luigi, et al.
Published: (2026)
Unfolding Videos Dynamics via Taylor Expansion
by: Chen, Siyi, et al.
Published: (2024)
by: Chen, Siyi, et al.
Published: (2024)
What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
by: Kung, Chi-Hsi, et al.
Published: (2025)
by: Kung, Chi-Hsi, et al.
Published: (2025)
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
by: Li, Jinghan, et al.
Published: (2025)
by: Li, Jinghan, et al.
Published: (2025)
Frequency-aware Neural Representation for Videos
by: Zhu, Jun, et al.
Published: (2026)
by: Zhu, Jun, et al.
Published: (2026)
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
by: Dong, Lu, et al.
Published: (2025)
by: Dong, Lu, et al.
Published: (2025)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
by: Shi, Lei, et al.
Published: (2024)
by: Shi, Lei, et al.
Published: (2024)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
EgoVIS@CVPR: What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
by: Kung, Chi-Hsi, et al.
Published: (2025)
by: Kung, Chi-Hsi, et al.
Published: (2025)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
VADMamba: Exploring State Space Models for Fast Video Anomaly Detection
by: Lyu, Jiahao, et al.
Published: (2025)
by: Lyu, Jiahao, et al.
Published: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)
by: Seminara, Luigi, et al.
Published: (2024)
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Latent Radiance Fields with 3D-aware 2D Representations
by: Zhou, Chaoyi, et al.
Published: (2025)
by: Zhou, Chaoyi, et al.
Published: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Uncertainty-aware Prototype Learning with Variational Inference for Few-shot Point Cloud Segmentation
by: Zhao, Yifei, et al.
Published: (2026)
by: Zhao, Yifei, et al.
Published: (2026)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
by: Schoonbeek, Tim J., et al.
Published: (2025)
by: Schoonbeek, Tim J., et al.
Published: (2025)
Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
Pose-Specific 3D Fingerprint Unfolding
by: Guan, Xiongjun, et al.
Published: (2024)
by: Guan, Xiongjun, et al.
Published: (2024)
MVGD-Net: A Novel Motion-aware Video Glass Surface Detection Network
by: Lu, Yiwei, et al.
Published: (2026)
by: Lu, Yiwei, et al.
Published: (2026)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
by: Ilaslan, Muhammet Furkan, et al.
Published: (2024)
by: Ilaslan, Muhammet Furkan, et al.
Published: (2024)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
by: Dong, Yuhao, et al.
Published: (2026)
by: Dong, Yuhao, et al.
Published: (2026)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
by: Yu, Fangxu, et al.
Published: (2026)
by: Yu, Fangxu, et al.
Published: (2026)
Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference
by: Ahmed, Sk Miraj, et al.
Published: (2026)
by: Ahmed, Sk Miraj, et al.
Published: (2026)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
RED: Robust Environmental Design
by: Yang, Jinghan
Published: (2024)
by: Yang, Jinghan
Published: (2024)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
Clapper: Compact Learning and Video Representation in VLMs
by: Kong, Lingyu, et al.
Published: (2025)
by: Kong, Lingyu, et al.
Published: (2025)
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
by: Liu, Pinxin, et al.
Published: (2025)
by: Liu, Pinxin, et al.
Published: (2025)
LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders
by: Naiman, Ilan, et al.
Published: (2025)
by: Naiman, Ilan, et al.
Published: (2025)
Similar Items
-
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
by: Samel, Karan, et al.
Published: (2025) -
RECIPE: Procedural Planning via Grounding in Instructional Video
by: Seminara, Luigi, et al.
Published: (2026) -
Unfolding Videos Dynamics via Taylor Expansion
by: Chen, Siyi, et al.
Published: (2024) -
What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
by: Kung, Chi-Hsi, et al.
Published: (2025) -
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
by: Li, Jinghan, et al.
Published: (2025)