Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Jinghan, Huang, Yifei, Lu, Feng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
von: Samel, Karan, et al.
Veröffentlicht: (2025)
von: Samel, Karan, et al.
Veröffentlicht: (2025)
RECIPE: Procedural Planning via Grounding in Instructional Video
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
von: Seminara, Luigi, et al.
Veröffentlicht: (2026)
Unfolding Videos Dynamics via Taylor Expansion
von: Chen, Siyi, et al.
Veröffentlicht: (2024)
von: Chen, Siyi, et al.
Veröffentlicht: (2024)
What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
von: Kung, Chi-Hsi, et al.
Veröffentlicht: (2025)
von: Kung, Chi-Hsi, et al.
Veröffentlicht: (2025)
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
von: Li, Jinghan, et al.
Veröffentlicht: (2025)
von: Li, Jinghan, et al.
Veröffentlicht: (2025)
Frequency-aware Neural Representation for Videos
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
Learning Streaming Video Representation via Multitask Training
von: Yan, Yibin, et al.
Veröffentlicht: (2025)
von: Yan, Yibin, et al.
Veröffentlicht: (2025)
VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
von: Dong, Lu, et al.
Veröffentlicht: (2025)
von: Dong, Lu, et al.
Veröffentlicht: (2025)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
von: Shi, Lei, et al.
Veröffentlicht: (2024)
von: Shi, Lei, et al.
Veröffentlicht: (2024)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
von: Lin, Han, et al.
Veröffentlicht: (2024)
von: Lin, Han, et al.
Veröffentlicht: (2024)
EgoVIS@CVPR: What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
von: Kung, Chi-Hsi, et al.
Veröffentlicht: (2025)
von: Kung, Chi-Hsi, et al.
Veröffentlicht: (2025)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
von: Dai, Ming, et al.
Veröffentlicht: (2025)
von: Dai, Ming, et al.
Veröffentlicht: (2025)
Anticipating Object State Changes in Long Procedural Videos
von: Manousaki, Victoria, et al.
Veröffentlicht: (2024)
von: Manousaki, Victoria, et al.
Veröffentlicht: (2024)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
VADMamba: Exploring State Space Models for Fast Video Anomaly Detection
von: Lyu, Jiahao, et al.
Veröffentlicht: (2025)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2025)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
von: Seminara, Luigi, et al.
Veröffentlicht: (2024)
von: Seminara, Luigi, et al.
Veröffentlicht: (2024)
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
von: Li, Yifei, et al.
Veröffentlicht: (2025)
von: Li, Yifei, et al.
Veröffentlicht: (2025)
Latent Radiance Fields with 3D-aware 2D Representations
von: Zhou, Chaoyi, et al.
Veröffentlicht: (2025)
von: Zhou, Chaoyi, et al.
Veröffentlicht: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
Uncertainty-aware Prototype Learning with Variational Inference for Few-shot Point Cloud Segmentation
von: Zhao, Yifei, et al.
Veröffentlicht: (2026)
von: Zhao, Yifei, et al.
Veröffentlicht: (2026)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
von: Lu, Feng, et al.
Veröffentlicht: (2024)
von: Lu, Feng, et al.
Veröffentlicht: (2024)
Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
von: Schoonbeek, Tim J., et al.
Veröffentlicht: (2025)
von: Schoonbeek, Tim J., et al.
Veröffentlicht: (2025)
Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model
von: Shi, Yuheng, et al.
Veröffentlicht: (2024)
von: Shi, Yuheng, et al.
Veröffentlicht: (2024)
Pose-Specific 3D Fingerprint Unfolding
von: Guan, Xiongjun, et al.
Veröffentlicht: (2024)
von: Guan, Xiongjun, et al.
Veröffentlicht: (2024)
MVGD-Net: A Novel Motion-aware Video Glass Surface Detection Network
von: Lu, Yiwei, et al.
Veröffentlicht: (2026)
von: Lu, Yiwei, et al.
Veröffentlicht: (2026)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
Multi-sentence Video Grounding for Long Video Generation
von: Feng, Wei, et al.
Veröffentlicht: (2024)
von: Feng, Wei, et al.
Veröffentlicht: (2024)
Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
von: Yu, Fangxu, et al.
Veröffentlicht: (2026)
von: Yu, Fangxu, et al.
Veröffentlicht: (2026)
Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference
von: Ahmed, Sk Miraj, et al.
Veröffentlicht: (2026)
von: Ahmed, Sk Miraj, et al.
Veröffentlicht: (2026)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
RED: Robust Environmental Design
von: Yang, Jinghan
Veröffentlicht: (2024)
von: Yang, Jinghan
Veröffentlicht: (2024)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
Clapper: Compact Learning and Video Representation in VLMs
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
von: Liu, Pinxin, et al.
Veröffentlicht: (2025)
von: Liu, Pinxin, et al.
Veröffentlicht: (2025)
LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders
von: Naiman, Ilan, et al.
Veröffentlicht: (2025)
von: Naiman, Ilan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
von: Samel, Karan, et al.
Veröffentlicht: (2025) -
RECIPE: Procedural Planning via Grounding in Instructional Video
von: Seminara, Luigi, et al.
Veröffentlicht: (2026) -
Unfolding Videos Dynamics via Taylor Expansion
von: Chen, Siyi, et al.
Veröffentlicht: (2024) -
What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
von: Kung, Chi-Hsi, et al.
Veröffentlicht: (2025) -
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
von: Li, Jinghan, et al.
Veröffentlicht: (2025)