CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Lei, Bulling, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
by: Shi, Lei, et al.
Published: (2024)
by: Shi, Lei, et al.
Published: (2024)
Multi-Modal Video Dialog State Tracking in the Wild
by: Abdessaied, Adnen, et al.
Published: (2024)
by: Abdessaied, Adnen, et al.
Published: (2024)
LAP: A Language-Aware Planning Model For Procedure Planning In Instructional Videos
by: Shi, Lei, et al.
Published: (2026)
by: Shi, Lei, et al.
Published: (2026)
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images
by: Penzkofer, Anna, et al.
Published: (2024)
by: Penzkofer, Anna, et al.
Published: (2024)
GazeMoDiff: Gaze-guided Diffusion Model for Stochastic Human Motion Prediction
by: Yan, Haodong, et al.
Published: (2023)
by: Yan, Haodong, et al.
Published: (2023)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
UP-FacE: User-predictable Fine-grained Face Shape Editing
by: Strohm, Florian, et al.
Published: (2024)
by: Strohm, Florian, et al.
Published: (2024)
HAIFAI: Human-AI Interaction for Mental Face Reconstruction
by: Strohm, Florian, et al.
Published: (2024)
by: Strohm, Florian, et al.
Published: (2024)
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
by: Guruprasad, Pranav, et al.
Published: (2025)
by: Guruprasad, Pranav, et al.
Published: (2025)
Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models
by: Cai, Zhejia, et al.
Published: (2025)
by: Cai, Zhejia, et al.
Published: (2025)
PDPP: Projected Diffusion for Procedure Planning in Instructional Videos
by: Wang, Hanlin, et al.
Published: (2023)
by: Wang, Hanlin, et al.
Published: (2023)
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026)
by: Wang, Fengshun, et al.
Published: (2026)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026)
by: Govind, Manish Kumar, et al.
Published: (2026)
OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog
by: Abdessaied, Adnen, et al.
Published: (2024)
by: Abdessaied, Adnen, et al.
Published: (2024)
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
by: Zhou, Yufan, et al.
Published: (2025)
by: Zhou, Yufan, et al.
Published: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
by: Sun, Jingwen, et al.
Published: (2026)
by: Sun, Jingwen, et al.
Published: (2026)
VLMaterial: Procedural Material Generation with Large Vision-Language Models
by: Li, Beichen, et al.
Published: (2025)
by: Li, Beichen, et al.
Published: (2025)
HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head Coordination
by: Hu, Zhiming, et al.
Published: (2025)
by: Hu, Zhiming, et al.
Published: (2025)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
by: Abdessaied, Adnen, et al.
Published: (2025)
by: Abdessaied, Adnen, et al.
Published: (2025)
Pose2Gaze: Eye-body Coordination during Daily Activities for Gaze Prediction from Full-body Poses
by: Hu, Zhiming, et al.
Published: (2023)
by: Hu, Zhiming, et al.
Published: (2023)
GazeMotion: Gaze-guided Human Motion Forecasting
by: Hu, Zhiming, et al.
Published: (2024)
by: Hu, Zhiming, et al.
Published: (2024)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
by: Tang, Zuojin, et al.
Published: (2026)
by: Tang, Zuojin, et al.
Published: (2026)
SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
by: Liu, Haowen, et al.
Published: (2025)
by: Liu, Haowen, et al.
Published: (2025)
LLaDA-VLA: Vision Language Diffusion Action Models
by: Wen, Yuqing, et al.
Published: (2025)
by: Wen, Yuqing, et al.
Published: (2025)
HAGI++: Head-Assisted Gaze Imputation and Generation
by: Jiao, Chuhan, et al.
Published: (2025)
by: Jiao, Chuhan, et al.
Published: (2025)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
by: Luo, Yuechen, et al.
Published: (2026)
by: Luo, Yuechen, et al.
Published: (2026)
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
by: Ghoddoosian, Reza, et al.
Published: (2024)
by: Ghoddoosian, Reza, et al.
Published: (2024)
PRET: Planning with Directed Fidelity Trajectory for Vision and Language Navigation
by: Lu, Renjie, et al.
Published: (2024)
by: Lu, Renjie, et al.
Published: (2024)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
by: Ye, Jiacheng, et al.
Published: (2025)
by: Ye, Jiacheng, et al.
Published: (2025)
Ontology-Guided Diffusion for Zero-Shot Visual Sim2Real Transfer
by: Youssef, Mohamed, et al.
Published: (2026)
by: Youssef, Mohamed, et al.
Published: (2026)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2026)
by: Huang, Chi-Pin, et al.
Published: (2026)
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
by: Si, Shengyu, et al.
Published: (2026)
by: Si, Shengyu, et al.
Published: (2026)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
by: Nie, Dujun, et al.
Published: (2026)
by: Nie, Dujun, et al.
Published: (2026)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
by: Han, Chaolei, et al.
Published: (2025)
by: Han, Chaolei, et al.
Published: (2025)
LatentCRF: Continuous CRF for Efficient Latent Diffusion
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
DiffGaze: A Diffusion Model for Continuous Gaze Sequence Generation on 360° Images
by: Jiao, Chuhan, et al.
Published: (2024)
by: Jiao, Chuhan, et al.
Published: (2024)
Similar Items
-
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
by: Shi, Lei, et al.
Published: (2024) -
Multi-Modal Video Dialog State Tracking in the Wild
by: Abdessaied, Adnen, et al.
Published: (2024) -
LAP: A Language-Aware Planning Model For Procedure Planning In Instructional Videos
by: Shi, Lei, et al.
Published: (2026) -
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images
by: Penzkofer, Anna, et al.
Published: (2024) -
GazeMoDiff: Gaze-guided Diffusion Model for Stochastic Human Motion Prediction
by: Yan, Haodong, et al.
Published: (2023)