Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yuxiao, Li, Kai, Bao, Wentao, Patel, Deep, Kong, Yu, Min, Martin Renqiang, Metaxas, Dimitris N. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024)
by: Bao, Wentao, et al.
Published: (2024)
Group Relative Augmentation for Data Efficient Action Detection
by: Patel, Deep Anil, et al.
Published: (2025)
by: Patel, Deep Anil, et al.
Published: (2025)
DiscussLLM: Teaching Large Language Models When to Speak
by: Patel, Deep Anil, et al.
Published: (2025)
by: Patel, Deep Anil, et al.
Published: (2025)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
Object-Aware 4D Human Motion Generation
by: Gui, Shurui, et al.
Published: (2025)
by: Gui, Shurui, et al.
Published: (2025)
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
by: Xian, Ruiqi, et al.
Published: (2026)
by: Xian, Ruiqi, et al.
Published: (2026)
Pooling and Semantic Shift: The Fundamental Challenges in Long Text Embedding and Retrieval
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
by: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Published: (2024)
by: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Published: (2024)
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
HVD: Human Vision-Driven Video Representation Learning for Text-Video Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
by: Dafnis, Konstantinos M., et al.
Published: (2025)
by: Dafnis, Konstantinos M., et al.
Published: (2025)
Aligning Large Language Models with Healthcare Stakeholders: A Pathway to Trustworthy AI Integration
by: Ding, Kexin, et al.
Published: (2025)
by: Ding, Kexin, et al.
Published: (2025)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
by: Wang, Zhenting, et al.
Published: (2023)
by: Wang, Zhenting, et al.
Published: (2023)
AVID: Any-Length Video Inpainting with Diffusion Model
by: Zhang, Zhixing, et al.
Published: (2023)
by: Zhang, Zhixing, et al.
Published: (2023)
Learning to Rank Caption Chains for Video-Text Alignment
by: Blume, Ansel, et al.
Published: (2026)
by: Blume, Ansel, et al.
Published: (2026)
Learning Skills from Action-Free Videos
by: Fang, Hung-Chieh, et al.
Published: (2025)
by: Fang, Hung-Chieh, et al.
Published: (2025)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
New Capability to Look Up an ASL Sign from a Video Example
by: Neidle, Carol, et al.
Published: (2024)
by: Neidle, Carol, et al.
Published: (2024)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
by: Yang, Zhuoyi, et al.
Published: (2026)
by: Yang, Zhuoyi, et al.
Published: (2026)
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
by: Gu, Difei, et al.
Published: (2025)
by: Gu, Difei, et al.
Published: (2025)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
EditGRPO: Reinforcement Learning with Post-Rollout Edits for Clinically Accurate Chest X-Ray Report Generation
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
by: Guo, Wenliang, et al.
Published: (2025)
by: Guo, Wenliang, et al.
Published: (2025)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
by: Zhang, Zhixing, et al.
Published: (2022)
by: Zhang, Zhixing, et al.
Published: (2022)
Signal or Noise in Multi-Agent LLM-based Stock Recommendations?
by: Fatouros, George, et al.
Published: (2026)
by: Fatouros, George, et al.
Published: (2026)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)
by: Kim, Joochan, et al.
Published: (2025)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
by: Souček, Tomáš, et al.
Published: (2023)
by: Souček, Tomáš, et al.
Published: (2023)
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024)
by: Dao, Quan, et al.
Published: (2024)
Joint Self-Supervised Video Alignment and Action Segmentation
by: Ali, Ali Shah, et al.
Published: (2025)
by: Ali, Ali Shah, et al.
Published: (2025)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
by: Pu, Yujiang, et al.
Published: (2025)
by: Pu, Yujiang, et al.
Published: (2025)
Deep Video Codec Control for Vision Models
by: Reich, Christoph, et al.
Published: (2023)
by: Reich, Christoph, et al.
Published: (2023)
LAIP: Learning Local Alignment from Image-Phrase Modeling for Text-based Person Search
by: Wang, Haiguang, et al.
Published: (2024)
by: Wang, Haiguang, et al.
Published: (2024)
Learning Volumetric Neural Deformable Models to Recover 3D Regional Heart Wall Motion from Multi-Planar Tagged MRI
by: Ye, Meng, et al.
Published: (2024)
by: Ye, Meng, et al.
Published: (2024)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025)
by: Zhao, Zixu, et al.
Published: (2025)
Similar Items
-
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024) -
Group Relative Augmentation for Data Efficient Action Detection
by: Patel, Deep Anil, et al.
Published: (2025) -
DiscussLLM: Teaching Large Language Models When to Speak
by: Patel, Deep Anil, et al.
Published: (2025) -
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026) -
Object-Aware 4D Human Motion Generation
by: Gui, Shurui, et al.
Published: (2025)