Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rajendiran, Ramanathan, Roy, Debaditya, Fernando, Basura |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interaction Region Visual Transformer for Egocentric Action Anticipation
von: Roy, Debaditya, et al.
Veröffentlicht: (2022)
von: Roy, Debaditya, et al.
Veröffentlicht: (2022)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
von: Rajendiran, Ramanathan, et al.
Veröffentlicht: (2023)
von: Rajendiran, Ramanathan, et al.
Veröffentlicht: (2023)
Predicting the Next Action by Modeling the Abstract Goal
von: Roy, Debaditya, et al.
Veröffentlicht: (2022)
von: Roy, Debaditya, et al.
Veröffentlicht: (2022)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
von: Verma, Dhruv, et al.
Veröffentlicht: (2024)
von: Verma, Dhruv, et al.
Veröffentlicht: (2024)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
von: Jaiswal, Shantanu, et al.
Veröffentlicht: (2024)
von: Jaiswal, Shantanu, et al.
Veröffentlicht: (2024)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
von: Ee, Yeo Keat, et al.
Veröffentlicht: (2026)
von: Ee, Yeo Keat, et al.
Veröffentlicht: (2026)
Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning
von: Bangde, Yashwant Pravinrao, et al.
Veröffentlicht: (2026)
von: Bangde, Yashwant Pravinrao, et al.
Veröffentlicht: (2026)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
von: Li, Chen, et al.
Veröffentlicht: (2025)
von: Li, Chen, et al.
Veröffentlicht: (2025)
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
Situational Scene Graph for Structured Human-centric Situation Understanding
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2024)
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2024)
Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework
von: Yang, Zhengwei, et al.
Veröffentlicht: (2024)
von: Yang, Zhengwei, et al.
Veröffentlicht: (2024)
GLaRE: A Graph-based Landmark Region Embedding Network for Emotion Recognition
von: Maji, Debasis, et al.
Veröffentlicht: (2025)
von: Maji, Debasis, et al.
Veröffentlicht: (2025)
Learning to Visually Connect Actions and their Effects
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
von: Keat, Ee Yeo, et al.
Veröffentlicht: (2024)
von: Keat, Ee Yeo, et al.
Veröffentlicht: (2024)
Inferring Past Human Actions in Homes with Abductive Reasoning
von: Tan, Clement, et al.
Veröffentlicht: (2022)
von: Tan, Clement, et al.
Veröffentlicht: (2022)
Generating Key Postures of Bharatanatyam Adavus with Pose Estimation
von: Kamble, Jagadish Kashinath, et al.
Veröffentlicht: (2026)
von: Kamble, Jagadish Kashinath, et al.
Veröffentlicht: (2026)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
von: Stracke, Nick, et al.
Veröffentlicht: (2026)
von: Stracke, Nick, et al.
Veröffentlicht: (2026)
Context-Aware Pesticide Recommendation via Few-Shot Pest Recognition for Precision Agriculture
von: Ghosh, Anirudha, et al.
Veröffentlicht: (2026)
von: Ghosh, Anirudha, et al.
Veröffentlicht: (2026)
LongLive: Real-time Interactive Long Video Generation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
VideoAuteur: Towards Long Narrative Video Generation
von: Xiao, Junfei, et al.
Veröffentlicht: (2025)
von: Xiao, Junfei, et al.
Veröffentlicht: (2025)
DescribeEarth: Describe Anything for Remote Sensing Images
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
Diagram-Driven Course Questions Generation
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2024)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
PovNet+: A Deep Learning Architecture for Socially Assistive Robots to Learn and Assist with Multiple Activities of Daily Living
von: Robinson, Fraser, et al.
Veröffentlicht: (2026)
von: Robinson, Fraser, et al.
Veröffentlicht: (2026)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Lagrangian Motion Fields for Long-term Motion Generation
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
Recognition of Daily Activities through Multi-Modal Deep Learning: A Video, Pose, and Object-Aware Approach for Ambient Assisted Living
von: Hashemifard, Kooshan, et al.
Veröffentlicht: (2026)
von: Hashemifard, Kooshan, et al.
Veröffentlicht: (2026)
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
ADL4D: Towards A Contextually Rich Dataset for 4D Activities of Daily Living
von: Zakour, Marsil, et al.
Veröffentlicht: (2024)
von: Zakour, Marsil, et al.
Veröffentlicht: (2024)
Describe Anything in Medical Images
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
An Attention Infused Deep Learning System with Grad-CAM Visualization for Early Screening of Glaucoma
von: Swaminathan, Ramanathan
Veröffentlicht: (2025)
von: Swaminathan, Ramanathan
Veröffentlicht: (2025)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
von: Parmar, Paritosh, et al.
Veröffentlicht: (2025)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2025)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
von: Durante, Zane, et al.
Veröffentlicht: (2024)
von: Durante, Zane, et al.
Veröffentlicht: (2024)
Controllable Long-term Motion Generation with Extended Joint Targets
von: Lee, Eunjong, et al.
Veröffentlicht: (2025)
von: Lee, Eunjong, et al.
Veröffentlicht: (2025)
NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation
von: Feng, X., et al.
Veröffentlicht: (2025)
von: Feng, X., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interaction Region Visual Transformer for Egocentric Action Anticipation
von: Roy, Debaditya, et al.
Veröffentlicht: (2022) -
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
von: Rajendiran, Ramanathan, et al.
Veröffentlicht: (2023) -
Predicting the Next Action by Modeling the Abstract Goal
von: Roy, Debaditya, et al.
Veröffentlicht: (2022) -
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
von: Verma, Dhruv, et al.
Veröffentlicht: (2024) -
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
von: Jaiswal, Shantanu, et al.
Veröffentlicht: (2024)