SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sinha, Arkaprava, Reilly, Dominick, Bremond, Francois, Wang, Pu, Das, Srijan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
Introducing Gating and Context into Temporal Action Detection
von: Reka, Aglind, et al.
Veröffentlicht: (2024)
von: Reka, Aglind, et al.
Veröffentlicht: (2024)
DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
von: Bondurant, Weston, et al.
Veröffentlicht: (2025)
von: Bondurant, Weston, et al.
Veröffentlicht: (2025)
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
von: Rahimian, Ali K., et al.
Veröffentlicht: (2024)
von: Rahimian, Ali K., et al.
Veröffentlicht: (2024)
Are Visual-Language Models Effective in Action Recognition? A Comparative Study
von: Ali, Mahmoud, et al.
Veröffentlicht: (2024)
von: Ali, Mahmoud, et al.
Veröffentlicht: (2024)
LAC: Latent Action Composition for Skeleton-based Action Segmentation
von: Yang, Di, et al.
Veröffentlicht: (2023)
von: Yang, Di, et al.
Veröffentlicht: (2023)
Temporally Propagated Masks and Bounding Boxes: Combining the Best of Both Worlds for Multi-Object Tracking
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2024)
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2024)
Identifying Surgical Instruments in Pedagogical Cataract Surgery Videos through an Optimized Aggregation Network
von: Sinha, Sanya, et al.
Veröffentlicht: (2025)
von: Sinha, Sanya, et al.
Veröffentlicht: (2025)
Affinity Contrastive Learning for Skeleton-based Human Activity Understanding
von: Liu, Hongda, et al.
Veröffentlicht: (2026)
von: Liu, Hongda, et al.
Veröffentlicht: (2026)
Sigma: Semantically Informative Pre-training for Skeleton-based Sign Language Understanding
von: Pu, Muxin, et al.
Veröffentlicht: (2025)
von: Pu, Muxin, et al.
Veröffentlicht: (2025)
No Train Yet Gain: Towards Generic Multi-Object Tracking in Sports and Beyond
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2025)
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2025)
BAMM: Bidirectional Autoregressive Motion Model
von: Pinyoanuntapong, Ekkasit, et al.
Veröffentlicht: (2024)
von: Pinyoanuntapong, Ekkasit, et al.
Veröffentlicht: (2024)
Foundation Model for Skeleton-Based Human Action Understanding
von: Wang, Hongsong, et al.
Veröffentlicht: (2025)
von: Wang, Hongsong, et al.
Veröffentlicht: (2025)
DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity Understanding
von: Kreutz, Thomas, et al.
Veröffentlicht: (2025)
von: Kreutz, Thomas, et al.
Veröffentlicht: (2025)
Understanding the Vulnerability of Skeleton-based Human Activity Recognition via Black-box Attack
von: Diao, Yunfeng, et al.
Veröffentlicht: (2022)
von: Diao, Yunfeng, et al.
Veröffentlicht: (2022)
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
von: Reddy, Shukesh, et al.
Veröffentlicht: (2026)
von: Reddy, Shukesh, et al.
Veröffentlicht: (2026)
Self-supervised Auxiliary Learning for Texture and Model-based Hybrid Robust and Fair Featuring in Face Analysis
von: Reddy, Shukesh, et al.
Veröffentlicht: (2024)
von: Reddy, Shukesh, et al.
Veröffentlicht: (2024)
Towards Understanding Best Practices for Quantization of Vision-Language Models
von: Das, Gautom, et al.
Veröffentlicht: (2026)
von: Das, Gautom, et al.
Veröffentlicht: (2026)
Siformer: Feature-isolated Transformer for Efficient Skeleton-based Sign Language Recognition
von: Pu, Muxin, et al.
Veröffentlicht: (2025)
von: Pu, Muxin, et al.
Veröffentlicht: (2025)
What Matters in Autonomous Driving Anomaly Detection: A Weakly Supervised Horizon
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2024)
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2024)
AM Flow: Adapters for Temporal Processing in Action Recognition
von: Agrawal, Tanay, et al.
Veröffentlicht: (2024)
von: Agrawal, Tanay, et al.
Veröffentlicht: (2024)
Understanding Degradation with Vision Language Model
von: Lan, Guanzhou, et al.
Veröffentlicht: (2026)
von: Lan, Guanzhou, et al.
Veröffentlicht: (2026)
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
KHMP: Frequency-Domain Kalman Refinement for High-Fidelity Human Motion Prediction
von: Wu, Wenhan, et al.
Veröffentlicht: (2026)
von: Wu, Wenhan, et al.
Veröffentlicht: (2026)
SGA-INTERACT: A 3D Skeleton-based Benchmark for Group Activity Understanding in Modern Basketball Tactic
von: Yang, Yuchen, et al.
Veröffentlicht: (2025)
von: Yang, Yuchen, et al.
Veröffentlicht: (2025)
Technical Report: Masked Skeleton Sequence Modeling for Learning Larval Zebrafish Behavior Latent Embeddings
von: Xu, Lanxin, et al.
Veröffentlicht: (2024)
von: Xu, Lanxin, et al.
Veröffentlicht: (2024)
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models
von: Yang, Siyuan, et al.
Veröffentlicht: (2026)
von: Yang, Siyuan, et al.
Veröffentlicht: (2026)
Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
von: Wang, Xinshun, et al.
Veröffentlicht: (2023)
von: Wang, Xinshun, et al.
Veröffentlicht: (2023)
LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition
von: Pu, Muxin, et al.
Veröffentlicht: (2026)
von: Pu, Muxin, et al.
Veröffentlicht: (2026)
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
von: Poddar, Nishit, et al.
Veröffentlicht: (2026)
von: Poddar, Nishit, et al.
Veröffentlicht: (2026)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
DensePercept-NCSSD: Vision Mamba towards Real-time Dense Visual Perception with Non-Causal State Space Duality
von: Anand, Tushar, et al.
Veröffentlicht: (2025)
von: Anand, Tushar, et al.
Veröffentlicht: (2025)
EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision Language Models
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
von: Wang, Xinshun, et al.
Veröffentlicht: (2026)
von: Wang, Xinshun, et al.
Veröffentlicht: (2026)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2024) -
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2025) -
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026) -
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
von: Reilly, Dominick, et al.
Veröffentlicht: (2025) -
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)