CLIP-MG: Guiding Semantic Attention with Skeletal Pose Features and RGB Data for Micro-Gesture Recognition on the iMiGUE Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Patapati, Santosh, Srinivasan, Trisanth, Adiraju, Amith |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition
by: Martirosyan, Arman, et al.
Published: (2025)
by: Martirosyan, Arman, et al.
Published: (2025)
PhysNav-DG: A Novel Adaptive Framework for Robust VLM-Sensor Fusion in Navigation Applications
by: Srinivasan, Trisanth, et al.
Published: (2025)
by: Srinivasan, Trisanth, et al.
Published: (2025)
Goal-Based Vision-Language Driving
by: Patapati, Santosh, et al.
Published: (2025)
by: Patapati, Santosh, et al.
Published: (2025)
iMiGUE-3K: A Large-Scale Benchmark for Micro-Gesture Analysis with Self-Supervised Learning
by: Wang, Chengyan, et al.
Published: (2026)
by: Wang, Chengyan, et al.
Published: (2026)
Vision-Language Cross-Attention for Real-Time Autonomous Driving
by: Patapati, Santosh, et al.
Published: (2025)
by: Patapati, Santosh, et al.
Published: (2025)
WebNav: An Intelligent Agent for Voice-Controlled Web Navigation
by: Srinivasan, Trisanth, et al.
Published: (2025)
by: Srinivasan, Trisanth, et al.
Published: (2025)
Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
by: Srinivasan, Trisanth, et al.
Published: (2025)
by: Srinivasan, Trisanth, et al.
Published: (2025)
iMiGUE-Speech: A Spontaneous Speech Dataset for Affective Analysis
by: Kakouros, Sofoklis, et al.
Published: (2026)
by: Kakouros, Sofoklis, et al.
Published: (2026)
CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities
by: Patapati, Santosh
Published: (2025)
by: Patapati, Santosh
Published: (2025)
Integrating Large Language Models into a Tri-Modal Architecture for Automated Depression Classification on the DAIC-WOZ
by: Patapati, Santosh V.
Published: (2024)
by: Patapati, Santosh V.
Published: (2024)
Graph Coloring for Multi-Task Learning
by: Patapati, Santosh
Published: (2025)
by: Patapati, Santosh
Published: (2025)
Data-Free Class-Incremental Gesture Recognition with Prototype-Guided Pseudo Feature Replay
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
Active Inference for Micro-Gesture Recognition: EFE-Guided Temporal Sampling and Adaptive Learning
by: Feng, Weijia, et al.
Published: (2026)
by: Feng, Weijia, et al.
Published: (2026)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
FG-SGL: Fine-Grained Semantic Guidance Learning via Motion Process Decomposition for Micro-Gesture Recognition
by: Wei, Jinsheng, et al.
Published: (2026)
by: Wei, Jinsheng, et al.
Published: (2026)
Gesture Matters: Pedestrian Gesture Recognition for AVs Through Skeleton Pose Evaluation
by: Mahdi, Alif Rizqullah, et al.
Published: (2026)
by: Mahdi, Alif Rizqullah, et al.
Published: (2026)
HaGRID - HAnd Gesture Recognition Image Dataset
by: Kapitanov, Alexander, et al.
Published: (2022)
by: Kapitanov, Alexander, et al.
Published: (2022)
SASG-DA: Sparse-Aware Semantic-Guided Diffusion Augmentation For Myoelectric Gesture Recognition
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Resource-Efficient Gesture Recognition through Convexified Attention
by: Schwartz, Daniel, et al.
Published: (2026)
by: Schwartz, Daniel, et al.
Published: (2026)
DURA-CPS: A Multi-Role Orchestrator for Dependability Assurance in LLM-Enabled Cyber-Physical Systems
by: Srinivasan, Trisanth, et al.
Published: (2025)
by: Srinivasan, Trisanth, et al.
Published: (2025)
Towards Fine-Grained Emotion Understanding via Skeleton-Based Micro-Gesture Recognition
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
by: Li, Deng, et al.
Published: (2025)
by: Li, Deng, et al.
Published: (2025)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Frequency-Guided Fusion For RGB-Thermal Semantic Segmentation
by: Canıtez, İsmail Emre, et al.
Published: (2026)
by: Canıtez, İsmail Emre, et al.
Published: (2026)
ROBUST-MIPS: A Combined Skeletal Pose and Instance Segmentation Dataset for Laparoscopic Surgical Instruments
by: Han, Zhe, et al.
Published: (2025)
by: Han, Zhe, et al.
Published: (2025)
Online Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention
by: Liu, Pengyu, et al.
Published: (2025)
by: Liu, Pengyu, et al.
Published: (2025)
Boosting Gesture Recognition with an Automatic Gesture Annotation Framework
by: Shen, Junxiao, et al.
Published: (2024)
by: Shen, Junxiao, et al.
Published: (2024)
Hololens 2 Hand Joints Dataset for Hand Gesture Recognition
by: Pistola, Theodora
Published: (2025)
by: Pistola, Theodora
Published: (2025)
EgoEV-HandPose: Egocentric 3D Hand Pose Estimation and Gesture Recognition with Stereo Event Cameras
by: Wang, Luming, et al.
Published: (2026)
by: Wang, Luming, et al.
Published: (2026)
GSGTrack: Gaussian Splatting-Guided Object Pose Tracking from RGB Videos
by: Chen, Zhiyuan, et al.
Published: (2024)
by: Chen, Zhiyuan, et al.
Published: (2024)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
by: Zhou, Qiongyi, et al.
Published: (2024)
by: Zhou, Qiongyi, et al.
Published: (2024)
CLIP-Guided Unsupervised Semantic-Aware Exposure Correction
by: Wu, Puzhen, et al.
Published: (2026)
by: Wu, Puzhen, et al.
Published: (2026)
Multi-scale Attention Guided Pose Transfer
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
Fusing Monocular RGB Images with AIS Data to Create a 6D Pose Estimation Dataset for Marine Vessels
by: Holst, Fabian, et al.
Published: (2025)
by: Holst, Fabian, et al.
Published: (2025)
Enhancing Micro Gesture Recognition for Emotion Understanding via Context-aware Visual-Text Contrastive Learning
by: Li, Deng, et al.
Published: (2024)
by: Li, Deng, et al.
Published: (2024)
Learning Flow-Guided Registration for RGB-Event Semantic Segmentation
by: Yao, Zhen, et al.
Published: (2025)
by: Yao, Zhen, et al.
Published: (2025)
WiFi-based Cross-Domain Gesture Recognition Using Attention Mechanism
by: Liu, Ruijing, et al.
Published: (2025)
by: Liu, Ruijing, et al.
Published: (2025)
Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer
by: Bui, Minh, et al.
Published: (2024)
by: Bui, Minh, et al.
Published: (2024)
Hierarchical Space-Time Attention for Micro-Expression Recognition
by: Hao, Haihong, et al.
Published: (2024)
by: Hao, Haihong, et al.
Published: (2024)
Multi-Modal Gesture Recognition from Video and Surgical Tool Pose Information via Motion Invariants
by: Atoum, Jumanh, et al.
Published: (2025)
by: Atoum, Jumanh, et al.
Published: (2025)
Similar Items
-
Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition
by: Martirosyan, Arman, et al.
Published: (2025) -
PhysNav-DG: A Novel Adaptive Framework for Robust VLM-Sensor Fusion in Navigation Applications
by: Srinivasan, Trisanth, et al.
Published: (2025) -
Goal-Based Vision-Language Driving
by: Patapati, Santosh, et al.
Published: (2025) -
iMiGUE-3K: A Large-Scale Benchmark for Micro-Gesture Analysis with Self-Supervised Learning
by: Wang, Chengyan, et al.
Published: (2026) -
Vision-Language Cross-Attention for Real-Time Autonomous Driving
by: Patapati, Santosh, et al.
Published: (2025)