Two-Stream Spatial-Temporal Transformer Framework for Person Identification via Natural Conversational Keypoints
Fuente:
arXiv
Saved in:
| Main Authors: | Chapariniya, Masoumeh, Ranjbar, Hossein, Vukovic, Teodora, Ebling, Sarah, Dellwo, Volker |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics
by: Chapariniya, Masoumeh, et al.
Published: (2025)
by: Chapariniya, Masoumeh, et al.
Published: (2025)
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Foundation Model Embeddings Meet Blended Emotions: A Multimodal Fusion Approach for the BLEMORE Challenge
by: Chapariniya, Masoumeh, et al.
Published: (2026)
by: Chapariniya, Masoumeh, et al.
Published: (2026)
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
by: Farhadipour, Aref, et al.
Published: (2024)
by: Farhadipour, Aref, et al.
Published: (2024)
Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
by: Chapariniya, Masoumeh, et al.
Published: (2026)
by: Chapariniya, Masoumeh, et al.
Published: (2026)
Investigating Identity Signals in Conversational Facial Dynamics via Disentangled Expression Features
by: Chapariniya, Masoumeh, et al.
Published: (2025)
by: Chapariniya, Masoumeh, et al.
Published: (2025)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
by: Neururer, Daniel, et al.
Published: (2023)
by: Neururer, Daniel, et al.
Published: (2023)
Keypoint Promptable Re-Identification
by: Somers, Vladimir, et al.
Published: (2024)
by: Somers, Vladimir, et al.
Published: (2024)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Continuous Sign Language Recognition Using Intra-inter Gloss Attention
by: Ranjbar, Hossein, et al.
Published: (2024)
by: Ranjbar, Hossein, et al.
Published: (2024)
LatentKeypointGAN: Controlling Images via Latent Keypoints
by: He, Xingzhe, et al.
Published: (2021)
by: He, Xingzhe, et al.
Published: (2021)
An Open-World, Diverse, Cross-Spatial-Temporal Benchmark for Dynamic Wild Person Re-Identification
by: Zhang, Lei, et al.
Published: (2024)
by: Zhang, Lei, et al.
Published: (2024)
Multi-Stream Keypoint Attention Network for Sign Language Recognition and Translation
by: Guan, Mo, et al.
Published: (2024)
by: Guan, Mo, et al.
Published: (2024)
Expressive Keypoints for Skeleton-based Action Recognition via Skeleton Transformation
by: Yang, Yijie, et al.
Published: (2024)
by: Yang, Yijie, et al.
Published: (2024)
Design and Identification of Keypoint Patches in Unstructured Environments
by: Park, Taewook, et al.
Published: (2024)
by: Park, Taewook, et al.
Published: (2024)
Fractional Correspondence Framework in Detection Transformer
by: Zareapoor, Masoumeh, et al.
Published: (2025)
by: Zareapoor, Masoumeh, et al.
Published: (2025)
Skeleton-Guided Spatial-Temporal Feature Learning for Video-Based Visible-Infrared Person Re-Identification
by: Jiang, Wenjia, et al.
Published: (2024)
by: Jiang, Wenjia, et al.
Published: (2024)
Categorical Keypoint Positional Embedding for Robust Animal Re-Identification
by: Lin, Yuhao, et al.
Published: (2024)
by: Lin, Yuhao, et al.
Published: (2024)
Good Keypoints for the Two-View Geometry Estimation Problem
by: Pakulev, Konstantin, et al.
Published: (2025)
by: Pakulev, Konstantin, et al.
Published: (2025)
ISLR101: an Iranian Word-Level Sign Language Recognition Dataset
by: Ranjbar, Hossein, et al.
Published: (2025)
by: Ranjbar, Hossein, et al.
Published: (2025)
KeyRe-ID: Keypoint-Guided Person Re-Identification using Part-Aware Representation in Videos
by: Kim, Jinseong, et al.
Published: (2025)
by: Kim, Jinseong, et al.
Published: (2025)
StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video
by: Ke, Zhihui, et al.
Published: (2025)
by: Ke, Zhihui, et al.
Published: (2025)
FLAMe: Federated Learning with Attention Mechanism using Spatio-Temporal Keypoint Transformers for Pedestrian Fall Detection in Smart Cities
by: Kim, Byeonghun, et al.
Published: (2024)
by: Kim, Byeonghun, et al.
Published: (2024)
SRPose: Two-view Relative Pose Estimation with Sparse Keypoints
by: Yin, Rui, et al.
Published: (2024)
by: Yin, Rui, et al.
Published: (2024)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
Learning Structure-Supporting Dependencies via Keypoint Interactive Transformer for General Mammal Pose Estimation
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
AAformer: Auto-Aligned Transformer for Person Re-Identification
by: Zhu, Kuan, et al.
Published: (2021)
by: Zhu, Kuan, et al.
Published: (2021)
Automatic Temporal Segmentation for Post-Stroke Rehabilitation: A Keypoint Detection and Temporal Segmentation Approach for Small Datasets
by: Lee, Jisoo, et al.
Published: (2025)
by: Lee, Jisoo, et al.
Published: (2025)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
by: Liao, Ruotong, et al.
Published: (2024)
by: Liao, Ruotong, et al.
Published: (2024)
VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
by: Liu, Pengyiang, et al.
Published: (2026)
by: Liu, Pengyiang, et al.
Published: (2026)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
by: Hu, Bin, et al.
Published: (2024)
by: Hu, Bin, et al.
Published: (2024)
Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal Interactions
by: Zheng, Yuanhong, et al.
Published: (2025)
by: Zheng, Yuanhong, et al.
Published: (2025)
TSDW: A Tri-Stream Dynamic Weight Network for Cloth-Changing Person Re-Identification
by: He, Ruiqi, et al.
Published: (2025)
by: He, Ruiqi, et al.
Published: (2025)
Test-Time Adaptation via Cache Personalization for Facial Expression Recognition in Videos
by: Sharafi, Masoumeh, et al.
Published: (2026)
by: Sharafi, Masoumeh, et al.
Published: (2026)
Motion Manipulation via Unsupervised Keypoint Positioning in Face Animation
by: Li, Hong, et al.
Published: (2026)
by: Li, Hong, et al.
Published: (2026)
Keypoint Counting Classifiers: Turning Vision Transformers into Self-Explainable Models Without Training
by: Wickstrøm, Kristoffer, et al.
Published: (2025)
by: Wickstrøm, Kristoffer, et al.
Published: (2025)
Exploring Stronger Transformer Representation Learning for Occluded Person Re-Identification
by: Ji, Zhangjian, et al.
Published: (2024)
by: Ji, Zhangjian, et al.
Published: (2024)
Dynamic Patch-aware Enrichment Transformer for Occluded Person Re-Identification
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Similar Items
-
Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics
by: Chapariniya, Masoumeh, et al.
Published: (2025) -
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
by: Farhadipour, Aref, et al.
Published: (2025) -
Foundation Model Embeddings Meet Blended Emotions: A Multimodal Fusion Approach for the BLEMORE Challenge
by: Chapariniya, Masoumeh, et al.
Published: (2026) -
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
by: Farhadipour, Aref, et al.
Published: (2024) -
Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
by: Chapariniya, Masoumeh, et al.
Published: (2026)