Gespeichert in:
| Hauptverfasser: | Gao, Rong, Liu, Xin, Hu, Zhuozhao, Xing, Bohao, Xia, Baiqiang, Yu, Zitong, Kälviäinen, Heikki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2504.19514 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DEEMO: De-identity Multimodal Emotion Recognition and Reasoning
von: Li, Deng, et al.
Veröffentlicht: (2025)
von: Li, Deng, et al.
Veröffentlicht: (2025)
Identity-free Artificial Emotional Intelligence via Micro-Gesture Understanding
von: Gao, Rong, et al.
Veröffentlicht: (2024)
von: Gao, Rong, et al.
Veröffentlicht: (2024)
Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer
von: Xing, Bohao, et al.
Veröffentlicht: (2026)
von: Xing, Bohao, et al.
Veröffentlicht: (2026)
AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
von: Li, Deng, et al.
Veröffentlicht: (2024)
von: Li, Deng, et al.
Veröffentlicht: (2024)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
von: Li, Deng, et al.
Veröffentlicht: (2025)
von: Li, Deng, et al.
Veröffentlicht: (2025)
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
DiffFAS: Face Anti-Spoofing via Generative Diffusion Models
von: Ge, Xinxu, et al.
Veröffentlicht: (2024)
von: Ge, Xinxu, et al.
Veröffentlicht: (2024)
YourSkatingCoach: A Figure Skating Video Benchmark for Fine-Grained Element Analysis
von: Chen, Wei-Yi, et al.
Veröffentlicht: (2024)
von: Chen, Wei-Yi, et al.
Veröffentlicht: (2024)
FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
von: Hu, Zhuozhao, et al.
Veröffentlicht: (2025)
von: Hu, Zhuozhao, et al.
Veröffentlicht: (2025)
Enhancing Micro Gesture Recognition for Emotion Understanding via Context-aware Visual-Text Contrastive Learning
von: Li, Deng, et al.
Veröffentlicht: (2024)
von: Li, Deng, et al.
Veröffentlicht: (2024)
Fine-grained Action Analysis: A Multi-modality and Multi-task Dataset of Figure Skating
von: Liu, Sheng-Lan, et al.
Veröffentlicht: (2023)
von: Liu, Sheng-Lan, et al.
Veröffentlicht: (2023)
Self-Supervised Pretraining for Fine-Grained Plankton Recognition
von: Kareinen, Joona, et al.
Veröffentlicht: (2025)
von: Kareinen, Joona, et al.
Veröffentlicht: (2025)
Understanding the Impact of Training Set Size on Animal Re-identification
von: Algasov, Aleksandr, et al.
Veröffentlicht: (2024)
von: Algasov, Aleksandr, et al.
Veröffentlicht: (2024)
DAPlankton: Benchmark Dataset for Multi-instrument Plankton Recognition via Fine-grained Domain Adaptation
von: Batrakhanov, Daniel, et al.
Veröffentlicht: (2024)
von: Batrakhanov, Daniel, et al.
Veröffentlicht: (2024)
The SkatingVerse Workshop & Challenge: Methods and Results
von: Zhao, Jian, et al.
Veröffentlicht: (2024)
von: Zhao, Jian, et al.
Veröffentlicht: (2024)
Unsupervised Pelage Pattern Unwrapping for Animal Re-identification
von: Algasov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Algasov, Aleksandr, et al.
Veröffentlicht: (2025)
VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
Open-Set Plankton Recognition
von: Kareinen, Joona, et al.
Veröffentlicht: (2025)
von: Kareinen, Joona, et al.
Veröffentlicht: (2025)
Cross-modal learning for plankton recognition
von: Kareinen, Joona, et al.
Veröffentlicht: (2026)
von: Kareinen, Joona, et al.
Veröffentlicht: (2026)
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
von: Ye, Qilang, et al.
Veröffentlicht: (2024)
von: Ye, Qilang, et al.
Veröffentlicht: (2024)
Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating Assessment
von: Wang, Fengshun, et al.
Veröffentlicht: (2025)
von: Wang, Fengshun, et al.
Veröffentlicht: (2025)
Intelligent Artistic Typography: A Comprehensive Review of Artistic Text Design and Generation
von: Bai, Yuhang, et al.
Veröffentlicht: (2024)
von: Bai, Yuhang, et al.
Veröffentlicht: (2024)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
Paintings and Drawings Aesthetics Assessment with Rich Attributes for Various Artistic Categories
von: Jin, Xin, et al.
Veröffentlicht: (2024)
von: Jin, Xin, et al.
Veröffentlicht: (2024)
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
von: Do, Jeonghyeok, et al.
Veröffentlicht: (2024)
von: Do, Jeonghyeok, et al.
Veröffentlicht: (2024)
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
von: Xia, Haotian, et al.
Veröffentlicht: (2025)
von: Xia, Haotian, et al.
Veröffentlicht: (2025)
TennisExpert: Towards Expert-Level Analytical Sports Video Understanding
von: Liu, Zhaoyu, et al.
Veröffentlicht: (2026)
von: Liu, Zhaoyu, et al.
Veröffentlicht: (2026)
Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3D
von: Chen, Ping, et al.
Veröffentlicht: (2026)
von: Chen, Ping, et al.
Veröffentlicht: (2026)
AnyArtisticGlyph: Multilingual Controllable Artistic Glyph Generation
von: Lu, Xiongbo, et al.
Veröffentlicht: (2025)
von: Lu, Xiongbo, et al.
Veröffentlicht: (2025)
1st Place Solution to the 1st SkatingVerse Challenge
von: Sun, Tao, et al.
Veröffentlicht: (2024)
von: Sun, Tao, et al.
Veröffentlicht: (2024)
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2025)
Training-Free Efficient Video Generation via Dynamic Token Carving
von: Zhang, Yuechen, et al.
Veröffentlicht: (2025)
von: Zhang, Yuechen, et al.
Veröffentlicht: (2025)
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
YUV20K: A Complexity-Driven Benchmark and Trajectory-Aware Alignment Model for Video Camouflaged Object Detection
von: Liu, Yiyu, et al.
Veröffentlicht: (2026)
von: Liu, Yiyu, et al.
Veröffentlicht: (2026)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
von: Zhang, Da, et al.
Veröffentlicht: (2025)
von: Zhang, Da, et al.
Veröffentlicht: (2025)
Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure
von: Li, Zitong, et al.
Veröffentlicht: (2026)
von: Li, Zitong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DEEMO: De-identity Multimodal Emotion Recognition and Reasoning
von: Li, Deng, et al.
Veröffentlicht: (2025) -
Identity-free Artificial Emotional Intelligence via Micro-Gesture Understanding
von: Gao, Rong, et al.
Veröffentlicht: (2024) -
Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer
von: Xing, Bohao, et al.
Veröffentlicht: (2026) -
AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection
von: Xing, Bohao, et al.
Veröffentlicht: (2025) -
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
von: Li, Deng, et al.
Veröffentlicht: (2024)