Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
Fuente:
arXiv
Saved in:
| Main Authors: | Kucherenko, Taras, Wolfert, Pieter, Yoon, Youngwoo, Viegas, Carla, Nikolov, Teodor, Tsakov, Mihail, Henter, Gustav Eje |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards a GENEA Leaderboard -- an Extended, Living Benchmark for Evaluating and Advancing Conversational Motion Synthesis
by: Nagy, Rajmund, et al.
Published: (2024)
by: Nagy, Rajmund, et al.
Published: (2024)
Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
by: Nagy, Rajmund, et al.
Published: (2025)
by: Nagy, Rajmund, et al.
Published: (2025)
Unified speech and gesture synthesis using flow matching
by: Mehta, Shivam, et al.
Published: (2023)
by: Mehta, Shivam, et al.
Published: (2023)
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
by: Mehta, Shivam, et al.
Published: (2024)
by: Mehta, Shivam, et al.
Published: (2024)
MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios
by: Li, Xiangxian, et al.
Published: (2025)
by: Li, Xiangxian, et al.
Published: (2025)
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
by: Chen, Zeyu, et al.
Published: (2024)
by: Chen, Zeyu, et al.
Published: (2024)
Racism in the Machine: Visualization Ethics in Digital Humanities Projects
by: Hepworth, K. J., et al.
Published: (2024)
by: Hepworth, K. J., et al.
Published: (2024)
Seeing The Words: Evaluating AI-generated Biblical Art
by: Makimei, Hidde, et al.
Published: (2025)
by: Makimei, Hidde, et al.
Published: (2025)
Digital analysis of early color photographs taken using regular color screen processes
by: Hubička, Jan, et al.
Published: (2023)
by: Hubička, Jan, et al.
Published: (2023)
A Low-Latency 3D Live Remote Visualization System for Tourist Sites Integrating Dynamic and Pre-captured Static Point Clouds
by: Matsumoto, Takahiro, et al.
Published: (2025)
by: Matsumoto, Takahiro, et al.
Published: (2025)
Expert Insight-Enhanced Follow-up Chest X-Ray Summary Generation
by: Wang, Zhichuan, et al.
Published: (2024)
by: Wang, Zhichuan, et al.
Published: (2024)
Matcha-TTS: A fast TTS architecture with conditional flow matching
by: Mehta, Shivam, et al.
Published: (2023)
by: Mehta, Shivam, et al.
Published: (2023)
Leum-VL Technical Report
by: He, Yuxuan, et al.
Published: (2026)
by: He, Yuxuan, et al.
Published: (2026)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
by: Mehta, Shivam, et al.
Published: (2024)
by: Mehta, Shivam, et al.
Published: (2024)
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
by: Wu, Kunlin, et al.
Published: (2026)
by: Wu, Kunlin, et al.
Published: (2026)
Saliency-Aware Diffusion Reconstruction for Effective Invisible Watermark Removal
by: Alam, Inzamamul, et al.
Published: (2025)
by: Alam, Inzamamul, et al.
Published: (2025)
Step-Aware Residual-Guided Diffusion for EEG Spatial Super-Resolution
by: Liu, Hongjun, et al.
Published: (2025)
by: Liu, Hongjun, et al.
Published: (2025)
P$^2$U: Progressive Precision Update For Efficient Model Distribution
by: Afrabandpey, Homayun, et al.
Published: (2025)
by: Afrabandpey, Homayun, et al.
Published: (2025)
MetaErr: Towards Predicting Error Patterns in Deep Neural Networks
by: Totakura, Varun, et al.
Published: (2026)
by: Totakura, Varun, et al.
Published: (2026)
Lens Distortion Encoding System Version 1.0
by: Fober, Jakub Maksymilian
Published: (2024)
by: Fober, Jakub Maksymilian
Published: (2024)
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
by: Li, Lin, et al.
Published: (2026)
by: Li, Lin, et al.
Published: (2026)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
by: Korolkov, Vasilii
Published: (2025)
by: Korolkov, Vasilii
Published: (2025)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
by: Zhong, Zesen, et al.
Published: (2025)
by: Zhong, Zesen, et al.
Published: (2025)
Towards Efficient 3D Gaussian Human Avatar Compression: A Prior-Guided Framework
by: Yin, Shanzhi, et al.
Published: (2025)
by: Yin, Shanzhi, et al.
Published: (2025)
AVControl: Efficient Framework for Training Audio-Visual Controls
by: Ben-Yosef, Matan, et al.
Published: (2026)
by: Ben-Yosef, Matan, et al.
Published: (2026)
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
by: Lu, Wanglong, et al.
Published: (2024)
by: Lu, Wanglong, et al.
Published: (2024)
FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing
by: Lu, Wanglong, et al.
Published: (2024)
by: Lu, Wanglong, et al.
Published: (2024)
Multi-level SSL Feature Gating for Audio Deepfake Detection
by: Tran, Hoan My, et al.
Published: (2025)
by: Tran, Hoan My, et al.
Published: (2025)
Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
by: Neptune, Nathalie, et al.
Published: (2025)
by: Neptune, Nathalie, et al.
Published: (2025)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
by: Boumber, Dainis, et al.
Published: (2024)
by: Boumber, Dainis, et al.
Published: (2024)
Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular Video
by: Choi, Seonghwa, et al.
Published: (2025)
by: Choi, Seonghwa, et al.
Published: (2025)
Improving Visual Object Tracking through Visual Prompting
by: Chen, Shih-Fang, et al.
Published: (2024)
by: Chen, Shih-Fang, et al.
Published: (2024)
Two-step Authentication: Multi-biometric System Using Voice and Facial Recognition
by: Chen, Kuan Wei, et al.
Published: (2026)
by: Chen, Kuan Wei, et al.
Published: (2026)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
by: Aghdam, Amir, et al.
Published: (2025)
by: Aghdam, Amir, et al.
Published: (2025)
Lightweight Complementary-Cue Fusion for Robust Video Face Forgery Detection
by: Baek, Sunghwan, et al.
Published: (2026)
by: Baek, Sunghwan, et al.
Published: (2026)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
by: Tu, Songjun, et al.
Published: (2025)
by: Tu, Songjun, et al.
Published: (2025)
MetaDigiHuman: Haptic Interfaces for Digital Humans in Metaverse
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity
by: Chen, Yuzhuo, et al.
Published: (2025)
by: Chen, Yuzhuo, et al.
Published: (2025)
Similar Items
-
Towards a GENEA Leaderboard -- an Extended, Living Benchmark for Evaluating and Advancing Conversational Motion Synthesis
by: Nagy, Rajmund, et al.
Published: (2024) -
Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
by: Nagy, Rajmund, et al.
Published: (2025) -
Unified speech and gesture synthesis using flow matching
by: Mehta, Shivam, et al.
Published: (2023) -
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
by: Mehta, Shivam, et al.
Published: (2024) -
MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios
by: Li, Xiangxian, et al.
Published: (2025)