Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Mehta, Shivam, Deichler, Anna, O'Regan, Jim, Moëll, Birger, Beskow, Jonas, Henter, Gustav Eje, Alexanderson, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified speech and gesture synthesis using flow matching
by: Mehta, Shivam, et al.
Published: (2023)
by: Mehta, Shivam, et al.
Published: (2023)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
by: Mehta, Shivam, et al.
Published: (2024)
by: Mehta, Shivam, et al.
Published: (2024)
Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Grounded Gesture Generation: Language, Motion, and Space
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Matcha-TTS: A fast TTS architecture with conditional flow matching
by: Mehta, Shivam, et al.
Published: (2023)
by: Mehta, Shivam, et al.
Published: (2023)
Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Gesture Evaluation in Virtual Reality
by: Werner, Axel Wiebe, et al.
Published: (2025)
by: Werner, Axel Wiebe, et al.
Published: (2025)
Incorporating Spatial Awareness in Data-Driven Gesture Generation for Virtual Agents
by: Deichler, Anna, et al.
Published: (2024)
by: Deichler, Anna, et al.
Published: (2024)
Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions
by: Yadav, Arin Gopalan, et al.
Published: (2026)
by: Yadav, Arin Gopalan, et al.
Published: (2026)
Faked Speech Detection with Zero Prior Knowledge
by: Ajmi, Sahar Al, et al.
Published: (2022)
by: Ajmi, Sahar Al, et al.
Published: (2022)
A Landmark-Aware Visual Navigation Dataset
by: Johnson, Faith, et al.
Published: (2024)
by: Johnson, Faith, et al.
Published: (2024)
DF-DM: A foundational process model for multimodal data fusion in the artificial intelligence era
by: Restrepo, David, et al.
Published: (2024)
by: Restrepo, David, et al.
Published: (2024)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
by: Platzer, André
Published: (2024)
by: Platzer, André
Published: (2024)
Few-Class Arena: A Benchmark for Efficient Selection of Vision Models and Dataset Difficulty Measurement
by: Cao, Bryan Bo, et al.
Published: (2024)
by: Cao, Bryan Bo, et al.
Published: (2024)
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery
by: Su, Jiarui, et al.
Published: (2026)
by: Su, Jiarui, et al.
Published: (2026)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
by: Huo, Dongjie, et al.
Published: (2026)
by: Huo, Dongjie, et al.
Published: (2026)
On measuring grounding and generalizing grounding problems
by: Quigley, Daniel, et al.
Published: (2025)
by: Quigley, Daniel, et al.
Published: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
by: Jia, Xiao
Published: (2026)
by: Jia, Xiao
Published: (2026)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior
by: Kim, Dongmin, et al.
Published: (2025)
by: Kim, Dongmin, et al.
Published: (2025)
Semantic Modeling for World-Centered Architectures
by: Mantsivoda, Andrei, et al.
Published: (2026)
by: Mantsivoda, Andrei, et al.
Published: (2026)
A Cost-Effective Eye-Tracker for Early Detection of Mild Cognitive Impairment
by: Greco, Danilo, et al.
Published: (2024)
by: Greco, Danilo, et al.
Published: (2024)
YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception
by: Impraimakis, Marios, et al.
Published: (2026)
by: Impraimakis, Marios, et al.
Published: (2026)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
by: Guichoux, Téo, et al.
Published: (2025)
by: Guichoux, Téo, et al.
Published: (2025)
PhysMorph-GS: Render-Guided Volumetric Morphing with Differentiable Physics
by: Song, Chang-Yong, et al.
Published: (2025)
by: Song, Chang-Yong, et al.
Published: (2025)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
by: Koh, Hyunseo, et al.
Published: (2026)
by: Koh, Hyunseo, et al.
Published: (2026)
Structured Basis Function Networks: Loss-Centric Multi-Hypothesis Ensembles with Controllable Diversity
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
by: Koubaa, Anis, et al.
Published: (2025)
by: Koubaa, Anis, et al.
Published: (2025)
Swish-T : Enhancing Swish Activation with Tanh Bias for Improved Neural Network Performance
by: Seo, Youngmin, et al.
Published: (2024)
by: Seo, Youngmin, et al.
Published: (2024)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
by: Chen, Kewei, et al.
Published: (2025)
by: Chen, Kewei, et al.
Published: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
by: Huang, Yingbing, et al.
Published: (2025)
by: Huang, Yingbing, et al.
Published: (2025)
Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
by: Cao, Deyu, et al.
Published: (2025)
by: Cao, Deyu, et al.
Published: (2025)
LLM-supported document separation for printed reviews from zbMATH Open
by: Pluzhnikov, Ivan, et al.
Published: (2026)
by: Pluzhnikov, Ivan, et al.
Published: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
by: Kalušev, Vladimir, et al.
Published: (2025)
by: Kalušev, Vladimir, et al.
Published: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
Similar Items
-
Unified speech and gesture synthesis using flow matching
by: Mehta, Shivam, et al.
Published: (2023) -
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
by: Mehta, Shivam, et al.
Published: (2024) -
Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
by: Deichler, Anna, et al.
Published: (2025) -
Grounded Gesture Generation: Language, Motion, and Space
by: Deichler, Anna, et al.
Published: (2025) -
Matcha-TTS: A fast TTS architecture with conditional flow matching
by: Mehta, Shivam, et al.
Published: (2023)