2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guichoux, Téo, Soulier, Laure, Obin, Nicolas, Pelachaud, Catherine |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
Graph Modelling Analysis of Speech-Gesture Interaction for Aphasia Severity Estimation
von: Kollapally, Navya Martin, et al.
Veröffentlicht: (2026)
von: Kollapally, Navya Martin, et al.
Veröffentlicht: (2026)
Investigating the impact of 2D gesture representation on co-speech gesture generation
von: Guichoux, Teo, et al.
Veröffentlicht: (2024)
von: Guichoux, Teo, et al.
Veröffentlicht: (2024)
Rethinking Discrete Speech Representation Tokens for Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
UniCoM: A Universal Code-Switching Speech Generator
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
von: Pang, Haozhou, et al.
Veröffentlicht: (2024)
von: Pang, Haozhou, et al.
Veröffentlicht: (2024)
Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
von: Zhang, Zeyi, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyi, et al.
Veröffentlicht: (2024)
Convexity-based Pruning of Speech Representation Models
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
von: Xu, Chun, et al.
Veröffentlicht: (2024)
von: Xu, Chun, et al.
Veröffentlicht: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
USAD: Universal Speech and Audio Representation via Distillation
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2025)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2025)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Task-Agnostic Structured Pruning of Speech Representation Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2023)
von: Wang, Haoyu, et al.
Veröffentlicht: (2023)
Investigation of Speaker Representation for Target-Speaker Speech Processing
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
What Are They Doing? Joint Audio-Speech Co-Reasoning
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
von: Wang, Yingzhi, et al.
Veröffentlicht: (2024)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Crossmodal ASR Error Correction with Discrete Speech Units
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
How Much Context Does My Attention-Based ASR System Need?
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
von: Guichoux, Téo, et al.
Veröffentlicht: (2025) -
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024) -
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024) -
Graph Modelling Analysis of Speech-Gesture Interaction for Aphasia Severity Estimation
von: Kollapally, Navya Martin, et al.
Veröffentlicht: (2026) -
Investigating the impact of 2D gesture representation on co-speech gesture generation
von: Guichoux, Teo, et al.
Veröffentlicht: (2024)