Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Zhen, Teye, Mattias, Yadgaroff, Derek, Bütepage, Judith |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
di: Bhattacharya, Aneesh, et al.
Pubblicazione: (2023)
di: Bhattacharya, Aneesh, et al.
Pubblicazione: (2023)
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
di: Yang, Kaixing, et al.
Pubblicazione: (2025)
di: Yang, Kaixing, et al.
Pubblicazione: (2025)
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
di: Pina, Leonardo, et al.
Pubblicazione: (2024)
di: Pina, Leonardo, et al.
Pubblicazione: (2024)
MusicScore: A Dataset for Music Score Modeling and Generation
di: Lin, Yuheng, et al.
Pubblicazione: (2024)
di: Lin, Yuheng, et al.
Pubblicazione: (2024)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
di: Zhang, Fan, et al.
Pubblicazione: (2023)
di: Zhang, Fan, et al.
Pubblicazione: (2023)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
di: Xie, Yifan, et al.
Pubblicazione: (2024)
di: Xie, Yifan, et al.
Pubblicazione: (2024)
LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
di: Jung, Jongmin, et al.
Pubblicazione: (2025)
di: Jung, Jongmin, et al.
Pubblicazione: (2025)
Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video Deepfake Dataset
di: Kaur, Sukhandeep, et al.
Pubblicazione: (2024)
di: Kaur, Sukhandeep, et al.
Pubblicazione: (2024)
InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
di: Liu, Xiaoxing, et al.
Pubblicazione: (2025)
di: Liu, Xiaoxing, et al.
Pubblicazione: (2025)
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
di: Ji, Xiaozhong, et al.
Pubblicazione: (2024)
di: Ji, Xiaozhong, et al.
Pubblicazione: (2024)
It Takes Two: Real-time Co-Speech Two-person's Interaction Generation via Reactive Auto-regressive Diffusion Model
di: Shi, Mingyi, et al.
Pubblicazione: (2024)
di: Shi, Mingyi, et al.
Pubblicazione: (2024)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
UniMuMo: Unified Text, Music and Motion Generation
di: Yang, Han, et al.
Pubblicazione: (2024)
di: Yang, Han, et al.
Pubblicazione: (2024)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
di: Zhang, Yu, et al.
Pubblicazione: (2025)
di: Zhang, Yu, et al.
Pubblicazione: (2025)
Dance2MIDI: Dance-driven multi-instruments music generation
di: Han, Bo, et al.
Pubblicazione: (2023)
di: Han, Bo, et al.
Pubblicazione: (2023)
CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
di: Yang, Kaixing, et al.
Pubblicazione: (2024)
di: Yang, Kaixing, et al.
Pubblicazione: (2024)
QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation
di: Zhou, Zhizhen, et al.
Pubblicazione: (2024)
di: Zhou, Zhizhen, et al.
Pubblicazione: (2024)
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
di: Pham, Kien T., et al.
Pubblicazione: (2025)
di: Pham, Kien T., et al.
Pubblicazione: (2025)
EnchantDance: Unveiling the Potential of Music-Driven Dance Movement
di: Han, Bo, et al.
Pubblicazione: (2023)
di: Han, Bo, et al.
Pubblicazione: (2023)
Flexible Control in Symbolic Music Generation via Musical Metadata
di: Han, Sangjun, et al.
Pubblicazione: (2024)
di: Han, Sangjun, et al.
Pubblicazione: (2024)
A Unified Framework for Modality-Agnostic Deepfakes Detection
di: Yu, Cai, et al.
Pubblicazione: (2023)
di: Yu, Cai, et al.
Pubblicazione: (2023)
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
di: Kim, Haven, et al.
Pubblicazione: (2025)
di: Kim, Haven, et al.
Pubblicazione: (2025)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
di: Gu, Ke, et al.
Pubblicazione: (2025)
di: Gu, Ke, et al.
Pubblicazione: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
di: Cui, Yang, et al.
Pubblicazione: (2025)
di: Cui, Yang, et al.
Pubblicazione: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
Building Audio-Visual Digital Twins with Smartphones
di: Lan, Zitong, et al.
Pubblicazione: (2025)
di: Lan, Zitong, et al.
Pubblicazione: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
di: Gong, Jingyao
Pubblicazione: (2026)
di: Gong, Jingyao
Pubblicazione: (2026)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
di: Yu, Fan, et al.
Pubblicazione: (2024)
di: Yu, Fan, et al.
Pubblicazione: (2024)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
di: Salvi, Davide, et al.
Pubblicazione: (2024)
di: Salvi, Davide, et al.
Pubblicazione: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
di: Kwak, Doyeop, et al.
Pubblicazione: (2026)
di: Kwak, Doyeop, et al.
Pubblicazione: (2026)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
di: Li, Yupei, et al.
Pubblicazione: (2024)
di: Li, Yupei, et al.
Pubblicazione: (2024)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
Intelligent Text-Conditioned Music Generation
di: Xie, Zhouyao, et al.
Pubblicazione: (2024)
di: Xie, Zhouyao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
di: Bhattacharya, Aneesh, et al.
Pubblicazione: (2023) -
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
di: Yang, Kaixing, et al.
Pubblicazione: (2025) -
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
di: Pina, Leonardo, et al.
Pubblicazione: (2024) -
MusicScore: A Dataset for Music Score Modeling and Generation
di: Lin, Yuheng, et al.
Pubblicazione: (2024) -
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
di: Zhang, Fan, et al.
Pubblicazione: (2023)