CAFA: a Controllable Automatic Foley Artist
Fuente:
arXiv
Guardado en:
| Autores principales: | Benita, Roi, Finkelson, Michael, Halperin, Tavi, Sterkin, Gleb, Adi, Yossi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
NAST: Noise Aware Speech Tokenization for Speech Language Models
por: Messica, Shoval, et al.
Publicado: (2024)
por: Messica, Shoval, et al.
Publicado: (2024)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
por: Benita, Roi, et al.
Publicado: (2023)
por: Benita, Roi, et al.
Publicado: (2023)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
por: Zeldes, Ella, et al.
Publicado: (2024)
por: Zeldes, Ella, et al.
Publicado: (2024)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
por: Tal, Or, et al.
Publicado: (2024)
por: Tal, Or, et al.
Publicado: (2024)
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
por: Aziz, Shiran, et al.
Publicado: (2024)
por: Aziz, Shiran, et al.
Publicado: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
por: Colombo, Marco Furio, et al.
Publicado: (2024)
por: Colombo, Marco Furio, et al.
Publicado: (2024)
LAST: Language Model Aware Speech Tokenization
por: Turetzky, Arnon, et al.
Publicado: (2024)
por: Turetzky, Arnon, et al.
Publicado: (2024)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
por: Rouard, Simon, et al.
Publicado: (2025)
por: Rouard, Simon, et al.
Publicado: (2025)
Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis
por: Wang, Junnuo
Publicado: (2025)
por: Wang, Junnuo
Publicado: (2025)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
por: Rouard, Simon, et al.
Publicado: (2024)
por: Rouard, Simon, et al.
Publicado: (2024)
The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
por: Coelho, Guilherme
Publicado: (2025)
por: Coelho, Guilherme
Publicado: (2025)
Latent Watermarking of Audio Generative Models
por: Roman, Robin San, et al.
Publicado: (2024)
por: Roman, Robin San, et al.
Publicado: (2024)
WHISTRESS: Enriching Transcriptions with Sentence Stress Detection
por: Yosha, Iddo, et al.
Publicado: (2025)
por: Yosha, Iddo, et al.
Publicado: (2025)
StressTest: Can YOUR Speech LM Handle the Stress?
por: Yosha, Iddo, et al.
Publicado: (2025)
por: Yosha, Iddo, et al.
Publicado: (2025)
Salmon: A Suite for Acoustic Language Model Evaluation
por: Maimon, Gallil, et al.
Publicado: (2024)
por: Maimon, Gallil, et al.
Publicado: (2024)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
por: Roth, Amit, et al.
Publicado: (2024)
por: Roth, Amit, et al.
Publicado: (2024)
Scaling Analysis of Interleaved Speech-Text Language Models
por: Maimon, Gallil, et al.
Publicado: (2025)
por: Maimon, Gallil, et al.
Publicado: (2025)
StereoFoley: Object-Aware Stereo Audio Generation from Video
por: Karchkhadze, Tornike, et al.
Publicado: (2025)
por: Karchkhadze, Tornike, et al.
Publicado: (2025)
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
por: Tal, Or, et al.
Publicado: (2025)
por: Tal, Or, et al.
Publicado: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
por: Huang, Zhiqi, et al.
Publicado: (2024)
por: Huang, Zhiqi, et al.
Publicado: (2024)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
por: Chang, Xuankai, et al.
Publicado: (2024)
por: Chang, Xuankai, et al.
Publicado: (2024)
FoleyBench: A Benchmark For Video-to-Audio Models
por: Dixit, Satvik, et al.
Publicado: (2025)
por: Dixit, Satvik, et al.
Publicado: (2025)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
por: Hsu, Po-chun, et al.
Publicado: (2023)
por: Hsu, Po-chun, et al.
Publicado: (2023)
Low-Resource Audio Codec (LRAC): 2025 Challenge Description
por: Wojcicki, Kamil, et al.
Publicado: (2025)
por: Wojcicki, Kamil, et al.
Publicado: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
por: Har-Tuv, Nadav, et al.
Publicado: (2025)
por: Har-Tuv, Nadav, et al.
Publicado: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
por: Zhang, Yaoyun, et al.
Publicado: (2024)
por: Zhang, Yaoyun, et al.
Publicado: (2024)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
por: Shan, Sizhe, et al.
Publicado: (2025)
por: Shan, Sizhe, et al.
Publicado: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
por: Tian, Jingguang, et al.
Publicado: (2024)
por: Tian, Jingguang, et al.
Publicado: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
por: Chen, Ziyang, et al.
Publicado: (2024)
por: Chen, Ziyang, et al.
Publicado: (2024)
Automatic Music Mixing using a Generative Model of Effect Embeddings
por: Moliner, Eloi, et al.
Publicado: (2025)
por: Moliner, Eloi, et al.
Publicado: (2025)
Automatic Detection and Annotation of Sperm Whale Codas
por: Gubnitsky, Guy, et al.
Publicado: (2024)
por: Gubnitsky, Guy, et al.
Publicado: (2024)
Phonetic Richness for Improved Automatic Speaker Verification
por: Klein, Nicholas, et al.
Publicado: (2024)
por: Klein, Nicholas, et al.
Publicado: (2024)
BWSNet: Automatic Perceptual Assessment of Audio Signals
por: Veillon, Clément Le Moine, et al.
Publicado: (2023)
por: Veillon, Clément Le Moine, et al.
Publicado: (2023)
Automatic Melody Reduction via Shortest Path Finding
por: Wang, Ziyu, et al.
Publicado: (2025)
por: Wang, Ziyu, et al.
Publicado: (2025)
A Dataset for Automatic Assessment of TTS Quality in Spanish
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
por: Ma, Guobin, et al.
Publicado: (2026)
por: Ma, Guobin, et al.
Publicado: (2026)
Slamming: Training a Speech Language Model on One GPU in a Day
por: Maimon, Gallil, et al.
Publicado: (2025)
por: Maimon, Gallil, et al.
Publicado: (2025)
Exploiting Music Source Separation for Automatic Lyrics Transcription with Whisper
por: Syed, Jaza, et al.
Publicado: (2025)
por: Syed, Jaza, et al.
Publicado: (2025)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
por: Bondaruk, Łukasz, et al.
Publicado: (2024)
por: Bondaruk, Łukasz, et al.
Publicado: (2024)
Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2024)
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2024)
Ejemplares similares
-
NAST: Noise Aware Speech Tokenization for Speech Language Models
por: Messica, Shoval, et al.
Publicado: (2024) -
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
por: Benita, Roi, et al.
Publicado: (2023) -
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
por: Zeldes, Ella, et al.
Publicado: (2024) -
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
por: Tal, Or, et al.
Publicado: (2024) -
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
por: Aziz, Shiran, et al.
Publicado: (2024)