FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yiming, Gu, Yicheng, Zeng, Yanhong, Xing, Zhening, Wang, Yuancheng, Wu, Zhizheng, Chen, Kai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MambaFoley: Foley Sound Generation using Selective State-Space Models
por: Colombo, Marco Furio, et al.
Publicado: (2024)
por: Colombo, Marco Furio, et al.
Publicado: (2024)
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
por: Gu, Yicheng, et al.
Publicado: (2025)
por: Gu, Yicheng, et al.
Publicado: (2025)
Solid State Bus-Comp: A Large-Scale and Diverse Dataset for Dynamic Range Compressor Virtual Analog Modeling
por: Gu, Yicheng, et al.
Publicado: (2025)
por: Gu, Yicheng, et al.
Publicado: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
por: Zhang, Yaoyun, et al.
Publicado: (2024)
por: Zhang, Yaoyun, et al.
Publicado: (2024)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
por: Zhang, Xueyao, et al.
Publicado: (2025)
por: Zhang, Xueyao, et al.
Publicado: (2025)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
por: Gu, Yicheng, et al.
Publicado: (2024)
por: Gu, Yicheng, et al.
Publicado: (2024)
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
por: Gu, Yicheng, et al.
Publicado: (2025)
por: Gu, Yicheng, et al.
Publicado: (2025)
Aliasing-Free Neural Audio Synthesis
por: Gu, Yicheng, et al.
Publicado: (2025)
por: Gu, Yicheng, et al.
Publicado: (2025)
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
por: Gu, Yicheng, et al.
Publicado: (2024)
por: Gu, Yicheng, et al.
Publicado: (2024)
Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
por: Zhang, Xueyao, et al.
Publicado: (2023)
por: Zhang, Xueyao, et al.
Publicado: (2023)
CAFA: a Controllable Automatic Foley Artist
por: Benita, Roi, et al.
Publicado: (2025)
por: Benita, Roi, et al.
Publicado: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
por: Huang, Zhiqi, et al.
Publicado: (2024)
por: Huang, Zhiqi, et al.
Publicado: (2024)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2025)
por: Li, Jiaqi, et al.
Publicado: (2025)
StereoFoley: Object-Aware Stereo Audio Generation from Video
por: Karchkhadze, Tornike, et al.
Publicado: (2025)
por: Karchkhadze, Tornike, et al.
Publicado: (2025)
FoleyBench: A Benchmark For Video-to-Audio Models
por: Dixit, Satvik, et al.
Publicado: (2025)
por: Dixit, Satvik, et al.
Publicado: (2025)
Video-Guided Foley Sound Generation with Multimodal Controls
por: Chen, Ziyang, et al.
Publicado: (2024)
por: Chen, Ziyang, et al.
Publicado: (2024)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
por: Zhang, Xueyao, et al.
Publicado: (2023)
por: Zhang, Xueyao, et al.
Publicado: (2023)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
por: Yemini, Yochai, et al.
Publicado: (2023)
por: Yemini, Yochai, et al.
Publicado: (2023)
Optimized Loudspeaker Panning for Adaptive Sound-Field Correction and Non-stationary Listening Areas
por: Luo, Yuancheng
Publicado: (2025)
por: Luo, Yuancheng
Publicado: (2025)
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
por: Lee, Junwon, et al.
Publicado: (2024)
por: Lee, Junwon, et al.
Publicado: (2024)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
por: Wang, Yuancheng, et al.
Publicado: (2025)
por: Wang, Yuancheng, et al.
Publicado: (2025)
Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis
por: Wang, Junnuo
Publicado: (2025)
por: Wang, Junnuo
Publicado: (2025)
Constant Directivity Loudspeaker Beamforming
por: Luo, Yuancheng
Publicado: (2024)
por: Luo, Yuancheng
Publicado: (2024)
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
por: He, Haorui, et al.
Publicado: (2025)
por: He, Haorui, et al.
Publicado: (2025)
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
por: Ren, Yong, et al.
Publicado: (2025)
por: Ren, Yong, et al.
Publicado: (2025)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
por: Xie, Zeyu, et al.
Publicado: (2023)
por: Xie, Zeyu, et al.
Publicado: (2023)
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
por: Gramaccioni, Riccardo Fosco, et al.
Publicado: (2024)
por: Gramaccioni, Riccardo Fosco, et al.
Publicado: (2024)
Leveraging Sound Source Trajectories for Universal Sound Separation
por: Wu, Donghang, et al.
Publicado: (2024)
por: Wu, Donghang, et al.
Publicado: (2024)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
por: Shan, Sizhe, et al.
Publicado: (2025)
por: Shan, Sizhe, et al.
Publicado: (2025)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
por: Zhou, Xuehao, et al.
Publicado: (2024)
por: Zhou, Xuehao, et al.
Publicado: (2024)
An Initial Investigation of Neural Replay Simulator for Over-the-Air Adversarial Perturbations to Automatic Speaker Verification
por: Li, Jiaqi, et al.
Publicado: (2023)
por: Li, Jiaqi, et al.
Publicado: (2023)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
por: Gu, Ke, et al.
Publicado: (2025)
por: Gu, Ke, et al.
Publicado: (2025)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
por: Rai, Aashish, et al.
Publicado: (2024)
por: Rai, Aashish, et al.
Publicado: (2024)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
por: Kim, Ji-Hoon, et al.
Publicado: (2023)
por: Kim, Ji-Hoon, et al.
Publicado: (2023)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
por: Fang, Zihao, et al.
Publicado: (2026)
por: Fang, Zihao, et al.
Publicado: (2026)
SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
por: Jiang, Shaohan, et al.
Publicado: (2025)
por: Jiang, Shaohan, et al.
Publicado: (2025)
The CCF AATC 2025 Speech Restoration Challenge: A Retrospective
por: Zhang, Junan, et al.
Publicado: (2025)
por: Zhang, Junan, et al.
Publicado: (2025)
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
por: Xu, Xuenan, et al.
Publicado: (2024)
por: Xu, Xuenan, et al.
Publicado: (2024)
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
por: Li, Baihan, et al.
Publicado: (2024)
por: Li, Baihan, et al.
Publicado: (2024)
Ejemplares similares
-
MambaFoley: Foley Sound Generation using Selective State-Space Models
por: Colombo, Marco Furio, et al.
Publicado: (2024) -
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
por: Gu, Yicheng, et al.
Publicado: (2025) -
Solid State Bus-Comp: A Large-Scale and Diverse Dataset for Dynamic Range Compressor Virtual Analog Modeling
por: Gu, Yicheng, et al.
Publicado: (2025) -
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
por: Zhang, Yaoyun, et al.
Publicado: (2024) -
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
por: Zhang, Xueyao, et al.
Publicado: (2025)