Sound effects in media:A comparative analysis of recorded and synthetic samples in live-action and animation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garcia, Nelly, Reiss, Joshua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
Visual-based spatial audio generation system for multi-speaker environments
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
Soundscapes in Spectrograms: Pioneering Multilabel Classification for South Asian Sounds
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2026)
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2026)
SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
von: Sun, Luoyi, et al.
Veröffentlicht: (2026)
von: Sun, Luoyi, et al.
Veröffentlicht: (2026)
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
von: Wang, Chunshi, et al.
Veröffentlicht: (2025)
von: Wang, Chunshi, et al.
Veröffentlicht: (2025)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2024)
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2024)
The Sound of Entanglement
von: Rodríguez, Enar de Dios, et al.
Veröffentlicht: (2025)
von: Rodríguez, Enar de Dios, et al.
Veröffentlicht: (2025)
Manipulated Regions Localization For Partially Deepfake Audio: A Survey
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
A Distribution Matching Approach to Neural Piano Transcription with Optimal Transport
von: Wei, Weixing, et al.
Veröffentlicht: (2026)
von: Wei, Weixing, et al.
Veröffentlicht: (2026)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
von: Qiu, Ke, et al.
Veröffentlicht: (2026)
von: Qiu, Ke, et al.
Veröffentlicht: (2026)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
SonicVisionLM: Playing Sound with Vision Language Models
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation
von: Zhu, Jian, et al.
Veröffentlicht: (2026)
von: Zhu, Jian, et al.
Veröffentlicht: (2026)
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
Can We Hear from Events? Generating Speech from Event Camera
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
RenCon 2025: Revival of the Expressive Performance Rendering Competition
von: Zhang, Huan, et al.
Veröffentlicht: (2026)
von: Zhang, Huan, et al.
Veröffentlicht: (2026)
Research on Piano Timbre Transformation System Based on Diffusion Model
von: Hsu, Chun-Chieh, et al.
Veröffentlicht: (2026)
von: Hsu, Chun-Chieh, et al.
Veröffentlicht: (2026)
Generative Audio Extension and Morphing
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding
von: Wang, Jieyi, et al.
Veröffentlicht: (2026)
von: Wang, Jieyi, et al.
Veröffentlicht: (2026)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
von: Yang, Meng, et al.
Veröffentlicht: (2026)
von: Yang, Meng, et al.
Veröffentlicht: (2026)
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
von: Fan, Congyi, et al.
Veröffentlicht: (2026)
von: Fan, Congyi, et al.
Veröffentlicht: (2026)
Timed text extraction from Taiwanese Kua-á-hì TV series
von: Huang, Tzu-Hung, et al.
Veröffentlicht: (2026)
von: Huang, Tzu-Hung, et al.
Veröffentlicht: (2026)
Private Speech Classification without Collapse: Stabilized DP Training and Offline Distillation
von: Wen, Yadi, et al.
Veröffentlicht: (2026)
von: Wen, Yadi, et al.
Veröffentlicht: (2026)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
von: Hu, Han, et al.
Veröffentlicht: (2025)
von: Hu, Han, et al.
Veröffentlicht: (2025)
Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
von: Wei, Weixing, et al.
Veröffentlicht: (2025)
von: Wei, Weixing, et al.
Veröffentlicht: (2025)
MusicWeaver: Composer-Style Structural Editing and Minute-Scale Coherent Music Generation
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
Towards Practical Real-Time Low-Latency Music Source Separation
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
von: Feng, Zhou, et al.
Veröffentlicht: (2025)
von: Feng, Zhou, et al.
Veröffentlicht: (2025)
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024) -
Visual-based spatial audio generation system for multi-speaker environments
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025) -
Soundscapes in Spectrograms: Pioneering Multilabel Classification for South Asian Sounds
von: Chakrabarty, Sudip, et al.
Veröffentlicht: (2026) -
SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
von: Sun, Luoyi, et al.
Veröffentlicht: (2026) -
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
von: Wang, Chunshi, et al.
Veröffentlicht: (2025)