Diseño de sonido para producciones audiovisuales e historias sonoras en el aula. Hacia una docencia creativa mediante el uso de herramientas inteligentes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Civit, Miguel, Cuadrado, Francisco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Effectively obtaining acoustic, visual and textual data from videos
von: León, Jorge E., et al.
Veröffentlicht: (2025)
von: León, Jorge E., et al.
Veröffentlicht: (2025)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025)
von: Kim, Haven, et al.
Veröffentlicht: (2025)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025)
von: Gu, Ke, et al.
Veröffentlicht: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Building Audio-Visual Digital Twins with Smartphones
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
von: Gong, Jingyao
Veröffentlicht: (2026)
von: Gong, Jingyao
Veröffentlicht: (2026)
Dance2MIDI: Dance-driven multi-instruments music generation
von: Han, Bo, et al.
Veröffentlicht: (2023)
von: Han, Bo, et al.
Veröffentlicht: (2023)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
Intelligent Text-Conditioned Music Generation
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
von: Ren, Yong, et al.
Veröffentlicht: (2024)
von: Ren, Yong, et al.
Veröffentlicht: (2024)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
von: Yao, Dong, et al.
Veröffentlicht: (2023)
von: Yao, Dong, et al.
Veröffentlicht: (2023)
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Trusted Fake Audio Detection Based on Dirichlet Distribution
von: Ding, Chi, et al.
Veröffentlicht: (2025)
von: Ding, Chi, et al.
Veröffentlicht: (2025)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
von: Huang, Lian, et al.
Veröffentlicht: (2024)
von: Huang, Lian, et al.
Veröffentlicht: (2024)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
SonicVisionLM: Playing Sound with Vision Language Models
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
von: Xie, Zhifeng, et al.
Veröffentlicht: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations
von: Wachter, Maximilian, et al.
Veröffentlicht: (2026)
von: Wachter, Maximilian, et al.
Veröffentlicht: (2026)
Dance-to-Music Generation with Encoder-based Textual Inversion
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
von: Yang, Kaixing, et al.
Veröffentlicht: (2024)
von: Yang, Kaixing, et al.
Veröffentlicht: (2024)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
von: Sun, Chang, et al.
Veröffentlicht: (2024)
von: Sun, Chang, et al.
Veröffentlicht: (2024)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
Flexible Control in Symbolic Music Generation via Musical Metadata
von: Han, Sangjun, et al.
Veröffentlicht: (2024)
von: Han, Sangjun, et al.
Veröffentlicht: (2024)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
StereoFoley: Object-Aware Stereo Audio Generation from Video
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation
von: Chan, Nolan, et al.
Veröffentlicht: (2026)
von: Chan, Nolan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Effectively obtaining acoustic, visual and textual data from videos
von: León, Jorge E., et al.
Veröffentlicht: (2025) -
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023) -
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025) -
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025) -
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)