Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen
Fuente:
arXiv
Saved in:
| Main Authors: | Sánchez, María, Fernández, Laura, Arias, Julián, Cámara, Mateo, Comini, Giulia, Gabrys, Adam, Blanco, José Luis, Godino, Juan Ignacio, Hernández, Luis Alfonso |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vivo : une approche multimodale de la synthese concatenative par corpus dans le cadre d'une oeuvre audiovisuelle immersive
by: Fayet, Mateo
Published: (2024)
by: Fayet, Mateo
Published: (2024)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
by: Vecino, Biel Tura, et al.
Published: (2025)
by: Vecino, Biel Tura, et al.
Published: (2025)
Application of Whisper in Clinical Practice: the Post-Stroke Speech Assessment during a Naming Task
by: Davudova, Milena, et al.
Published: (2025)
by: Davudova, Milena, et al.
Published: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
by: Lin, Zhaofeng, et al.
Published: (2024)
by: Lin, Zhaofeng, et al.
Published: (2024)
Angular Distance Distribution Loss for Audio Classification
by: Almudévar, Antonio, et al.
Published: (2024)
by: Almudévar, Antonio, et al.
Published: (2024)
Online Audio-Visual Autoregressive Speaker Extraction
by: Pan, Zexu, et al.
Published: (2025)
by: Pan, Zexu, et al.
Published: (2025)
Vision Transformer Segmentation for Visual Bird Sound Denoising
by: Kumar, Sahil, et al.
Published: (2024)
by: Kumar, Sahil, et al.
Published: (2024)
Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
by: Piñeiro-Martín, Andrés, et al.
Published: (2024)
by: Piñeiro-Martín, Andrés, et al.
Published: (2024)
Decoding Vocal Articulations from Acoustic Latent Representations
by: Cámara, Mateo, et al.
Published: (2024)
by: Cámara, Mateo, et al.
Published: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
by: Liu, Huadai, et al.
Published: (2023)
by: Liu, Huadai, et al.
Published: (2023)
AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
by: Sharma, Sarthak, et al.
Published: (2024)
by: Sharma, Sarthak, et al.
Published: (2024)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
by: Sang, Wendi, et al.
Published: (2025)
by: Sang, Wendi, et al.
Published: (2025)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
by: Tao, Ruijie, et al.
Published: (2024)
by: Tao, Ruijie, et al.
Published: (2024)
Autonomous Soundscape Augmentation with Multimodal Fusion of Visual and Participant-linked Inputs
by: Ooi, Kenneth, et al.
Published: (2023)
by: Ooi, Kenneth, et al.
Published: (2023)
BickGraphing: Web-Based Application for Visual Inspection of Audio Recordings
by: Seow, Kayley, et al.
Published: (2026)
by: Seow, Kayley, et al.
Published: (2026)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
by: Chao, Rong, et al.
Published: (2025)
by: Chao, Rong, et al.
Published: (2025)
Sound-Based Spin Estimation in Table Tennis: Dataset and Real-Time Classification Pipeline
by: Gossard, Thomas, et al.
Published: (2024)
by: Gossard, Thomas, et al.
Published: (2024)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
by: Roman, Adrian S., et al.
Published: (2025)
by: Roman, Adrian S., et al.
Published: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
by: Jung, Chaeyoung, et al.
Published: (2024)
by: Jung, Chaeyoung, et al.
Published: (2024)
Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and Machine Learning Approaches
by: Pierotti, Francesco, et al.
Published: (2025)
by: Pierotti, Francesco, et al.
Published: (2025)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
by: Hussain, Tassadaq, et al.
Published: (2024)
by: Hussain, Tassadaq, et al.
Published: (2024)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
by: Jin, Zhan, et al.
Published: (2025)
by: Jin, Zhan, et al.
Published: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
by: Zhang, Jing-Xuan, et al.
Published: (2025)
by: Zhang, Jing-Xuan, et al.
Published: (2025)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
by: Tian, Wenjie, et al.
Published: (2026)
by: Tian, Wenjie, et al.
Published: (2026)
BWSNet: Automatic Perceptual Assessment of Audio Signals
by: Veillon, Clément Le Moine, et al.
Published: (2023)
by: Veillon, Clément Le Moine, et al.
Published: (2023)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
by: Ren, Wenze, et al.
Published: (2024)
by: Ren, Wenze, et al.
Published: (2024)
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features
by: Li, Sifei, et al.
Published: (2024)
by: Li, Sifei, et al.
Published: (2024)
Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods
by: Ferreira, Aníbal J. S., et al.
Published: (2025)
by: Ferreira, Aníbal J. S., et al.
Published: (2025)
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
by: Li, Zixuan, et al.
Published: (2025)
by: Li, Zixuan, et al.
Published: (2025)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
by: Daeglau, Mareike, et al.
Published: (2025)
by: Daeglau, Mareike, et al.
Published: (2025)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
by: Chen, Xueyuan, et al.
Published: (2024)
by: Chen, Xueyuan, et al.
Published: (2024)
AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
by: Zhang, Daning, et al.
Published: (2025)
by: Zhang, Daning, et al.
Published: (2025)
A dataset and model for auditory scene recognition for hearing devices: AHEAD-DS and OpenYAMNet
by: Zhong, Henry, et al.
Published: (2025)
by: Zhong, Henry, et al.
Published: (2025)
Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods
by: Lechler, Laura, et al.
Published: (2025)
by: Lechler, Laura, et al.
Published: (2025)
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
by: Chen, Yu, et al.
Published: (2025)
by: Chen, Yu, et al.
Published: (2025)
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
by: Xu, Weihan, et al.
Published: (2024)
by: Xu, Weihan, et al.
Published: (2024)
Efficient Sound Field Reconstruction with Conditional Invertible Neural Networks
by: Karakonstantis, Xenofon, et al.
Published: (2024)
by: Karakonstantis, Xenofon, et al.
Published: (2024)
Music to Dance as Language Translation using Sequence Models
by: Correia, André, et al.
Published: (2024)
by: Correia, André, et al.
Published: (2024)
Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
by: Wechsler, Julian, et al.
Published: (2024)
by: Wechsler, Julian, et al.
Published: (2024)
Similar Items
-
Vivo : une approche multimodale de la synthese concatenative par corpus dans le cadre d'une oeuvre audiovisuelle immersive
by: Fayet, Mateo
Published: (2024) -
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
by: Vecino, Biel Tura, et al.
Published: (2025) -
Application of Whisper in Clinical Practice: the Post-Stroke Speech Assessment during a Naming Task
by: Davudova, Milena, et al.
Published: (2025) -
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
by: Lin, Zhaofeng, et al.
Published: (2024) -
Angular Distance Distribution Loss for Audio Classification
by: Almudévar, Antonio, et al.
Published: (2024)