Dialogue Understandability: Why are we streaming movies with subtitles?
Fuente:
arXiv
Salvato in:
| Autori principali: | Martinez, Helard Becerra, Ragano, Alessandro, Debnath, Diptasree, Ullah, Asad, Lucas, Crisron Rudolf, Walsh, Martin, Hines, Andrew |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
di: Ragano, Alessandro, et al.
Pubblicazione: (2024)
di: Ragano, Alessandro, et al.
Pubblicazione: (2024)
Reduce, Reuse, Recycle: Is Perturbed Data better than Other Language augmentation for Low Resource Self-Supervised Speech Models
di: Ullah, Asad, et al.
Pubblicazione: (2023)
di: Ullah, Asad, et al.
Pubblicazione: (2023)
NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
di: Ragano, Alessandro, et al.
Pubblicazione: (2023)
di: Ragano, Alessandro, et al.
Pubblicazione: (2023)
SCOREQ: Speech Quality Assessment with Contrastive Regression
di: Ragano, Alessandro, et al.
Pubblicazione: (2024)
di: Ragano, Alessandro, et al.
Pubblicazione: (2024)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
di: Zhao, Qihao, et al.
Pubblicazione: (2026)
di: Zhao, Qihao, et al.
Pubblicazione: (2026)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
di: Liu, Shansong, et al.
Pubblicazione: (2023)
di: Liu, Shansong, et al.
Pubblicazione: (2023)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
di: Liu, Shansong, et al.
Pubblicazione: (2024)
di: Liu, Shansong, et al.
Pubblicazione: (2024)
StereoFoley: Object-Aware Stereo Audio Generation from Video
di: Karchkhadze, Tornike, et al.
Pubblicazione: (2025)
di: Karchkhadze, Tornike, et al.
Pubblicazione: (2025)
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
di: Panah, Davoud Shariat, et al.
Pubblicazione: (2025)
di: Panah, Davoud Shariat, et al.
Pubblicazione: (2025)
Binamix -- A Python Library for Generating Binaural Audio Datasets
di: Barry, Dan, et al.
Pubblicazione: (2025)
di: Barry, Dan, et al.
Pubblicazione: (2025)
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
di: Yang, Chenyu, et al.
Pubblicazione: (2025)
di: Yang, Chenyu, et al.
Pubblicazione: (2025)
Target Speech Diarization with Multimodal Prompts
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
Iola Walker: A Mobile Footfall Detection System for Music Composition
di: James, William B.
Pubblicazione: (2025)
di: James, William B.
Pubblicazione: (2025)
LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
di: Zhang, Huan, et al.
Pubblicazione: (2024)
di: Zhang, Huan, et al.
Pubblicazione: (2024)
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
di: Wu, Shilong
Pubblicazione: (2025)
di: Wu, Shilong
Pubblicazione: (2025)
MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
RenderBox: Expressive Performance Rendering with Text Control
di: Zhang, Huan, et al.
Pubblicazione: (2025)
di: Zhang, Huan, et al.
Pubblicazione: (2025)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
di: Morrone, Giovanni, et al.
Pubblicazione: (2024)
di: Morrone, Giovanni, et al.
Pubblicazione: (2024)
Can Large Language Models Predict Audio Effects Parameters from Natural Language?
di: Doh, Seungheon, et al.
Pubblicazione: (2025)
di: Doh, Seungheon, et al.
Pubblicazione: (2025)
A Survey of Foundation Models for Music Understanding
di: Li, Wenjun, et al.
Pubblicazione: (2024)
di: Li, Wenjun, et al.
Pubblicazione: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
di: Kim, Haven, et al.
Pubblicazione: (2025)
di: Kim, Haven, et al.
Pubblicazione: (2025)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
di: Gu, Ke, et al.
Pubblicazione: (2025)
di: Gu, Ke, et al.
Pubblicazione: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
di: Cui, Yang, et al.
Pubblicazione: (2025)
di: Cui, Yang, et al.
Pubblicazione: (2025)
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
di: Guo, Ziang, et al.
Pubblicazione: (2026)
di: Guo, Ziang, et al.
Pubblicazione: (2026)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
Building Audio-Visual Digital Twins with Smartphones
di: Lan, Zitong, et al.
Pubblicazione: (2025)
di: Lan, Zitong, et al.
Pubblicazione: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
di: Gong, Jingyao
Pubblicazione: (2026)
di: Gong, Jingyao
Pubblicazione: (2026)
Dance2MIDI: Dance-driven multi-instruments music generation
di: Han, Bo, et al.
Pubblicazione: (2023)
di: Han, Bo, et al.
Pubblicazione: (2023)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
di: Yu, Fan, et al.
Pubblicazione: (2024)
di: Yu, Fan, et al.
Pubblicazione: (2024)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
di: Salvi, Davide, et al.
Pubblicazione: (2024)
di: Salvi, Davide, et al.
Pubblicazione: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
di: Kwak, Doyeop, et al.
Pubblicazione: (2026)
di: Kwak, Doyeop, et al.
Pubblicazione: (2026)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
di: Li, Yupei, et al.
Pubblicazione: (2024)
di: Li, Yupei, et al.
Pubblicazione: (2024)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
Intelligent Text-Conditioned Music Generation
di: Xie, Zhouyao, et al.
Pubblicazione: (2024)
di: Xie, Zhouyao, et al.
Pubblicazione: (2024)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
di: Li, Xiaolou, et al.
Pubblicazione: (2024)
di: Li, Xiaolou, et al.
Pubblicazione: (2024)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
di: Ren, Yong, et al.
Pubblicazione: (2024)
di: Ren, Yong, et al.
Pubblicazione: (2024)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
di: Ragano, Alessandro, et al.
Pubblicazione: (2024) -
Reduce, Reuse, Recycle: Is Perturbed Data better than Other Language augmentation for Low Resource Self-Supervised Speech Models
di: Ullah, Asad, et al.
Pubblicazione: (2023) -
NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
di: Ragano, Alessandro, et al.
Pubblicazione: (2023) -
SCOREQ: Speech Quality Assessment with Contrastive Regression
di: Ragano, Alessandro, et al.
Pubblicazione: (2024) -
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
di: Zhao, Qihao, et al.
Pubblicazione: (2026)