A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pham, Lam, Lam, Phat, Nguyen, Tin, Tang, Hieu, Schindler, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
von: Nguyen, Tin, et al.
Veröffentlicht: (2024)
von: Nguyen, Tin, et al.
Veröffentlicht: (2024)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
DIN-CTS: Low-Complexity Depthwise-Inception Neural Network with Contrastive Training Strategy for Deepfake Speech Detection
von: Pham, Lam, et al.
Veröffentlicht: (2025)
von: Pham, Lam, et al.
Veröffentlicht: (2025)
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
von: Lam, Phat, et al.
Veröffentlicht: (2024)
von: Lam, Phat, et al.
Veröffentlicht: (2024)
Aud-Sur: An Audio Analyzer Assistant for Audio Surveillance Applications
von: Lam, Phat, et al.
Veröffentlicht: (2025)
von: Lam, Phat, et al.
Veröffentlicht: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
Autonomous Soundscape Augmentation with Multimodal Fusion of Visual and Participant-linked Inputs
von: Ooi, Kenneth, et al.
Veröffentlicht: (2023)
von: Ooi, Kenneth, et al.
Veröffentlicht: (2023)
Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and Machine Learning Approaches
von: Pierotti, Francesco, et al.
Veröffentlicht: (2025)
von: Pierotti, Francesco, et al.
Veröffentlicht: (2025)
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
von: Wang, Pingjie, et al.
Veröffentlicht: (2024)
von: Wang, Pingjie, et al.
Veröffentlicht: (2024)
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
WeDefense: A Toolkit to Defend Against Fake Audio
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
AudioGenie-Reasoner: A Training-Free Multi-Agent Framework for Coarse-to-Fine Audio Deep Reasoning
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
von: Chen, Gehui, et al.
Veröffentlicht: (2024)
von: Chen, Gehui, et al.
Veröffentlicht: (2024)
Deep Learning for Personalized Binaural Audio Reproduction
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
Video-to-Audio Generation with Fine-grained Temporal Semantics
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Audio Avatar Fingerprinting: An Approach for Authorized Use of Voice Cloning in the Era of Synthetic Audio
von: Gerstner, Candice R.
Veröffentlicht: (2026)
von: Gerstner, Candice R.
Veröffentlicht: (2026)
PromptSep: Generative Audio Separation via Multimodal Prompting
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
von: Kheir, Yassine El, et al.
Veröffentlicht: (2026)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2026)
Quantifying Spatial Audio Quality Impairment
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
Room Impulse Responses help attackers to evade Deep Fake Detection
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
Generalized Fake Audio Detection via Deep Stable Learning
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
The ICASSP 2024 Audio Deep Packet Loss Concealment Challenge
von: Diener, Lorenz, et al.
Veröffentlicht: (2024)
von: Diener, Lorenz, et al.
Veröffentlicht: (2024)
Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals
von: Chen, Zhang, et al.
Veröffentlicht: (2026)
von: Chen, Zhang, et al.
Veröffentlicht: (2026)
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
von: Kang, Wei, et al.
Veröffentlicht: (2023)
von: Kang, Wei, et al.
Veröffentlicht: (2023)
Continuous Learning of Transformer-based Audio Deepfake Detection
von: Le, Tuan Duy Nguyen, et al.
Veröffentlicht: (2024)
von: Le, Tuan Duy Nguyen, et al.
Veröffentlicht: (2024)
Enhancing Generalization in Audio Deepfake Detection: A Neural Collapse based Sampling and Training Approach
von: Yousif, Mohammed, et al.
Veröffentlicht: (2024)
von: Yousif, Mohammed, et al.
Veröffentlicht: (2024)
Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach
von: Abeßer, Jakob, et al.
Veröffentlicht: (2025)
von: Abeßer, Jakob, et al.
Veröffentlicht: (2025)
Automatic acoustic detection of birds through deep learning: the first Bird Audio Detection challenge
von: Stowell, Dan, et al.
Veröffentlicht: (2018)
von: Stowell, Dan, et al.
Veröffentlicht: (2018)
Are audio DeepFake detection models polyglots?
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
A Real-Time Platform for Portable and Scalable Active Noise Mitigation for Construction Machinery
von: Gan, Woon-Seng, et al.
Veröffentlicht: (2024)
von: Gan, Woon-Seng, et al.
Veröffentlicht: (2024)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
Advancing Continual Learning for Robust Deepfake Audio Classification
von: Dong, Feiyi, et al.
Veröffentlicht: (2024)
von: Dong, Feiyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
von: Nguyen, Tin, et al.
Veröffentlicht: (2024) -
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
von: Pham, Lam, et al.
Veröffentlicht: (2024) -
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
von: Pham, Lam, et al.
Veröffentlicht: (2024) -
DIN-CTS: Low-Complexity Depthwise-Inception Neural Network with Contrastive Training Strategy for Deepfake Speech Detection
von: Pham, Lam, et al.
Veröffentlicht: (2025) -
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
von: Lam, Phat, et al.
Veröffentlicht: (2024)