Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Yasuda, Masahiro, Harada, Noboru, Ohishi, Yasunori, Saito, Shoichiro, Nakayama, Akira, Ono, Nobutaka |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on self-motioning human
di: Yasuda, Masahiro, et al.
Pubblicazione: (2024)
di: Yasuda, Masahiro, et al.
Pubblicazione: (2024)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
di: Niizumi, Daisuke, et al.
Pubblicazione: (2024)
di: Niizumi, Daisuke, et al.
Pubblicazione: (2024)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
di: Niizumi, Daisuke, et al.
Pubblicazione: (2026)
di: Niizumi, Daisuke, et al.
Pubblicazione: (2026)
Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2025)
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2025)
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
di: Takeuchi, Daiki, et al.
Pubblicazione: (2025)
di: Takeuchi, Daiki, et al.
Pubblicazione: (2025)
Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
di: Niizumi, Daisuke, et al.
Pubblicazione: (2025)
di: Niizumi, Daisuke, et al.
Pubblicazione: (2025)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
di: Niizumi, Daisuke, et al.
Pubblicazione: (2024)
di: Niizumi, Daisuke, et al.
Pubblicazione: (2024)
Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
di: Yasuda, Masahiro, et al.
Pubblicazione: (2025)
di: Yasuda, Masahiro, et al.
Pubblicazione: (2025)
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
di: Niizumi, Daisuke, et al.
Pubblicazione: (2025)
di: Niizumi, Daisuke, et al.
Pubblicazione: (2025)
Towards Pre-training an Effective Respiratory Audio Foundation Model
di: Niizumi, Daisuke, et al.
Pubblicazione: (2025)
di: Niizumi, Daisuke, et al.
Pubblicazione: (2025)
RenderBox: Expressive Performance Rendering with Text Control
di: Zhang, Huan, et al.
Pubblicazione: (2025)
di: Zhang, Huan, et al.
Pubblicazione: (2025)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
di: Tsubaki, Shunsuke, et al.
Pubblicazione: (2024)
di: Tsubaki, Shunsuke, et al.
Pubblicazione: (2024)
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
di: Niizumi, Daisuke, et al.
Pubblicazione: (2024)
di: Niizumi, Daisuke, et al.
Pubblicazione: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
di: Cui, Yang, et al.
Pubblicazione: (2025)
di: Cui, Yang, et al.
Pubblicazione: (2025)
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
di: Ragano, Alessandro, et al.
Pubblicazione: (2024)
di: Ragano, Alessandro, et al.
Pubblicazione: (2024)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
Trusted Fake Audio Detection Based on Dirichlet Distribution
di: Ding, Chi, et al.
Pubblicazione: (2025)
di: Ding, Chi, et al.
Pubblicazione: (2025)
Amanous: Distribution-Switching for Superhuman Piano Density on Disklavier
di: Bae, Joonhyung
Pubblicazione: (2026)
di: Bae, Joonhyung
Pubblicazione: (2026)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
di: Kim, Haven, et al.
Pubblicazione: (2025)
di: Kim, Haven, et al.
Pubblicazione: (2025)
Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2026)
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2026)
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
di: Huang, Lian, et al.
Pubblicazione: (2024)
di: Huang, Lian, et al.
Pubblicazione: (2024)
Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2026)
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2026)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
di: Su, Fei, et al.
Pubblicazione: (2026)
di: Su, Fei, et al.
Pubblicazione: (2026)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
di: Deng, Jiajun, et al.
Pubblicazione: (2025)
di: Deng, Jiajun, et al.
Pubblicazione: (2025)
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
di: Yang, Chenyu, et al.
Pubblicazione: (2025)
di: Yang, Chenyu, et al.
Pubblicazione: (2025)
Target Speech Diarization with Multimodal Prompts
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
Iola Walker: A Mobile Footfall Detection System for Music Composition
di: James, William B.
Pubblicazione: (2025)
di: James, William B.
Pubblicazione: (2025)
LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
di: Zhang, Huan, et al.
Pubblicazione: (2024)
di: Zhang, Huan, et al.
Pubblicazione: (2024)
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
di: Wu, Shilong
Pubblicazione: (2025)
di: Wu, Shilong
Pubblicazione: (2025)
MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
Dialogue Understandability: Why are we streaming movies with subtitles?
di: Martinez, Helard Becerra, et al.
Pubblicazione: (2024)
di: Martinez, Helard Becerra, et al.
Pubblicazione: (2024)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
di: Morrone, Giovanni, et al.
Pubblicazione: (2024)
di: Morrone, Giovanni, et al.
Pubblicazione: (2024)
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
di: Bhattacharya, Aneesh, et al.
Pubblicazione: (2023)
di: Bhattacharya, Aneesh, et al.
Pubblicazione: (2023)
VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect
di: Fei, Qingyuan, et al.
Pubblicazione: (2025)
di: Fei, Qingyuan, et al.
Pubblicazione: (2025)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
di: Bai, Yatong, et al.
Pubblicazione: (2023)
di: Bai, Yatong, et al.
Pubblicazione: (2023)
Double Mixture: Towards Continual Event Detection from Speech
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
Microphone Conversion: Mitigating Device Variability in Sound Event Classification
di: Ryu, Myeonghoon, et al.
Pubblicazione: (2024)
di: Ryu, Myeonghoon, et al.
Pubblicazione: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
Documenti analoghi
-
6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on self-motioning human
di: Yasuda, Masahiro, et al.
Pubblicazione: (2024) -
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
di: Niizumi, Daisuke, et al.
Pubblicazione: (2024) -
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
di: Niizumi, Daisuke, et al.
Pubblicazione: (2026) -
Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2025) -
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
di: Takeuchi, Daiki, et al.
Pubblicazione: (2025)