BickGraphing: Web-Based Application for Visual Inspection of Audio Recordings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seow, Kayley, Arovas, Alexander, Steinmetz, Grace, Bick, Emily |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Differentiable Black-box and Gray-box Modeling of Nonlinear Audio Effects
von: Comunità, Marco, et al.
Veröffentlicht: (2025)
von: Comunità, Marco, et al.
Veröffentlicht: (2025)
ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization
von: Steinmetz, Christian J., et al.
Veröffentlicht: (2024)
von: Steinmetz, Christian J., et al.
Veröffentlicht: (2024)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2024)
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2024)
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
von: Aziz, Shiran, et al.
Veröffentlicht: (2024)
von: Aziz, Shiran, et al.
Veröffentlicht: (2024)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
von: Hussain, Tassadaq, et al.
Veröffentlicht: (2024)
von: Hussain, Tassadaq, et al.
Veröffentlicht: (2024)
Review of MEMS Speakers for Audio Applications
von: Wittek, Nils, et al.
Veröffentlicht: (2025)
von: Wittek, Nils, et al.
Veröffentlicht: (2025)
Online Audio-Visual Autoregressive Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
Diffusion-Based Audio Inpainting
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
Quantifying Spatial Audio Quality Impairment
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
von: Sharma, Sarthak, et al.
Veröffentlicht: (2024)
von: Sharma, Sarthak, et al.
Veröffentlicht: (2024)
Low-Complexity Neural Wind Noise Reduction for Audio Recordings
von: Eftekhari, Hesam, et al.
Veröffentlicht: (2025)
von: Eftekhari, Hesam, et al.
Veröffentlicht: (2025)
AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
von: Jia, Yuhang, et al.
Veröffentlicht: (2024)
von: Jia, Yuhang, et al.
Veröffentlicht: (2024)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
Attention-Based Audio Embeddings for Query-by-Example
von: Singh, Anup, et al.
Veröffentlicht: (2022)
von: Singh, Anup, et al.
Veröffentlicht: (2022)
Online Single-Channel Audio-Based Sound Speed Estimation for Robust Multi-Channel Audio Control
von: Fuglsig, Andreas Jonas, et al.
Veröffentlicht: (2026)
von: Fuglsig, Andreas Jonas, et al.
Veröffentlicht: (2026)
ASPED: An Audio Dataset for Detecting Pedestrians
von: Seshadri, Pavan, et al.
Veröffentlicht: (2023)
von: Seshadri, Pavan, et al.
Veröffentlicht: (2023)
Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and Machine Learning Approaches
von: Pierotti, Francesco, et al.
Veröffentlicht: (2025)
von: Pierotti, Francesco, et al.
Veröffentlicht: (2025)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
An Adaptive CMSA for Solving the Longest Filled Common Subsequence Problem with an Application in Audio Querying
von: Djukanovic, Marko, et al.
Veröffentlicht: (2025)
von: Djukanovic, Marko, et al.
Veröffentlicht: (2025)
Audio Effect Estimation with DNN-Based Prediction and Search Algorithm
von: Okita, Youichi, et al.
Veröffentlicht: (2026)
von: Okita, Youichi, et al.
Veröffentlicht: (2026)
Streaming Audio Transformers for Online Audio Tagging
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Pengi: An Audio Language Model for Audio Tasks
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Differentiable Black-box and Gray-box Modeling of Nonlinear Audio Effects
von: Comunità, Marco, et al.
Veröffentlicht: (2025) -
ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization
von: Steinmetz, Christian J., et al.
Veröffentlicht: (2024) -
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2024) -
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
von: Aziz, Shiran, et al.
Veröffentlicht: (2024) -
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
von: Hussain, Tassadaq, et al.
Veröffentlicht: (2024)