From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Karbalaie, Abdolamir, Seoane, Fernando, Abtahi, Farhad |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
par: Abtahi, Farhad, et autres
Publié: (2026)
par: Abtahi, Farhad, et autres
Publié: (2026)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
par: Hori, Takaaki, et autres
Publié: (2025)
par: Hori, Takaaki, et autres
Publié: (2025)
Quantization for OpenAI's Whisper Models: A Comparative Analysis
par: Andreyev, Allison
Publié: (2025)
par: Andreyev, Allison
Publié: (2025)
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
par: Lasbordes, Maxence, et autres
Publié: (2025)
par: Lasbordes, Maxence, et autres
Publié: (2025)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
par: Papyan, Narek, et autres
Publié: (2024)
par: Papyan, Narek, et autres
Publié: (2024)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
par: Salvi, Giampiero
Publié: (2024)
par: Salvi, Giampiero
Publié: (2024)
Audio-based Kinship Verification Using Age Domain Conversion
par: Sun, Qiyang, et autres
Publié: (2024)
par: Sun, Qiyang, et autres
Publié: (2024)
Graph Connectionist Temporal Classification for Phoneme Recognition
par: Grafé, Henry, et autres
Publié: (2025)
par: Grafé, Henry, et autres
Publié: (2025)
Passive Underwater Acoustic Signal Separation based on Feature Decoupling Dual-path Network
par: Liu, Yucheng, et autres
Publié: (2025)
par: Liu, Yucheng, et autres
Publié: (2025)
Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life
par: Batliner, Anton, et autres
Publié: (2025)
par: Batliner, Anton, et autres
Publié: (2025)
STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
par: Firc, Anton, et autres
Publié: (2025)
par: Firc, Anton, et autres
Publié: (2025)
Detection and Classification of Cetacean Echolocation Clicks using Image-based Object Detection Methods applied to Advanced Wavelet-based Transformations
par: Hauer, Christopher
Publié: (2026)
par: Hauer, Christopher
Publié: (2026)
Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
par: Ethiraj, Vignesh, et autres
Publié: (2025)
par: Ethiraj, Vignesh, et autres
Publié: (2025)
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
par: Zhang, Haoyang, et autres
Publié: (2025)
par: Zhang, Haoyang, et autres
Publié: (2025)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
par: Cheripally, Sowmya
Publié: (2024)
par: Cheripally, Sowmya
Publié: (2024)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
par: Patel, Urjitkumar, et autres
Publié: (2025)
par: Patel, Urjitkumar, et autres
Publié: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
par: Viveiros, André G., et autres
Publié: (2025)
par: Viveiros, André G., et autres
Publié: (2025)
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
par: Chen, Zhehuai, et autres
Publié: (2024)
par: Chen, Zhehuai, et autres
Publié: (2024)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
par: McKinnon, Max, et autres
Publié: (2026)
par: McKinnon, Max, et autres
Publié: (2026)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
par: Jain, Sarthak, et autres
Publié: (2024)
par: Jain, Sarthak, et autres
Publié: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Prevailing Research Areas for Music AI in the Era of Foundation Models
par: Wei, Megan, et autres
Publié: (2024)
par: Wei, Megan, et autres
Publié: (2024)
Noise-Robust Keyword Spotting through Self-supervised Pretraining
par: Mørk, Jacob, et autres
Publié: (2024)
par: Mørk, Jacob, et autres
Publié: (2024)
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
par: Bovbjerg, Holger Severin, et autres
Publié: (2023)
par: Bovbjerg, Holger Severin, et autres
Publié: (2023)
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
par: Bovbjerg, Holger Severin, et autres
Publié: (2025)
par: Bovbjerg, Holger Severin, et autres
Publié: (2025)
Monaural Multi-Speaker Speech Separation Using Efficient Transformer Model
par: Rijal, S., et autres
Publié: (2023)
par: Rijal, S., et autres
Publié: (2023)
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
par: Bovbjerg, Holger Severin, et autres
Publié: (2025)
par: Bovbjerg, Holger Severin, et autres
Publié: (2025)
Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and Localization
par: Wu, Junyan, et autres
Publié: (2024)
par: Wu, Junyan, et autres
Publié: (2024)
SARA: Stress Test Reasoning in Audio Deepfake Detection
par: Nguyen, Binh, et autres
Publié: (2026)
par: Nguyen, Binh, et autres
Publié: (2026)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
par: Maben, Leander Melroy, et autres
Publié: (2025)
par: Maben, Leander Melroy, et autres
Publié: (2025)
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
par: Chopra, Anuradha, et autres
Publié: (2025)
par: Chopra, Anuradha, et autres
Publié: (2025)
Cross-modal Cognitive Consensus guided Audio-Visual Segmentation
par: Shi, Zhaofeng, et autres
Publié: (2023)
par: Shi, Zhaofeng, et autres
Publié: (2023)
NAAQA: A Neural Architecture for Acoustic Question Answering
par: Abdelnour, Jerome, et autres
Publié: (2021)
par: Abdelnour, Jerome, et autres
Publié: (2021)
Hybrid ASR for Resource-Constrained Robots: HMM - Deep Learning Fusion
par: Ranjan, Anshul, et autres
Publié: (2023)
par: Ranjan, Anshul, et autres
Publié: (2023)
Understanding the Algorithm Behind Audio Key Detection
par: Silva, Henrique Perez G.
Publié: (2025)
par: Silva, Henrique Perez G.
Publié: (2025)
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Documents similaires
-
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
par: Abtahi, Farhad, et autres
Publié: (2026) -
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
par: Hori, Takaaki, et autres
Publié: (2025) -
Quantization for OpenAI's Whisper Models: A Comparative Analysis
par: Andreyev, Allison
Publié: (2025) -
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
par: Lasbordes, Maxence, et autres
Publié: (2025) -
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
par: Papyan, Narek, et autres
Publié: (2024)