TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Xinyu, Zhao, Ziyu, Bai, Ke, Meng, Silin, Shen, Dongming, Chang, Xiao-Wen, HE, Yixuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Back to Basics: Revisiting ASR in the Age of Voice Agents
por: Tay, Geeyang, et al.
Publicado: (2026)
por: Tay, Geeyang, et al.
Publicado: (2026)
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
por: Li, Shaokai, et al.
Publicado: (2024)
por: Li, Shaokai, et al.
Publicado: (2024)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
por: Xu, Siyuan, et al.
Publicado: (2026)
por: Xu, Siyuan, et al.
Publicado: (2026)
CARAT: Contrastive Feature Reconstruction and Aggregation for Multi-Modal Multi-Label Emotion Recognition
por: Peng, Cheng, et al.
Publicado: (2023)
por: Peng, Cheng, et al.
Publicado: (2023)
AMD: Autoregressive Motion Diffusion
por: Han, Bo, et al.
Publicado: (2023)
por: Han, Bo, et al.
Publicado: (2023)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
por: Nishida, Naoto, et al.
Publicado: (2025)
por: Nishida, Naoto, et al.
Publicado: (2025)
Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition
por: Zhuang, Yan, et al.
Publicado: (2026)
por: Zhuang, Yan, et al.
Publicado: (2026)
Fully Automatic Content-Aware Tiling Pipeline for Pathology Whole Slide Images
por: Jabar, Falah, et al.
Publicado: (2024)
por: Jabar, Falah, et al.
Publicado: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
por: Liu, Qianhui, et al.
Publicado: (2024)
por: Liu, Qianhui, et al.
Publicado: (2024)
DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems
por: Guo, Qi, et al.
Publicado: (2026)
por: Guo, Qi, et al.
Publicado: (2026)
GTLR-GS: Geometry-Texture Aware LiDAR-Regularized 3D Gaussian Splatting for Realistic Scene Reconstruction
por: Fang, Yan, et al.
Publicado: (2026)
por: Fang, Yan, et al.
Publicado: (2026)
SRA: Semantic Relation-Aware Flowchart Question Answering
por: Li, Xinyu, et al.
Publicado: (2026)
por: Li, Xinyu, et al.
Publicado: (2026)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
por: Pan, Zhaoyan, et al.
Publicado: (2026)
por: Pan, Zhaoyan, et al.
Publicado: (2026)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
por: Zhang, Han, et al.
Publicado: (2025)
por: Zhang, Han, et al.
Publicado: (2025)
Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation
por: Phan, Nghia, et al.
Publicado: (2026)
por: Phan, Nghia, et al.
Publicado: (2026)
HADUA: Hierarchical Attention and Dynamic Uniform Alignment for Robust Cross-Subject Emotion Recognition
por: Tang, Jiahao, et al.
Publicado: (2026)
por: Tang, Jiahao, et al.
Publicado: (2026)
Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation
por: Yi, Zijian, et al.
Publicado: (2024)
por: Yi, Zijian, et al.
Publicado: (2024)
Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis
por: Dong, Guangyuan, et al.
Publicado: (2026)
por: Dong, Guangyuan, et al.
Publicado: (2026)
High-Fidelity 3D Gaussian Human Reconstruction via Region-Aware Initialization and Geometric Priors
por: Liu, Yang, et al.
Publicado: (2026)
por: Liu, Yang, et al.
Publicado: (2026)
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
por: Gong, Ziyu, et al.
Publicado: (2024)
por: Gong, Ziyu, et al.
Publicado: (2024)
Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch Inbetweening
por: Shen, Jiaming, et al.
Publicado: (2023)
por: Shen, Jiaming, et al.
Publicado: (2023)
LLM2Manim: Pedagogy-Aware AI Generation of STEM Animations
por: Joshi, Aastha, et al.
Publicado: (2026)
por: Joshi, Aastha, et al.
Publicado: (2026)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
por: Su, Fei, et al.
Publicado: (2026)
por: Su, Fei, et al.
Publicado: (2026)
SpeechEE: A Novel Benchmark for Speech Event Extraction
por: Wang, Bin, et al.
Publicado: (2024)
por: Wang, Bin, et al.
Publicado: (2024)
Towards Real-World Stickers Use: A New Dataset for Multi-Tag Sticker Recognition
por: Wang, Bingbing, et al.
Publicado: (2024)
por: Wang, Bingbing, et al.
Publicado: (2024)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
por: Fernandez-Lopez, Adriana, et al.
Publicado: (2024)
por: Fernandez-Lopez, Adriana, et al.
Publicado: (2024)
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
por: Wang, Jiadong, et al.
Publicado: (2026)
por: Wang, Jiadong, et al.
Publicado: (2026)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
por: Wang, Yuyue, et al.
Publicado: (2025)
por: Wang, Yuyue, et al.
Publicado: (2025)
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
por: Huynh, Ngoc Dung, et al.
Publicado: (2025)
por: Huynh, Ngoc Dung, et al.
Publicado: (2025)
Rethinking Bjøntegaard Delta for Compression Efficiency Evaluation: Are We Calculating It Precisely and Reliably?
por: Hang, Xinyu, et al.
Publicado: (2024)
por: Hang, Xinyu, et al.
Publicado: (2024)
Private Speech Classification without Collapse: Stabilized DP Training and Offline Distillation
por: Wen, Yadi, et al.
Publicado: (2026)
por: Wen, Yadi, et al.
Publicado: (2026)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
por: Cao, Jiajun, et al.
Publicado: (2025)
por: Cao, Jiajun, et al.
Publicado: (2025)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
por: Chen, Chen, et al.
Publicado: (2024)
por: Chen, Chen, et al.
Publicado: (2024)
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
por: Bai, Hayes, et al.
Publicado: (2026)
por: Bai, Hayes, et al.
Publicado: (2026)
Predictability-Aware Motion Prediction for Edge XR via High-Order Error-State Kalman Filtering
por: Zhong, Ziyu, et al.
Publicado: (2025)
por: Zhong, Ziyu, et al.
Publicado: (2025)
Towards Structure-aware Model for Multi-modal Knowledge Graph Completion
por: Li, Linyu, et al.
Publicado: (2025)
por: Li, Linyu, et al.
Publicado: (2025)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
por: Zhou, Qianrui, et al.
Publicado: (2023)
por: Zhou, Qianrui, et al.
Publicado: (2023)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
por: Wu, Linzhi, et al.
Publicado: (2026)
por: Wu, Linzhi, et al.
Publicado: (2026)
Voxel-GS: Quantized Scaffold Gaussian Splatting Compression with Run-Length Coding
por: Fu, Chunyang, et al.
Publicado: (2025)
por: Fu, Chunyang, et al.
Publicado: (2025)
Ejemplares similares
-
Back to Basics: Revisiting ASR in the Age of Voice Agents
por: Tay, Geeyang, et al.
Publicado: (2026) -
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
por: Li, Shaokai, et al.
Publicado: (2024) -
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
por: Xu, Siyuan, et al.
Publicado: (2026) -
CARAT: Contrastive Feature Reconstruction and Aggregation for Multi-Modal Multi-Label Emotion Recognition
por: Peng, Cheng, et al.
Publicado: (2023) -
AMD: Autoregressive Motion Diffusion
por: Han, Bo, et al.
Publicado: (2023)