Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yang, Lai, Liangzhen, Shangguan, Yuan, Iandola, Forrest N., Ni, Zhaoheng, Chang, Ernie, Shi, Yangyang, Chandra, Vikas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Real-Time Piano Note Frequency Detection Using FPGA and FFT Core
by: Anik, Shafayet M., et al.
Published: (2025)
by: Anik, Shafayet M., et al.
Published: (2025)
A Low-Power Streaming Speech Enhancement Accelerator For Edge Devices
by: Wu, Ci-Hao, et al.
Published: (2025)
by: Wu, Ci-Hao, et al.
Published: (2025)
Acoustic Local Positioning With Encoded Emission Beacons
by: Urena, Jesus, et al.
Published: (2024)
by: Urena, Jesus, et al.
Published: (2024)
Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
by: Liang, Huihong, et al.
Published: (2025)
by: Liang, Huihong, et al.
Published: (2025)
DeltaKWS: A 65nm 36nJ/Decision Bio-inspired Temporal-Sparsity-Aware Digital Keyword Spotting IC with 0.6V Near-Threshold SRAM
by: Chen, Qinyu, et al.
Published: (2024)
by: Chen, Qinyu, et al.
Published: (2024)
A 71.2-$μ$W Speech Recognition Accelerator with Recurrent Spiking Neural Network
by: Yang, Chih-Chyau, et al.
Published: (2025)
by: Yang, Chih-Chyau, et al.
Published: (2025)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
by: Lan, Gael Le, et al.
Published: (2024)
by: Lan, Gael Le, et al.
Published: (2024)
ASAP-FE: Energy-Efficient Feature Extraction Enabling Multi-Channel Keyword Spotting on Edge Processors
by: Choi, Jongin, et al.
Published: (2025)
by: Choi, Jongin, et al.
Published: (2025)
TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
by: Lin, Baizhou, et al.
Published: (2025)
by: Lin, Baizhou, et al.
Published: (2025)
Ultra-low power on-chip learning of speech commands with phase-change memories
by: Miriyala, Venkata Pavan Kumar, et al.
Published: (2020)
by: Miriyala, Venkata Pavan Kumar, et al.
Published: (2020)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
by: Shi, Hao, et al.
Published: (2024)
by: Shi, Hao, et al.
Published: (2024)
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
by: Zhang, Qiquan, et al.
Published: (2024)
by: Zhang, Qiquan, et al.
Published: (2024)
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
by: Mei, Xinhao, et al.
Published: (2026)
by: Mei, Xinhao, et al.
Published: (2026)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
by: Liu, Haohe, et al.
Published: (2024)
by: Liu, Haohe, et al.
Published: (2024)
Learnable Pulse Accumulation for On-Device Speech Recognition: How Much Attention Do You Need?
by: Shkolnikov, Yakov Pyotr
Published: (2026)
by: Shkolnikov, Yakov Pyotr
Published: (2026)
A 14uJ/Decision Keyword Spotting Accelerator with In-SRAM-Computing and On Chip Learning for Customization
by: Chiang, Yu-Hsiang, et al.
Published: (2022)
by: Chiang, Yu-Hsiang, et al.
Published: (2022)
Reverse Attention for Lightweight Speech Enhancement on Edge Devices
by: Ojha, Shuubham, et al.
Published: (2025)
by: Ojha, Shuubham, et al.
Published: (2025)
MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition
by: Xia, Yinfeng, et al.
Published: (2025)
by: Xia, Yinfeng, et al.
Published: (2025)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
by: Zeineldeen, Mohammad, et al.
Published: (2023)
by: Zeineldeen, Mohammad, et al.
Published: (2023)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
by: Chen, Peikun, et al.
Published: (2024)
by: Chen, Peikun, et al.
Published: (2024)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
Attention-Guided Adaptation for Code-Switching Speech Recognition
by: Aditya, Bobbi, et al.
Published: (2023)
by: Aditya, Bobbi, et al.
Published: (2023)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
by: Guo, Dake, et al.
Published: (2025)
by: Guo, Dake, et al.
Published: (2025)
Do we really need Self-Attention for Streaming Automatic Speech Recognition?
by: Dkhissi, Youness, et al.
Published: (2026)
by: Dkhissi, Youness, et al.
Published: (2026)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
by: Wan, Genshun, et al.
Published: (2026)
by: Wan, Genshun, et al.
Published: (2026)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
by: Zhang, Zixing, et al.
Published: (2024)
by: Zhang, Zixing, et al.
Published: (2024)
ICASSP 2026 URGENT Speech Enhancement Challenge
by: Li, Chenda, et al.
Published: (2026)
by: Li, Chenda, et al.
Published: (2026)
Bhasha-Rupantarika: Algorithm-Hardware Co-design approach for Multilingual Neural Machine Translation
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
by: Zhang, Wangyou, et al.
Published: (2024)
by: Zhang, Wangyou, et al.
Published: (2024)
Less is More: Data Curation Matters in Scaling Speech Enhancement
by: Li, Chenda, et al.
Published: (2025)
by: Li, Chenda, et al.
Published: (2025)
FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
by: Wang, Haoxu, et al.
Published: (2025)
by: Wang, Haoxu, et al.
Published: (2025)
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
by: Costa, Federico, et al.
Published: (2024)
by: Costa, Federico, et al.
Published: (2024)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
by: Ueda, Lucas, et al.
Published: (2025)
by: Ueda, Lucas, et al.
Published: (2025)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
by: Wang, He, et al.
Published: (2024)
by: Wang, He, et al.
Published: (2024)
In-Materia Speech Recognition
by: Zolfagharinejad, Mohamadreza, et al.
Published: (2024)
by: Zolfagharinejad, Mohamadreza, et al.
Published: (2024)
Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
by: Patil, Aditya, et al.
Published: (2024)
by: Patil, Aditya, et al.
Published: (2024)
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
by: Wang, Jiahe, et al.
Published: (2025)
by: Wang, Jiahe, et al.
Published: (2025)
StreamAAD: Decoding Spatial Auditory Attention with a Streaming Architecture
by: Qiu, Zelin, et al.
Published: (2024)
by: Qiu, Zelin, et al.
Published: (2024)
Similar Items
-
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions
by: Li, Yang, et al.
Published: (2024) -
Real-Time Piano Note Frequency Detection Using FPGA and FFT Core
by: Anik, Shafayet M., et al.
Published: (2025) -
A Low-Power Streaming Speech Enhancement Accelerator For Edge Devices
by: Wu, Ci-Hao, et al.
Published: (2025) -
Acoustic Local Positioning With Encoded Emission Beacons
by: Urena, Jesus, et al.
Published: (2024) -
Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
by: Liang, Huihong, et al.
Published: (2025)