RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
Fuente:
arXiv
Salvato in:
| Autori principali: | Choe, Wonkyo, Ji, Yangfeng, Lin, Felix Xiaozhu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Video RWKV:Video Action Recognition Based RWKV
di: Yin, Zhuowen, et al.
Pubblicazione: (2024)
di: Yin, Zhuowen, et al.
Pubblicazione: (2024)
SelectFormer: Private and Practical Data Selection for Transformers
di: Ouyang, Xu, et al.
Pubblicazione: (2023)
di: Ouyang, Xu, et al.
Pubblicazione: (2023)
Quantitative Analysis of Deeply Quantized Tiny Neural Networks Robust to Adversarial Attacks
di: Zakariyya, Idris, et al.
Pubblicazione: (2025)
di: Zakariyya, Idris, et al.
Pubblicazione: (2025)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
di: Benazir, Afsara, et al.
Pubblicazione: (2025)
di: Benazir, Afsara, et al.
Pubblicazione: (2025)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
di: Almurshed, Osama, et al.
Pubblicazione: (2025)
di: Almurshed, Osama, et al.
Pubblicazione: (2025)
Application Research On Real-Time Perception Of Device Performance Status
di: Wang, Zhe, et al.
Pubblicazione: (2024)
di: Wang, Zhe, et al.
Pubblicazione: (2024)
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices
di: Xu, Mingbin, et al.
Pubblicazione: (2023)
di: Xu, Mingbin, et al.
Pubblicazione: (2023)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
di: Duan, Shukai, et al.
Pubblicazione: (2024)
di: Duan, Shukai, et al.
Pubblicazione: (2024)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
State Tuning: State-based Test-Time Scaling on RWKV-7
di: Xiao, Liu, et al.
Pubblicazione: (2025)
di: Xiao, Liu, et al.
Pubblicazione: (2025)
Belief-State RWKV for Reinforcement Learning under Partial Observability
di: Xiao, Liu
Pubblicazione: (2026)
di: Xiao, Liu
Pubblicazione: (2026)
Profiling Apple Silicon Performance for ML Training
di: Feng, Dahua, et al.
Pubblicazione: (2025)
di: Feng, Dahua, et al.
Pubblicazione: (2025)
Optimizing Methane Detection On Board Satellites: Speed, Accuracy, and Low-Power Solutions for Resource-Constrained Hardware
di: Herec, Jonáš, et al.
Pubblicazione: (2025)
di: Herec, Jonáš, et al.
Pubblicazione: (2025)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
di: Yi, Qingao, et al.
Pubblicazione: (2025)
di: Yi, Qingao, et al.
Pubblicazione: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
di: Wang, Wenxiao, et al.
Pubblicazione: (2024)
di: Wang, Wenxiao, et al.
Pubblicazione: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
di: Hossain, Abrar, et al.
Pubblicazione: (2025)
di: Hossain, Abrar, et al.
Pubblicazione: (2025)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
RWKV-TS: Beyond Traditional Recurrent Neural Network for Time Series Tasks
di: Hou, Haowen, et al.
Pubblicazione: (2024)
di: Hou, Haowen, et al.
Pubblicazione: (2024)
Proto: A Guided Journey through Modern OS Construction
di: Choe, Wonkyo, et al.
Pubblicazione: (2025)
di: Choe, Wonkyo, et al.
Pubblicazione: (2025)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
di: Peng, Bo, et al.
Pubblicazione: (2025)
di: Peng, Bo, et al.
Pubblicazione: (2025)
Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
di: Xiao, Liu, et al.
Pubblicazione: (2025)
di: Xiao, Liu, et al.
Pubblicazione: (2025)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
di: Li, Jianhui, et al.
Pubblicazione: (2023)
di: Li, Jianhui, et al.
Pubblicazione: (2023)
NeRFlex: Resource-aware Real-time High-quality Rendering of Complex Scenes on Mobile Devices
di: Wang, Zhe, et al.
Pubblicazione: (2025)
di: Wang, Zhe, et al.
Pubblicazione: (2025)
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
di: Bergman, Shai, et al.
Pubblicazione: (2025)
di: Bergman, Shai, et al.
Pubblicazione: (2025)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
di: Xu, Chen, et al.
Pubblicazione: (2025)
di: Xu, Chen, et al.
Pubblicazione: (2025)
Vision-QRWKV: Exploring Quantum-Enhanced RWKV Models for Image Classification
di: Chen, Chi-Sheng
Pubblicazione: (2025)
di: Chen, Chi-Sheng
Pubblicazione: (2025)
CRAM: Large-scale Video Continual Learning with Bootstrapped Compression
di: Mall, Shivani, et al.
Pubblicazione: (2025)
di: Mall, Shivani, et al.
Pubblicazione: (2025)
A Study on Inference Latency for Vision Transformers on Mobile Devices
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
Rapid Augmentations for Time Series (RATS): A High-Performance Library for Time Series Augmentation
di: Skaf, Wadie, et al.
Pubblicazione: (2026)
di: Skaf, Wadie, et al.
Pubblicazione: (2026)
A Scalable k-Medoids Clustering via Whale Optimization Algorithm
di: Chenan, Huang, et al.
Pubblicazione: (2024)
di: Chenan, Huang, et al.
Pubblicazione: (2024)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
di: Lazcano, Raquel, et al.
Pubblicazione: (2024)
di: Lazcano, Raquel, et al.
Pubblicazione: (2024)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
di: Zhang, Yijia, et al.
Pubblicazione: (2024)
di: Zhang, Yijia, et al.
Pubblicazione: (2024)
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
di: Wang, Can, et al.
Pubblicazione: (2024)
di: Wang, Can, et al.
Pubblicazione: (2024)
SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator
di: Chen, Yukai, et al.
Pubblicazione: (2024)
di: Chen, Yukai, et al.
Pubblicazione: (2024)
Tabular and Deep Reinforcement Learning for Gittins Index
di: Dhankhar, Harshit, et al.
Pubblicazione: (2024)
di: Dhankhar, Harshit, et al.
Pubblicazione: (2024)
Towards an Integrated Performance Framework for Fire Science and Management Workflows
di: Ahmed, H., et al.
Pubblicazione: (2024)
di: Ahmed, H., et al.
Pubblicazione: (2024)
Forecasting GPU Performance for Deep Learning Training and Inference
di: Lee, Seonho, et al.
Pubblicazione: (2024)
di: Lee, Seonho, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Video RWKV:Video Action Recognition Based RWKV
di: Yin, Zhuowen, et al.
Pubblicazione: (2024) -
SelectFormer: Private and Practical Data Selection for Transformers
di: Ouyang, Xu, et al.
Pubblicazione: (2023) -
Quantitative Analysis of Deeply Quantized Tiny Neural Networks Robust to Adversarial Attacks
di: Zakariyya, Idris, et al.
Pubblicazione: (2025) -
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
di: Benazir, Afsara, et al.
Pubblicazione: (2025) -
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
di: Almurshed, Osama, et al.
Pubblicazione: (2025)