Tiny-QMoE
Fuente:
arXiv
Guardado en:
| Autores principales: | Cashman, Jack, Nie, Jiaqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Latency Based Tiling
por: Cashman, Jack
Publicado: (2025)
por: Cashman, Jack
Publicado: (2025)
KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer
por: Al-Qawlaq, Aness, et al.
Publicado: (2024)
por: Al-Qawlaq, Aness, et al.
Publicado: (2024)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
por: Jung, Victor J. B., et al.
Publicado: (2024)
por: Jung, Victor J. B., et al.
Publicado: (2024)
Optimizing LoRa for Edge Computing with TinyML Pipeline for Channel Hopping
por: Grunewald, Marla, et al.
Publicado: (2024)
por: Grunewald, Marla, et al.
Publicado: (2024)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
por: Arantes, Gabriel M., et al.
Publicado: (2025)
por: Arantes, Gabriel M., et al.
Publicado: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
por: Ivanov, Andrei, et al.
Publicado: (2025)
por: Ivanov, Andrei, et al.
Publicado: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
por: Prieto, Pablo, et al.
Publicado: (2025)
por: Prieto, Pablo, et al.
Publicado: (2025)
PixLift: Accelerating Web Browsing via AI Upscaling
por: Atinafu, Yonas, et al.
Publicado: (2025)
por: Atinafu, Yonas, et al.
Publicado: (2025)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
por: Pfister, Rolf, et al.
Publicado: (2025)
por: Pfister, Rolf, et al.
Publicado: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
por: Chai, Yuji, et al.
Publicado: (2025)
por: Chai, Yuji, et al.
Publicado: (2025)
XTC, A Research Platform for Optimizing AI Workload Operators
por: Hugo, Pompougnac, et al.
Publicado: (2025)
por: Hugo, Pompougnac, et al.
Publicado: (2025)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
por: Cai, Yanan, et al.
Publicado: (2025)
por: Cai, Yanan, et al.
Publicado: (2025)
WANDER: An Explainable Decision-Support Framework for HPC
por: Lahiry, Ankur, et al.
Publicado: (2025)
por: Lahiry, Ankur, et al.
Publicado: (2025)
Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
por: Zhao, Youpeng, et al.
Publicado: (2025)
por: Zhao, Youpeng, et al.
Publicado: (2025)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Training Transformers in Cosine Coefficient Space
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Reliability by design: quantifying and eliminating fabrication risk in LLMs. From generative to consultative AI: a comparative analysis in the legal domain and lessons for high-stakes knowledge bases
por: Dantart, Alex
Publicado: (2026)
por: Dantart, Alex
Publicado: (2026)
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2026)
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2026)
Learning, Potential, and Retention: An Approach for Evaluating Adaptive AI-Enabled Medical Devices
por: Burgon, Alexis, et al.
Publicado: (2026)
por: Burgon, Alexis, et al.
Publicado: (2026)
Faster LLM Inference using DBMS-Inspired Preemption and Cache Replacement Policies
por: Kim, Kyoungmin, et al.
Publicado: (2024)
por: Kim, Kyoungmin, et al.
Publicado: (2024)
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
por: Atinafu, Yonas, et al.
Publicado: (2026)
por: Atinafu, Yonas, et al.
Publicado: (2026)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
por: Liu, Xiaoxuan, et al.
Publicado: (2024)
por: Liu, Xiaoxuan, et al.
Publicado: (2024)
Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage
por: Ngabonziza, Bernard, et al.
Publicado: (2026)
por: Ngabonziza, Bernard, et al.
Publicado: (2026)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
por: Zhao, Youpeng, et al.
Publicado: (2024)
por: Zhao, Youpeng, et al.
Publicado: (2024)
DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning Workloads
por: Zhao, Qidong, et al.
Publicado: (2024)
por: Zhao, Qidong, et al.
Publicado: (2024)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
por: Liu, Yi
Publicado: (2026)
por: Liu, Yi
Publicado: (2026)
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
por: Ponzina, Flavio, et al.
Publicado: (2024)
por: Ponzina, Flavio, et al.
Publicado: (2024)
MoEITS: A Green AI approach for simplifying MoE-LLMs
por: Balderas, Luis, et al.
Publicado: (2026)
por: Balderas, Luis, et al.
Publicado: (2026)
Looking Forward: Challenges and Opportunities in Agentic AI Reliability
por: Xing, Liudong, et al.
Publicado: (2025)
por: Xing, Liudong, et al.
Publicado: (2025)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
por: Ma, Jeffrey Jian, et al.
Publicado: (2025)
por: Ma, Jeffrey Jian, et al.
Publicado: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
por: Shao, Zishan, et al.
Publicado: (2025)
por: Shao, Zishan, et al.
Publicado: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
por: Huang, Zixiao, et al.
Publicado: (2025)
por: Huang, Zixiao, et al.
Publicado: (2025)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
por: Garg, Spandan, et al.
Publicado: (2025)
por: Garg, Spandan, et al.
Publicado: (2025)
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
por: Hangun, Batuhan, et al.
Publicado: (2025)
por: Hangun, Batuhan, et al.
Publicado: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
por: Hossain, Md Arafat, et al.
Publicado: (2025)
por: Hossain, Md Arafat, et al.
Publicado: (2025)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
por: Yin, Yishu, et al.
Publicado: (2025)
por: Yin, Yishu, et al.
Publicado: (2025)
Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2025)
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2025)
Breaking the Loop: Detecting and Mitigating Denial-of-Service Vulnerabilities in Large Language Models
por: Yu, Junzhe, et al.
Publicado: (2025)
por: Yu, Junzhe, et al.
Publicado: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
por: Zhao, Yushang, et al.
Publicado: (2025)
por: Zhao, Yushang, et al.
Publicado: (2025)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
por: Kermani, Arshia, et al.
Publicado: (2025)
por: Kermani, Arshia, et al.
Publicado: (2025)
Ejemplares similares
-
Latency Based Tiling
por: Cashman, Jack
Publicado: (2025) -
KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer
por: Al-Qawlaq, Aness, et al.
Publicado: (2024) -
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
por: Jung, Victor J. B., et al.
Publicado: (2024) -
Optimizing LoRa for Edge Computing with TinyML Pipeline for Channel Hopping
por: Grunewald, Marla, et al.
Publicado: (2024) -
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
por: Arantes, Gabriel M., et al.
Publicado: (2025)