Litespark Inference on Consumer CPUs: Custom SIMD Kernels for Ternary Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Dade, Nii Osae Osae, Morri, Tony, Rahat, Moinul Hossain, Pal, Sayandip |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
by: Dade, Nii Osae Osae, et al.
Published: (2025)
by: Dade, Nii Osae Osae, et al.
Published: (2025)
SpinGQE: A Generative Quantum Eigensolver for Spin Hamiltonians
by: Holden, Alexander, et al.
Published: (2026)
by: Holden, Alexander, et al.
Published: (2026)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
by: Gope, Dibakar, et al.
Published: (2024)
by: Gope, Dibakar, et al.
Published: (2024)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Towards Intelligent Legal Document Analysis: CNN-Driven Classification of Case Law Texts
by: Hossain, Moinul, et al.
Published: (2026)
by: Hossain, Moinul, et al.
Published: (2026)
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
by: Wang, Jinheng, et al.
Published: (2025)
by: Wang, Jinheng, et al.
Published: (2025)
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree
by: Gao, Xiangxiang, et al.
Published: (2024)
by: Gao, Xiangxiang, et al.
Published: (2024)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Canvas: End-to-End Kernel Architecture Search in Neural Networks
by: Zhao, Chenggang, et al.
Published: (2023)
by: Zhao, Chenggang, et al.
Published: (2023)
Generating Attractive and Authentic Copywriting from Customer Reviews
by: Lin, Yu-Xiang, et al.
Published: (2024)
by: Lin, Yu-Xiang, et al.
Published: (2024)
Customizing Language Model Responses with Contrastive In-Context Learning
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
TODO: Enhancing LLM Alignment with Ternary Preferences
by: Guo, Yuxiang, et al.
Published: (2024)
by: Guo, Yuxiang, et al.
Published: (2024)
Enhancing Programming Error Messages in Real Time with Generative AI
by: Kimmel, Bailey, et al.
Published: (2024)
by: Kimmel, Bailey, et al.
Published: (2024)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Self-Improving Customer Review Response Generation Based on LLMs
by: Azov, Guy, et al.
Published: (2024)
by: Azov, Guy, et al.
Published: (2024)
ChatPattern: Layout Pattern Customization via Natural Language
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
Pruning General Large Language Models into Customized Expert Models
by: Zhao, Yirao, et al.
Published: (2025)
by: Zhao, Yirao, et al.
Published: (2025)
InsightNet: Structured Insight Mining from Customer Feedback
by: Mukku, Sandeep Sricharan, et al.
Published: (2024)
by: Mukku, Sandeep Sricharan, et al.
Published: (2024)
PHANTOM: PHysical ANamorphic Threats Obstructing Connected Vehicle Mobility
by: Shuvo, Md Nahid Hasan, et al.
Published: (2025)
by: Shuvo, Md Nahid Hasan, et al.
Published: (2025)
An Empirical Evaluation of Large Language Models on Consumer Health Questions
by: Abrar, Moaiz, et al.
Published: (2024)
by: Abrar, Moaiz, et al.
Published: (2024)
Patch-Effect Graph Kernels for LLM Interpretability
by: Fernandez-Boullon, Ruben, et al.
Published: (2026)
by: Fernandez-Boullon, Ruben, et al.
Published: (2026)
Fingerprinting Deep Learning Models via Network Traffic Patterns in Federated Learning
by: Shuvo, Md Nahid Hasan, et al.
Published: (2025)
by: Shuvo, Md Nahid Hasan, et al.
Published: (2025)
CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs
by: Shi, Jingzhe, et al.
Published: (2024)
by: Shi, Jingzhe, et al.
Published: (2024)
ICDM 2020 Knowledge Graph Contest: Consumer Event-Cause Extraction
by: He, Congqing, et al.
Published: (2021)
by: He, Congqing, et al.
Published: (2021)
Hitting "Probe"rty with Non-Linearity, and More
by: Pal, Avik, et al.
Published: (2024)
by: Pal, Avik, et al.
Published: (2024)
SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
by: Dudeja, Divij, et al.
Published: (2025)
by: Dudeja, Divij, et al.
Published: (2025)
Prompt-Based Value Steering of Large Language Models
by: Abbo, Giulio Antonio, et al.
Published: (2025)
by: Abbo, Giulio Antonio, et al.
Published: (2025)
TernaryLM: Memory-Efficient Language Modeling via Native 1.5-Bit Quantization with Adaptive Layer-wise Scaling
by: Nargund, Nisharg, et al.
Published: (2026)
by: Nargund, Nisharg, et al.
Published: (2026)
Customized Information and Domain-centric Knowledge Graph Construction with Large Language Models
by: Wawrzik, Frank, et al.
Published: (2024)
by: Wawrzik, Frank, et al.
Published: (2024)
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders
by: Zhu, Xiaofeng, et al.
Published: (2024)
by: Zhu, Xiaofeng, et al.
Published: (2024)
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
by: Kiulian, Artur, et al.
Published: (2024)
by: Kiulian, Artur, et al.
Published: (2024)
Beyond One-Size-Fits-All Summarization: Customizing Summaries for Diverse Users
by: Duran, Mehmet Samet, et al.
Published: (2025)
by: Duran, Mehmet Samet, et al.
Published: (2025)
Distill-C: Enhanced NL2SQL via Distilled Customization with LLMs
by: Hoang, Cong Duy Vu, et al.
Published: (2025)
by: Hoang, Cong Duy Vu, et al.
Published: (2025)
MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
by: Gong, Ming, et al.
Published: (2025)
by: Gong, Ming, et al.
Published: (2025)
The AI Consumer Index (ACE)
by: Benchek, Julien, et al.
Published: (2025)
by: Benchek, Julien, et al.
Published: (2025)
FastKernels: Benchmarking GPU Kernel Generation in Production
by: Oliaro, Gabriele, et al.
Published: (2026)
by: Oliaro, Gabriele, et al.
Published: (2026)
Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop
by: Wang, Yaxuan, et al.
Published: (2026)
by: Wang, Yaxuan, et al.
Published: (2026)
ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
Explainability in Neural Networks for Natural Language Processing Tasks
by: Mersha, Melkamu, et al.
Published: (2024)
by: Mersha, Melkamu, et al.
Published: (2024)
Similar Items
-
Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
by: Dade, Nii Osae Osae, et al.
Published: (2025) -
SpinGQE: A Generative Quantum Eigensolver for Spin Hamiltonians
by: Holden, Alexander, et al.
Published: (2026) -
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
by: Gope, Dibakar, et al.
Published: (2024) -
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
by: Zhang, Tianyi, et al.
Published: (2024) -
Towards Intelligent Legal Document Analysis: CNN-Driven Classification of Case Law Texts
by: Hossain, Moinul, et al.
Published: (2026)