Enabling Energy-Efficient Deployment of Large Language Models on Memristor Crossbar: A Synergy of Large and Small
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhehui, Luo, Tao, Liu, Cheng, Liu, Weichen, Goh, Rick Siow Mong, Wong, Weng-Fai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA Acceleration
von: Yin, Mengyuan, et al.
Veröffentlicht: (2025)
von: Yin, Mengyuan, et al.
Veröffentlicht: (2025)
Is Quantum Optimization Ready? An Effort Towards Neural Network Compression using Adiabatic Quantum Computing
von: Wang, Zhehui, et al.
Veröffentlicht: (2025)
von: Wang, Zhehui, et al.
Veröffentlicht: (2025)
Aligning Medical Conversational AI through Online Reinforcement Learning with Information-Theoretic Rewards
von: Verma, Tanvi, et al.
Veröffentlicht: (2026)
von: Verma, Tanvi, et al.
Veröffentlicht: (2026)
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
von: Gerlinghoff, Daniel, et al.
Veröffentlicht: (2024)
von: Gerlinghoff, Daniel, et al.
Veröffentlicht: (2024)
How Interpretable are Reasoning Explanations from Prompting Large Language Models?
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
Energy‐Efficient Knapsack Optimization Using Probabilistic Memristor Crossbars
von: Jinzhan Li, et al.
Veröffentlicht: (2025)
von: Jinzhan Li, et al.
Veröffentlicht: (2025)
Structured Semantic Cloaking for Jailbreak Attacks on Large Language Models
von: Sun, Xiaobing, et al.
Veröffentlicht: (2026)
von: Sun, Xiaobing, et al.
Veröffentlicht: (2026)
AiRacleX: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation
von: Gao, Bo, et al.
Veröffentlicht: (2025)
von: Gao, Bo, et al.
Veröffentlicht: (2025)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
Secure and Explainable Fraud Detection in Finance via Hierarchical Multi-source Dataset Distillation
von: Qian, Yiming, et al.
Veröffentlicht: (2025)
von: Qian, Yiming, et al.
Veröffentlicht: (2025)
UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
von: Yu, Kai, et al.
Veröffentlicht: (2024)
von: Yu, Kai, et al.
Veröffentlicht: (2024)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
von: Yang, Ge, et al.
Veröffentlicht: (2024)
von: Yang, Ge, et al.
Veröffentlicht: (2024)
Aligning Crowd-sourced Human Feedback for Reinforcement Learning on Code Generation by Large Language Models
von: Wong, Man Fai, et al.
Veröffentlicht: (2025)
von: Wong, Man Fai, et al.
Veröffentlicht: (2025)
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
von: Song, Yixin, et al.
Veröffentlicht: (2025)
von: Song, Yixin, et al.
Veröffentlicht: (2025)
Can Large Language Models Solve Robot Routing?
von: Huang, Zhehui, et al.
Veröffentlicht: (2024)
von: Huang, Zhehui, et al.
Veröffentlicht: (2024)
Partially Supervised Unpaired Multi-Modal Learning for Label-Efficient Medical Image Segmentation
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
An Ion-Intercalation Memristor for Enabling Full Parallel Writing in Crossbar Networks
von: Zhang, Tingwei, et al.
Veröffentlicht: (2026)
von: Zhang, Tingwei, et al.
Veröffentlicht: (2026)
A Comprehensive Evaluation on Event Reasoning of Large Language Models
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
Energy Efficient Knapsack Optimization Using Probabilistic Memristor Crossbars
von: Li, Jinzhan, et al.
Veröffentlicht: (2024)
von: Li, Jinzhan, et al.
Veröffentlicht: (2024)
Anchor-based Large Language Models
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning
von: Bai, Yang, et al.
Veröffentlicht: (2024)
von: Bai, Yang, et al.
Veröffentlicht: (2024)
History-Aware and Dynamic Client Contribution in Federated Learning
von: Ghosh, Bishwamittra, et al.
Veröffentlicht: (2024)
von: Ghosh, Bishwamittra, et al.
Veröffentlicht: (2024)
Semi-rPPG: Semi-Supervised Remote Physiological Measurement with Curriculum Pseudo-Labeling
von: Wu, Bingjie, et al.
Veröffentlicht: (2025)
von: Wu, Bingjie, et al.
Veröffentlicht: (2025)
Enhancing Community Vision Screening -- AI Driven Retinal Photography for Early Disease Detection and Patient Trust
von: Lei, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Lei, Xiaofeng, et al.
Veröffentlicht: (2024)
Compositional Coordination for Multi-Robot Teams with Large Language Models
von: Huang, Zhehui, et al.
Veröffentlicht: (2025)
von: Huang, Zhehui, et al.
Veröffentlicht: (2025)
Tandem: Riding Together with Large and Small Language Models for Efficient Reasoning
von: Fu, Zichuan, et al.
Veröffentlicht: (2026)
von: Fu, Zichuan, et al.
Veröffentlicht: (2026)
Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning
von: Zhang, Chuang, et al.
Veröffentlicht: (2026)
von: Zhang, Chuang, et al.
Veröffentlicht: (2026)
Deploying Multi-task Online Server with Large Language Model
von: Qu, Yincen, et al.
Veröffentlicht: (2024)
von: Qu, Yincen, et al.
Veröffentlicht: (2024)
SLMQuant:Benchmarking Small Language Model Quantization for Practical Deployment
von: Wang, Jiacheng, et al.
Veröffentlicht: (2025)
von: Wang, Jiacheng, et al.
Veröffentlicht: (2025)
DART: Difficulty-Adaptive Reasoning Truncation for Efficient Large Language Models
von: Zhang, Ruofan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruofan, et al.
Veröffentlicht: (2025)
Joint Knowledge Base Completion and Question Answering by Combining Large Language Models and Small Language Models
von: Liu, Yinan, et al.
Veröffentlicht: (2026)
von: Liu, Yinan, et al.
Veröffentlicht: (2026)
An Aggregation-Free Federated Learning for Tackling Data Heterogeneity
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
von: Shankar, Ravi, et al.
Veröffentlicht: (2025)
von: Shankar, Ravi, et al.
Veröffentlicht: (2025)
SpikySpace: A Spiking State Space Model for Energy-Efficient Time Series Forecasting
von: Tang, Kaiwen, et al.
Veröffentlicht: (2026)
von: Tang, Kaiwen, et al.
Veröffentlicht: (2026)
Federated Co-tuning Framework for Large and Small Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)
von: Fan, Tao, et al.
Veröffentlicht: (2024)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
von: Liu, Weichen, et al.
Veröffentlicht: (2025)
von: Liu, Weichen, et al.
Veröffentlicht: (2025)
Efficient Deployment of Large Language Models on Resource-constrained Devices
von: Yao, Zhiwei, et al.
Veröffentlicht: (2025)
von: Yao, Zhiwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA Acceleration
von: Yin, Mengyuan, et al.
Veröffentlicht: (2025) -
Is Quantum Optimization Ready? An Effort Towards Neural Network Compression using Adiabatic Quantum Computing
von: Wang, Zhehui, et al.
Veröffentlicht: (2025) -
Aligning Medical Conversational AI through Online Reinforcement Learning with Information-Theoretic Rewards
von: Verma, Tanvi, et al.
Veröffentlicht: (2026) -
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
von: Gerlinghoff, Daniel, et al.
Veröffentlicht: (2024) -
How Interpretable are Reasoning Explanations from Prompting Large Language Models?
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)