KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Lishuo, Xu, Shaojie, Chen, Jinwu, Yan, Changwei, Wang, Jiajie, Jiang, Zhe, Shan, Weiwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Single-Cell Universal Logic-in-Memory Using 2T-nC FeRAM: An Area and Energy-Efficient Approach for Bulk Bitwise Computation
von: Biswas, Rudra, et al.
Veröffentlicht: (2025)
von: Biswas, Rudra, et al.
Veröffentlicht: (2025)
Fault-Free Analog Computing with Imperfect Hardware
von: Xu, Zhicheng, et al.
Veröffentlicht: (2025)
von: Xu, Zhicheng, et al.
Veröffentlicht: (2025)
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
Towards Efficient Flash Caches with Emerging NVMe Flexible Data Placement SSDs
von: Allison, Michael, et al.
Veröffentlicht: (2025)
von: Allison, Michael, et al.
Veröffentlicht: (2025)
Analog-to-Stochastic Converter Using Magnetic Tunnel Junction Devices for Vision Chips
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026)
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026)
All-in-One Analog AI Hardware: On-Chip Training and Inference with Conductive-Metal-Oxide/HfOx ReRAM Devices
von: Falcone, Donato Francesco, et al.
Veröffentlicht: (2025)
von: Falcone, Donato Francesco, et al.
Veröffentlicht: (2025)
MASIM: An Efficient Multi-Array Scheduler for In-Memory SIMD Computation
von: Qian, Xingyue, et al.
Veröffentlicht: (2024)
von: Qian, Xingyue, et al.
Veröffentlicht: (2024)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
von: Pan, Yue, et al.
Veröffentlicht: (2025)
von: Pan, Yue, et al.
Veröffentlicht: (2025)
Partially-Precise Computing Paradigm for Efficient Hardware Implementation of Application-Specific Embedded Systems
von: Faryabi, Mohsen, et al.
Veröffentlicht: (2024)
von: Faryabi, Mohsen, et al.
Veröffentlicht: (2024)
ReCross: Efficient Embedding Reduction Scheme for In-Memory Computing using ReRAM-Based Crossbar
von: Lai, Yu-Hong, et al.
Veröffentlicht: (2025)
von: Lai, Yu-Hong, et al.
Veröffentlicht: (2025)
TetrisG-SDK: Efficient Convolutional Layer Mapping with Adaptive Windows and Grouped Convolutions for Fast In-Memory Computing
von: Dong, Ke, et al.
Veröffentlicht: (2026)
von: Dong, Ke, et al.
Veröffentlicht: (2026)
Towards Efficient Hyperdimensional Computing Using Photonics
von: Fayza, Farbin, et al.
Veröffentlicht: (2023)
von: Fayza, Farbin, et al.
Veröffentlicht: (2023)
AnalogToBi: Device-Level Analog Circuit Topology Generation via Bipartite Graph and Grammar Guided Decoding
von: Kim, Seungmin, et al.
Veröffentlicht: (2026)
von: Kim, Seungmin, et al.
Veröffentlicht: (2026)
Device-Circuit Co-Design of Variation-Resilient Read and Write Drivers for Antiferromagnetic Tunnel Junction (AFMTJ) Memories
von: Choudhary, Yousuf, et al.
Veröffentlicht: (2026)
von: Choudhary, Yousuf, et al.
Veröffentlicht: (2026)
Computing High-Degree Polynomial Gradients in Memory
von: Bhattacharya, T., et al.
Veröffentlicht: (2024)
von: Bhattacharya, T., et al.
Veröffentlicht: (2024)
Hybrid Temporal Computing for Lower Power Hardware Accelerators
von: Tasnim, Maliha, et al.
Veröffentlicht: (2024)
von: Tasnim, Maliha, et al.
Veröffentlicht: (2024)
PEARL: Power- and Energy-Aware Multicore Intermittent Computing
von: Akhunov, Khakim, et al.
Veröffentlicht: (2025)
von: Akhunov, Khakim, et al.
Veröffentlicht: (2025)
FPIA: Field-Programmable Ising Arrays with In-Memory Computing
von: Hutchinson, George Higgins, et al.
Veröffentlicht: (2024)
von: Hutchinson, George Higgins, et al.
Veröffentlicht: (2024)
All-in-Memory Stochastic Computing using ReRAM
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2025)
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2025)
WAGONN: Weight Bit Agglomeration in Crossbar Arrays for Reduced Impact of Interconnect Resistance on DNN Inference Accuracy
von: Victor, Jeffry, et al.
Veröffentlicht: (2024)
von: Victor, Jeffry, et al.
Veröffentlicht: (2024)
PICO-RAM: A PVT-Insensitive Analog Compute-In-Memory SRAM Macro with In-Situ Multi-Bit Charge Computing and 6T Thin-Cell-Compatible Layout
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
Antiferromagnetic Tunnel Junctions (AFMTJs) for In-Memory Computing: Modeling and Case Study
von: Choudhary, Yousuf, et al.
Veröffentlicht: (2026)
von: Choudhary, Yousuf, et al.
Veröffentlicht: (2026)
Bayes2IMC: In-Memory Computing for Bayesian Binary Neural Networks
von: Katti, Prabodh, et al.
Veröffentlicht: (2024)
von: Katti, Prabodh, et al.
Veröffentlicht: (2024)
SKYLIGHT: A Scalable Hundred-Channel 3D Photonic In-Memory Tensor Core Architecture for Real-time AI Inference
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
Sensitivity-Aware Mixed-Precision Quantization for ReRAM-based Computing-in-Memory
von: Chen, Guan-Cheng, et al.
Veröffentlicht: (2025)
von: Chen, Guan-Cheng, et al.
Veröffentlicht: (2025)
Cognition Engines: A Row-Scale HVDC Architecture for Computational Continuity of AI
von: Churnock, Paul
Veröffentlicht: (2025)
von: Churnock, Paul
Veröffentlicht: (2025)
The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
XL-HD: Extended Learning in Hyperdimensional Computing via Deterministic Projections for In-Memory Accelerators
von: Moon, Sabrina Hassan, et al.
Veröffentlicht: (2026)
von: Moon, Sabrina Hassan, et al.
Veröffentlicht: (2026)
Scalable Digital Compute-in-Memory Ising Machines for Robustness Verification of Binary Neural Networks
von: Vadlamani, Madhav, et al.
Veröffentlicht: (2026)
von: Vadlamani, Madhav, et al.
Veröffentlicht: (2026)
Efficient FIR filtering with Bit Layer Multiply Accumulator
von: Liguori, Vincenzo
Veröffentlicht: (2024)
von: Liguori, Vincenzo
Veröffentlicht: (2024)
Evaluating Computing Platforms for Sustainability: A Comparative Analysis of FPGAs against ASICs, GPUs, and CPUs
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2026)
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2026)
Stoch-IMC: A Bit-Parallel Stochastic In-Memory Computing Architecture Based on STT-MRAM
von: Hajisadeghi, Amir M., et al.
Veröffentlicht: (2024)
von: Hajisadeghi, Amir M., et al.
Veröffentlicht: (2024)
A Novel 8T SRAM-Based In-Memory Computing Architecture for MAC-Derived Logical Functions
von: M, Amogh K, et al.
Veröffentlicht: (2025)
von: M, Amogh K, et al.
Veröffentlicht: (2025)
Novel Efficient Scalable QCA XOR and Full Adder Designs
von: Safaiezadeh, Behrouz, et al.
Veröffentlicht: (2023)
von: Safaiezadeh, Behrouz, et al.
Veröffentlicht: (2023)
Function Approximation Using Analog Building Blocks in Flexible Electronics
von: Duarte, Paula Carolina Lozano, et al.
Veröffentlicht: (2025)
von: Duarte, Paula Carolina Lozano, et al.
Veröffentlicht: (2025)
HiSEP-Q: A Highly Scalable and Efficient Quantum Control Processor for Superconducting Qubits
von: Guo, Xiaorang, et al.
Veröffentlicht: (2023)
von: Guo, Xiaorang, et al.
Veröffentlicht: (2023)
Through Silicon Via Aware Design Planning for Thermally Efficient 3-D Integrated Circuits
von: Chen, Yibo, et al.
Veröffentlicht: (2025)
von: Chen, Yibo, et al.
Veröffentlicht: (2025)
Length-Matching Routing for Programmable Photonic Circuits Using Best-First Strategy
von: Wang, Xiaoke, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoke, et al.
Veröffentlicht: (2025)
A Reconfigurable Time-Domain In-Memory Computing Macro using FeFET-Based CAM with Multilevel Delay Calibration in 28 nm CMOS
von: Mattar, Jeries, et al.
Veröffentlicht: (2025)
von: Mattar, Jeries, et al.
Veröffentlicht: (2025)
Power-Area Efficient Serial IMPLY-based 4:2 Compressor Applied in Data-Intensive Applications
von: Bagheralmoosavi, Bahareh, et al.
Veröffentlicht: (2024)
von: Bagheralmoosavi, Bahareh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Single-Cell Universal Logic-in-Memory Using 2T-nC FeRAM: An Area and Energy-Efficient Approach for Bulk Bitwise Computation
von: Biswas, Rudra, et al.
Veröffentlicht: (2025) -
Fault-Free Analog Computing with Imperfect Hardware
von: Xu, Zhicheng, et al.
Veröffentlicht: (2025) -
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026) -
Towards Efficient Flash Caches with Emerging NVMe Flexible Data Placement SSDs
von: Allison, Michael, et al.
Veröffentlicht: (2025) -
Analog-to-Stochastic Converter Using Magnetic Tunnel Junction Devices for Vision Chips
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026)