A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Zhipeng, Shao, Kunming, Yu, Jiangnan, Zhao, Liang, Cheng, Tim Kwang-Ting, Tsui, Chi-Ying, Yang, Jie, Sawan, Mohamad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
by: Shao, Kunming, et al.
Published: (2026)
by: Shao, Kunming, et al.
Published: (2026)
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
by: Shao, Kunming, et al.
Published: (2025)
by: Shao, Kunming, et al.
Published: (2025)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
by: Zhao, Liang, et al.
Published: (2026)
by: Zhao, Liang, et al.
Published: (2026)
SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
by: Shao, Kunming, et al.
Published: (2024)
by: Shao, Kunming, et al.
Published: (2024)
A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
by: Zhao, Liang, et al.
Published: (2025)
by: Zhao, Liang, et al.
Published: (2025)
Full System Architecture Modeling for Wearable Egocentric Contextual AI
by: Lee, Vincent T., et al.
Published: (2025)
by: Lee, Vincent T., et al.
Published: (2025)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
by: Dong, Pingcheng, et al.
Published: (2026)
by: Dong, Pingcheng, et al.
Published: (2026)
VeriRAG: A Retrieval-Augmented Framework for Automated RTL Testability Repair
by: Qi, Haomin, et al.
Published: (2025)
by: Qi, Haomin, et al.
Published: (2025)
In-Memory Computing Architecture for Efficient Hardware Security
by: Ajmi, Hala, et al.
Published: (2024)
by: Ajmi, Hala, et al.
Published: (2024)
PUMA: Efficient and Low-Cost Memory Allocation and Alignment Support for Processing-Using-Memory Architectures
by: Oliveira, Geraldo F., et al.
Published: (2024)
by: Oliveira, Geraldo F., et al.
Published: (2024)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
by: Lee, Dongjae, et al.
Published: (2025)
by: Lee, Dongjae, et al.
Published: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Efficient Sparse Processing-in-Memory Architecture (ESPIM) for Machine Learning Inference
by: He, Mingxuan, et al.
Published: (2024)
by: He, Mingxuan, et al.
Published: (2024)
LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
by: Vungarala, Deepak, et al.
Published: (2025)
by: Vungarala, Deepak, et al.
Published: (2025)
Configurable Multi-Port Memory Architecture for High-Speed Data Communication
by: Dhakad, Narendra Singh, et al.
Published: (2024)
by: Dhakad, Narendra Singh, et al.
Published: (2024)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
by: Caon, Michele, et al.
Published: (2024)
by: Caon, Michele, et al.
Published: (2024)
Bitwise Logic Using Phase Change Memory Devices Based on the Pinatubo Architecture
by: Aflalo, Noa, et al.
Published: (2024)
by: Aflalo, Noa, et al.
Published: (2024)
DaPPA: A Data-Parallel Programming Framework for Processing-in-Memory Architectures
by: Oliveira, Geraldo F., et al.
Published: (2023)
by: Oliveira, Geraldo F., et al.
Published: (2023)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
by: Ko, Younghoon, et al.
Published: (2026)
by: Ko, Younghoon, et al.
Published: (2026)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026)
by: Xuan, Zihao, et al.
Published: (2026)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
by: Jiang, Aojie, et al.
Published: (2026)
by: Jiang, Aojie, et al.
Published: (2026)
Optimized Memory System Architecture for VESA VDC-M Decoder with Multi-Slice Support
by: Yang, Hannah, et al.
Published: (2025)
by: Yang, Hannah, et al.
Published: (2025)
APACHE: A Processing-Near-Memory Architecture for Multi-Scheme Fully Homomorphic Encryption
by: Ding, Lin, et al.
Published: (2024)
by: Ding, Lin, et al.
Published: (2024)
Near-Memory Architecture for Threshold-Ordinal Surface-Based Corner Detection of Event Cameras
by: Shang, Hongyang, et al.
Published: (2025)
by: Shang, Hongyang, et al.
Published: (2025)
CVA6S+: A Superscalar RISC-V Core with High-Throughput Memory Architecture
by: Tedeschi, Riccardo, et al.
Published: (2025)
by: Tedeschi, Riccardo, et al.
Published: (2025)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
by: Pun, Junius, et al.
Published: (2025)
by: Pun, Junius, et al.
Published: (2025)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
by: Sestito, Cristian, et al.
Published: (2025)
by: Sestito, Cristian, et al.
Published: (2025)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
by: Shi, Shangyi, et al.
Published: (2025)
by: Shi, Shangyi, et al.
Published: (2025)
In-Pipeline Integration of Digital In-Memory-Computing into RISC-V Vector Architecture to Accelerate Deep Learning
by: Spagnolo, Tommaso, et al.
Published: (2026)
by: Spagnolo, Tommaso, et al.
Published: (2026)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
Retrieve, Schedule, Reflect: LLM Agents for Chip QoR Optimization
by: ouyang, Yikang, et al.
Published: (2026)
by: ouyang, Yikang, et al.
Published: (2026)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
GEM3D CIM General Purpose Matrix Computation Using 3D Integrated SRAM eDRAM Hybrid Compute In Memory on Memory Architecture
by: Chakraborty, Subhradip, et al.
Published: (2026)
by: Chakraborty, Subhradip, et al.
Published: (2026)
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
by: Hao, Mingbo, et al.
Published: (2026)
by: Hao, Mingbo, et al.
Published: (2026)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
by: Abdelmaksoud, Ahmed J., et al.
Published: (2026)
by: Abdelmaksoud, Ahmed J., et al.
Published: (2026)
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
Similar Items
-
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
by: Shao, Kunming, et al.
Published: (2026) -
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
by: Shao, Kunming, et al.
Published: (2025) -
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
by: Zhao, Liang, et al.
Published: (2026) -
SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
by: Shao, Kunming, et al.
Published: (2024) -
A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
by: Zhao, Liang, et al.
Published: (2025)