Characterizing State Space Model and Hybrid Language Model Performance with Long Context
Fuente:
arXiv
Saved in:
| Main Authors: | Mitra, Saptarshi, Karami, Rachid, Xu, Haocheng, Huang, Sitao, Kwon, Hyoukjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026)
by: Qiao, Ye, et al.
Published: (2026)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Memristor-Based Neural Network Accelerators for Space Applications: Enhancing Performance with Temporal Averaging and SIRENs
by: Rudge, Zacharia A., et al.
Published: (2025)
by: Rudge, Zacharia A., et al.
Published: (2025)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
Large Language Models (LLMs) for Electronic Design Automation (EDA)
by: Xu, Kangwei, et al.
Published: (2025)
by: Xu, Kangwei, et al.
Published: (2025)
Guidance and Control Neural Network Acceleration using Memristors
by: Rudge, Zacharia A., et al.
Published: (2025)
by: Rudge, Zacharia A., et al.
Published: (2025)
Reliable Interval Prediction of Minimum Operating Voltage Based on On-chip Monitors via Conformalized Quantile Regression
by: Yin, Yuxuan, et al.
Published: (2024)
by: Yin, Yuxuan, et al.
Published: (2024)
Optimizing High-Level Synthesis Designs with Retrieval-Augmented Large Language Models
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
The Unseen AI Disruptions for Power Grids: LLM-Induced Transients
by: Li, Yuzhuo, et al.
Published: (2024)
by: Li, Yuzhuo, et al.
Published: (2024)
BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
by: Rakka, Mariam, et al.
Published: (2024)
by: Rakka, Mariam, et al.
Published: (2024)
ProTEA: Programmable Transformer Encoder Acceleration on FPGA
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
by: Hoefler, Torsten, et al.
Published: (2026)
by: Hoefler, Torsten, et al.
Published: (2026)
Deep Learning-Driven Black-Box Doherty Power Amplifier with Pixelated Output Combiner and Extended Efficiency Range
by: Zhou, Han, et al.
Published: (2026)
by: Zhou, Han, et al.
Published: (2026)
Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation
by: Kallakurik, Uttej, et al.
Published: (2025)
by: Kallakurik, Uttej, et al.
Published: (2025)
PGR-DRC: Pre-Global Routing DRC Violation Prediction Using Unsupervised Learning
by: Islam, Riadul, et al.
Published: (2025)
by: Islam, Riadul, et al.
Published: (2025)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
OpenGCRAM: An Open-Source Gain Cell Compiler Enabling Design-Space Exploration for AI Workloads
by: Wang, Xinxin, et al.
Published: (2025)
by: Wang, Xinxin, et al.
Published: (2025)
ReLMXEL: Adaptive RL-Based Memory Controller with Explainable Energy and Latency Optimization
by: Sai, Panuganti Chirag, et al.
Published: (2026)
by: Sai, Panuganti Chirag, et al.
Published: (2026)
Robotics Under Construction: Challenges on Job Sites
by: Uchiito, Haruki, et al.
Published: (2025)
by: Uchiito, Haruki, et al.
Published: (2025)
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
by: Chen, Chun-Ting, et al.
Published: (2025)
by: Chen, Chun-Ting, et al.
Published: (2025)
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
by: Feng, Yuannuo, et al.
Published: (2025)
by: Feng, Yuannuo, et al.
Published: (2025)
Photolithography Control System : A Case Study For Cyber-Physical System
by: Zhang, Youbao, et al.
Published: (2024)
by: Zhang, Youbao, et al.
Published: (2024)
Testing and Fault Tolerance Techniques for CNT-Based FPGAs
by: Lu, Siyuan, et al.
Published: (2025)
by: Lu, Siyuan, et al.
Published: (2025)
AnalogSeeker: An Open-source Foundation Language Model for Analog Circuit Design
by: Chen, Zihao, et al.
Published: (2025)
by: Chen, Zihao, et al.
Published: (2025)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications
by: Xing, Changwen, et al.
Published: (2025)
by: Xing, Changwen, et al.
Published: (2025)
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
by: Tian, Chunlin, et al.
Published: (2025)
by: Tian, Chunlin, et al.
Published: (2025)
MERIT: Multimodal Wearable Vital Sign Waveform Monitoring
by: Tang, Yongyang, et al.
Published: (2024)
by: Tang, Yongyang, et al.
Published: (2024)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
Surrogates, Spikes, and Sparsity: Performance Analysis and Characterization of SNN Hyperparameters on Hardware
by: Aliyev, Ilkin, et al.
Published: (2026)
by: Aliyev, Ilkin, et al.
Published: (2026)
Probabilistic Sensing: Intelligence in Data Sampling
by: Albulushi, Ibrahim, et al.
Published: (2026)
by: Albulushi, Ibrahim, et al.
Published: (2026)
Recent Advances in mm-Wave and Sub-THz/THz Oscillators for FutureG Technologies
by: Behmanesh, Baktash, et al.
Published: (2026)
by: Behmanesh, Baktash, et al.
Published: (2026)
Assessing Large Language Models in Generating RTL Design Specifications
by: Huang, Hung-Ming, et al.
Published: (2025)
by: Huang, Hung-Ming, et al.
Published: (2025)
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
by: Zhong, Ruizhe, et al.
Published: (2023)
by: Zhong, Ruizhe, et al.
Published: (2023)
ReasoningV: Efficient Verilog Code Generation with Adaptive Hybrid Reasoning Model
by: Qin, Haiyan, et al.
Published: (2025)
by: Qin, Haiyan, et al.
Published: (2025)
Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing
by: Liu, Pengju, et al.
Published: (2026)
by: Liu, Pengju, et al.
Published: (2026)
Memristor-Based Selective Convolutional Circuit for High-Density Salt-and-Pepper Noise Removal
by: Ding, Binghui, et al.
Published: (2024)
by: Ding, Binghui, et al.
Published: (2024)
MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging
by: Tang, Jinwei, et al.
Published: (2025)
by: Tang, Jinwei, et al.
Published: (2025)
Similar Items
-
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026) -
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024) -
Memristor-Based Neural Network Accelerators for Space Applications: Enhancing Performance with Temporal Averaging and SIRENs
by: Rudge, Zacharia A., et al.
Published: (2025) -
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024) -
Large Language Models (LLMs) for Electronic Design Automation (EDA)
by: Xu, Kangwei, et al.
Published: (2025)