Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baruah, Trinayan, Shivdikar, Kaustubh, Prescott, Sara, Kaeli, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enabling Accelerators for Graph Computing
von: Shivdikar, Kaustubh
Veröffentlicht: (2023)
von: Shivdikar, Kaustubh
Veröffentlicht: (2023)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
Orion: Characterizing and Programming Apple's Neural Engine for LLM Training and Inference
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
Scaling Laws for Floating Point Quantization Training
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
GRPO with State Mutations: Improving LLM-Based Hardware Test Plan Generation
von: Kochar, Dimple Vijay, et al.
Veröffentlicht: (2026)
von: Kochar, Dimple Vijay, et al.
Veröffentlicht: (2026)
DeepRTL2: A Versatile Model for RTL-Related Tasks
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation Model
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View
von: Wu, Yanran, et al.
Veröffentlicht: (2025)
von: Wu, Yanran, et al.
Veröffentlicht: (2025)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases
von: Zheng, Pingqing, et al.
Veröffentlicht: (2025)
von: Zheng, Pingqing, et al.
Veröffentlicht: (2025)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
Understanding and Mitigating Errors of LLM-Generated RTL Code
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
von: Wang, Erwei, et al.
Veröffentlicht: (2025)
von: Wang, Erwei, et al.
Veröffentlicht: (2025)
Speculative Decoding for Verilog: Speed and Quality, All in One
von: Xu, Changran, et al.
Veröffentlicht: (2025)
von: Xu, Changran, et al.
Veröffentlicht: (2025)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
von: Su, Huangyuan, et al.
Veröffentlicht: (2025)
von: Su, Huangyuan, et al.
Veröffentlicht: (2025)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
von: Wu, Yutong, et al.
Veröffentlicht: (2026)
von: Wu, Yutong, et al.
Veröffentlicht: (2026)
GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic Encryption
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2023)
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2023)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
von: Jiang, Wenqi, et al.
Veröffentlicht: (2023)
von: Jiang, Wenqi, et al.
Veröffentlicht: (2023)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
Allo: A Programming Model for Composable Accelerator Design
von: Chen, Hongzheng, et al.
Veröffentlicht: (2024)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2024)
Dato: A Task-Based Programming Model for Dataflow Accelerators
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
Leveraging High-Level Synthesis and Large Language Models to Generate, Simulate, and Deploy a Uniform Random Number Generator Hardware Design
von: Meech, James T.
Veröffentlicht: (2023)
von: Meech, James T.
Veröffentlicht: (2023)
Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning Approach
von: Saroliya, Urvij, et al.
Veröffentlicht: (2024)
von: Saroliya, Urvij, et al.
Veröffentlicht: (2024)
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
von: Sonnino, Lorenzo, et al.
Veröffentlicht: (2023)
von: Sonnino, Lorenzo, et al.
Veröffentlicht: (2023)
Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Lorecast: Layout-Aware Performance and Power Forecasting from Natural Language
von: Wang, Runzhi, et al.
Veröffentlicht: (2025)
von: Wang, Runzhi, et al.
Veröffentlicht: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
The Graph's Apprentice: Teaching an LLM Low Level Knowledge for Circuit Quality Estimation
von: Moravej, Reza, et al.
Veröffentlicht: (2024)
von: Moravej, Reza, et al.
Veröffentlicht: (2024)
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
von: Banasik, Spencer
Veröffentlicht: (2025)
von: Banasik, Spencer
Veröffentlicht: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
Characterizing State Space Model and Hybrid Language Model Performance with Long Context
von: Mitra, Saptarshi, et al.
Veröffentlicht: (2025)
von: Mitra, Saptarshi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enabling Accelerators for Graph Computing
von: Shivdikar, Kaustubh
Veröffentlicht: (2023) -
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024) -
Orion: Characterizing and Programming Apple's Neural Engine for LLM Training and Inference
von: Kumaresan, Ramchand
Veröffentlicht: (2026) -
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
von: Soi, Rupanshu, et al.
Veröffentlicht: (2025) -
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
von: Chen, Hongzheng, et al.
Veröffentlicht: (2025)