Technology solutions targeting the performance of gen-AI inference in resource constrained platforms
Fuente:
arXiv
Saved in:
| Main Authors: | Kundu, Joyjit, Klein, Joshua, Patel, Aakash, Biswas, Dwaipayan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
System-performance and cost modeling of Large Language Model training and inference
by: Guo, Wenzhe, et al.
Published: (2025)
by: Guo, Wenzhe, et al.
Published: (2025)
Splitwiser: Efficient LM inference with constrained resources
by: Aali, Asad, et al.
Published: (2025)
by: Aali, Asad, et al.
Published: (2025)
Physical Design Exploration of a Wire-Friendly Domain-Specific Processor for Angstrom-Era Nodes
by: Ruotolo, Lorenzo, et al.
Published: (2025)
by: Ruotolo, Lorenzo, et al.
Published: (2025)
A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence
by: Kundu, Yudhishthira, et al.
Published: (2025)
by: Kundu, Yudhishthira, et al.
Published: (2025)
Efficient LLM inference solution on Intel GPU
by: Wu, Hui, et al.
Published: (2023)
by: Wu, Hui, et al.
Published: (2023)
HAPM -- Hardware Aware Pruning Method for CNN hardware accelerators in resource constrained devices
by: Peccia, Federico Nicolas, et al.
Published: (2024)
by: Peccia, Federico Nicolas, et al.
Published: (2024)
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
by: Parvathy, Aradhana Mohan, et al.
Published: (2026)
by: Parvathy, Aradhana Mohan, et al.
Published: (2026)
NetGAP: A Graph-Grammar approach for concept design of networked platforms with extra-functional requirements
by: de Moraes, Rodrigo Saar, et al.
Published: (2023)
by: de Moraes, Rodrigo Saar, et al.
Published: (2023)
Spec2Assertion: Automatic Pre-RTL Assertion Generation using Large Language Models with Progressive Regularization
by: Wu, Fenghua, et al.
Published: (2025)
by: Wu, Fenghua, et al.
Published: (2025)
Multiplier Design Addressing Area-Delay Trade-offs by using DSP and Logic resources on FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification
by: Patel, Vihaan, et al.
Published: (2026)
by: Patel, Vihaan, et al.
Published: (2026)
Versatile silicon integrated photonic processor: a reconfigurable solution for next-generation AI clusters
by: Zhu, Ying, et al.
Published: (2025)
by: Zhu, Ying, et al.
Published: (2025)
Modeling and Simulating Emerging Memory Technologies: A Tutorial
by: Chen, Yun-Chih, et al.
Published: (2025)
by: Chen, Yun-Chih, et al.
Published: (2025)
Neuromorphic Processor Employing FPGA Technology with Universal Interconnections
by: Harlikar, Pracheta, et al.
Published: (2025)
by: Harlikar, Pracheta, et al.
Published: (2025)
Mapping Fusion: Improving FPGA Technology Mapping with ASIC Mapper
by: Yu, Cunxi
Published: (2025)
by: Yu, Cunxi
Published: (2025)
Mixed Structural Choice Operator: Enhancing Technology Mapping with Heterogeneous Representations
by: Hu, Zhang, et al.
Published: (2025)
by: Hu, Zhang, et al.
Published: (2025)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
by: Hyun, Bongjoon, et al.
Published: (2023)
by: Hyun, Bongjoon, et al.
Published: (2023)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
Fixed and Movable Antenna Technology for 6G Integrated Sensing and Communication
by: Zeng, Yong, et al.
Published: (2024)
by: Zeng, Yong, et al.
Published: (2024)
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
by: Ren, Yi, et al.
Published: (2025)
by: Ren, Yi, et al.
Published: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
by: Verhelst, Marian, et al.
Published: (2025)
by: Verhelst, Marian, et al.
Published: (2025)
Memristor Technologies for Dynamic Vision Sensors: A Critical Assessment and Research Roadmap
by: Sadoun, Mohamad Yazan, et al.
Published: (2026)
by: Sadoun, Mohamad Yazan, et al.
Published: (2026)
An Affordable Experimental Technique for SRAM Write Margin Characterization for Nanometer CMOS Technologies
by: Alorda, Bartomeu, et al.
Published: (2024)
by: Alorda, Bartomeu, et al.
Published: (2024)
Exploiting Control-flow Enforcement Technology for Sound and Precise Static Binary Disassembly
by: Zhao, Brian, et al.
Published: (2025)
by: Zhao, Brian, et al.
Published: (2025)
Benchmarking Deep Learning Convolutions on Energy-constrained CPUs
by: Galvez, Enrique, et al.
Published: (2025)
by: Galvez, Enrique, et al.
Published: (2025)
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
by: Kundu, Joyjit, et al.
Published: (2024)
by: Kundu, Joyjit, et al.
Published: (2024)
An integrated design of energy and indoor environmental quality monitoring system for effective building performance management
by: Zakka, Vincent Gbouna, et al.
Published: (2025)
by: Zakka, Vincent Gbouna, et al.
Published: (2025)
System-Technology Co-Optimization of Bitline Routing and Bonding Pathways in Monolithic 3D DRAM Architectures
by: Lee, Kiseok, et al.
Published: (2026)
by: Lee, Kiseok, et al.
Published: (2026)
MINISA: Minimal Instruction Set Architecture for Next-gen Reconfigurable Inference Accelerator
by: Tong, Jianming, et al.
Published: (2026)
by: Tong, Jianming, et al.
Published: (2026)
Memory Scraping Attack on Xilinx FPGAs: Private Data Extraction from Terminated Processes
by: Madabhushi, Bharadwaj, et al.
Published: (2024)
by: Madabhushi, Bharadwaj, et al.
Published: (2024)
A System Level Performance Evaluation for Superconducting Digital Systems
by: Kundu, Joyjit, et al.
Published: (2024)
by: Kundu, Joyjit, et al.
Published: (2024)
The xPU-athalon: Quantifying the Competition of AI Acceleration
by: Golden, Alicia, et al.
Published: (2026)
by: Golden, Alicia, et al.
Published: (2026)
Implementation of a 8-bit Wallace Tree Multiplier
by: Biswas, Ayan, et al.
Published: (2025)
by: Biswas, Ayan, et al.
Published: (2025)
A Reconfigurable Framework for AI-FPGA Agent Integration and Acceleration
by: Yunusoglu, Aybars, et al.
Published: (2026)
by: Yunusoglu, Aybars, et al.
Published: (2026)
Rethinking the Producer-Consumer Relationship in Modern DRAM-Based Systems
by: Patel, Minesh, et al.
Published: (2024)
by: Patel, Minesh, et al.
Published: (2024)
Toward Open-Source Chiplets for HPC and AI: Occamy and Beyond
by: Scheffler, Paul, et al.
Published: (2025)
by: Scheffler, Paul, et al.
Published: (2025)
Exploring the Versal AI Engine for 3D Gaussian Splatting
by: Shimamura, Kotaro, et al.
Published: (2025)
by: Shimamura, Kotaro, et al.
Published: (2025)
Full System Architecture Modeling for Wearable Egocentric Contextual AI
by: Lee, Vincent T., et al.
Published: (2025)
by: Lee, Vincent T., et al.
Published: (2025)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
Similar Items
-
System-performance and cost modeling of Large Language Model training and inference
by: Guo, Wenzhe, et al.
Published: (2025) -
Splitwiser: Efficient LM inference with constrained resources
by: Aali, Asad, et al.
Published: (2025) -
Physical Design Exploration of a Wire-Friendly Domain-Specific Processor for Angstrom-Era Nodes
by: Ruotolo, Lorenzo, et al.
Published: (2025) -
A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence
by: Kundu, Yudhishthira, et al.
Published: (2025) -
Efficient LLM inference solution on Intel GPU
by: Wu, Hui, et al.
Published: (2023)