An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
Fuente:
arXiv
Saved in:
| Main Authors: | Valpey, Benjamin, Li, Xinyi, Pai, Sreepathi, Gopalakrishnan, Ganesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
by: Behrendt, Emet, et al.
Published: (2025)
by: Behrendt, Emet, et al.
Published: (2025)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)
by: Parameshwara, Arya
Published: (2025)
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
Computing sharp and scalable bounds on errors in approximate zeros of univariate polynomials
by: Ramakrishna, P. H. D., et al.
Published: (2003)
by: Ramakrishna, P. H. D., et al.
Published: (2003)
Traffic-Aware Configuration of OPC UA PubSub in Industrial Automation Networks
by: Ekrad, Kasra, et al.
Published: (2026)
by: Ekrad, Kasra, et al.
Published: (2026)
What a Mesh: Formal Security Analysis of WPA3 SAE Wireless Authentication
by: Metere, Roberto, et al.
Published: (2026)
by: Metere, Roberto, et al.
Published: (2026)
Design and Implementation of a Takum Arithmetic Hardware Codec
by: Hunhold, Laslo
Published: (2024)
by: Hunhold, Laslo
Published: (2024)
Approximate Logic Synthesis Using BLASYS
by: Ma, Jingxiao, et al.
Published: (2025)
by: Ma, Jingxiao, et al.
Published: (2025)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
by: Jung, Myoungsoo
Published: (2025)
by: Jung, Myoungsoo
Published: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
AxOCS: Scaling FPGA-based Approximate Operators using Configuration Supersampling
by: Sahoo, Siva Satyendra, et al.
Published: (2023)
by: Sahoo, Siva Satyendra, et al.
Published: (2023)
DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
Tekum: Balanced Ternary Tapered Precision Real Arithmetic
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
An Improved Template for Approximate Computing
by: Rezaalipour, Morteza, et al.
Published: (2025)
by: Rezaalipour, Morteza, et al.
Published: (2025)
Orchestrating Data Collection and Computation in Green IoT Networks
by: Zhan, Junfei, et al.
Published: (2026)
by: Zhan, Junfei, et al.
Published: (2026)
BECS: A Privacy-Preserving Computing Resource Sharing Mechanism for 6G Computing Power Network
by: Yan, Kun, et al.
Published: (2024)
by: Yan, Kun, et al.
Published: (2024)
Solving Elliptic Finite Element Systems in Near-Linear Time with Support Preconditioners
by: Boman, Erik, et al.
Published: (2004)
by: Boman, Erik, et al.
Published: (2004)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
by: Sedukhin, Stanislav, et al.
Published: (2025)
by: Sedukhin, Stanislav, et al.
Published: (2025)
Space-time process algebra with asynchronous communication
by: Bergstra, J. A., et al.
Published: (2024)
by: Bergstra, J. A., et al.
Published: (2024)
Multi-diseases detection with memristive system on chip
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Mestra: Exploring Migration on Virtualized CGRAs
by: Kyriazis, Agamemnon, et al.
Published: (2026)
by: Kyriazis, Agamemnon, et al.
Published: (2026)
Leveraging LLMs for Formal Software Requirements -- Challenges and Prospects
by: Beg, Arshad, et al.
Published: (2025)
by: Beg, Arshad, et al.
Published: (2025)
How long can you sleep? Idle Time System Inefficiencies and Opportunities
by: Antoniou, Georgia, et al.
Published: (2025)
by: Antoniou, Georgia, et al.
Published: (2025)
Active Admission Control in a P2P Distributed Environment for Capacity Efficient Livestreaming in Mobile Wireless Networks
by: Negulescu, Andrei, et al.
Published: (2023)
by: Negulescu, Andrei, et al.
Published: (2023)
Bandwidth Efficient Livestreaming in Mobile Wireless Networks: A Peer-to-Peer ACIDE Solution
by: Negulescu, Andrei, et al.
Published: (2023)
by: Negulescu, Andrei, et al.
Published: (2023)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Static Communication Analysis for Hardware Design
by: Rosendahl, Mads, et al.
Published: (2025)
by: Rosendahl, Mads, et al.
Published: (2025)
Learning-Infused Formal Reasoning: From Contract Synthesis to Artifact Reuse and Formal Semantics
by: Beg, Arshad, et al.
Published: (2026)
by: Beg, Arshad, et al.
Published: (2026)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
Short Version of VERIFAI2026 Paper -- Learning Infused Formal Reasoning: Contract Synthesis, Artefact Reuse and Semantic Foundations
by: Beg, Arshad, et al.
Published: (2026)
by: Beg, Arshad, et al.
Published: (2026)
Minimal Neuron Circuits -- Part I: Resonators
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
Minimal Neuron Circuits: Bursters
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
Safety-Aware AoI Scheduling for LEO Satellite-Assisted Autonomous Driving
by: Sun, Kangkang, et al.
Published: (2026)
by: Sun, Kangkang, et al.
Published: (2026)
Synthesis of signal processing algorithms with constraints on minimal parallelism and memory space
by: Salishev, Sergey
Published: (2025)
by: Salishev, Sergey
Published: (2025)
Towards LLM-based Generation of Human-Readable Proofs in Polynomial Formal Verification
by: Drechsler, Rolf
Published: (2025)
by: Drechsler, Rolf
Published: (2025)
Evaluating LLM-Generated ACSL Annotations for Formal Verification
by: Beg, Arshad, et al.
Published: (2026)
by: Beg, Arshad, et al.
Published: (2026)
Deterministic and Reliable Software-Defined Vehicles: key building blocks, challenges, and vision
by: Teixeira, Pedro Veloso, et al.
Published: (2024)
by: Teixeira, Pedro Veloso, et al.
Published: (2024)
Similar Items
-
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024) -
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
by: Behrendt, Emet, et al.
Published: (2025) -
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025) -
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025) -
Computing sharp and scalable bounds on errors in approximate zeros of univariate polynomials
by: Ramakrishna, P. H. D., et al.
Published: (2003)