Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
Fuente:
arXiv
Saved in:
| Main Authors: | Behrendt, Emet, Pun, Shing Wai, Nair, Prashant J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
Mestra: Exploring Migration on Virtualized CGRAs
by: Kyriazis, Agamemnon, et al.
Published: (2026)
by: Kyriazis, Agamemnon, et al.
Published: (2026)
Evaluating the Impact of Packet Scheduling and Congestion Control Algorithms on MPTCP Performance over Heterogeneous Networks
by: Dimopoulos, Dimitrios, et al.
Published: (2025)
by: Dimopoulos, Dimitrios, et al.
Published: (2025)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
by: Giorgi, Roberto, et al.
Published: (2026)
by: Giorgi, Roberto, et al.
Published: (2026)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
FREESS: An Educational Simulator of a RISC-V-Inspired Superscalar Processor Based on Tomasulo's Algorithm
by: Giorgi, Roberto
Published: (2025)
by: Giorgi, Roberto
Published: (2025)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)
by: Parameshwara, Arya
Published: (2025)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
by: Jung, Myoungsoo
Published: (2025)
by: Jung, Myoungsoo
Published: (2025)
SISA: A Scale-In Systolic Array for GEMM Acceleration
by: Altamura, Luigi, et al.
Published: (2026)
by: Altamura, Luigi, et al.
Published: (2026)
Application-Driven Exascale: The JUPITER Benchmark Suite
by: Herten, Andreas, et al.
Published: (2024)
by: Herten, Andreas, et al.
Published: (2024)
HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
by: Suryadevara, Pranav
Published: (2025)
by: Suryadevara, Pranav
Published: (2025)
Glass-Box Analysis for Computer Systems: Transparency Index, Shapley Attribution, and Markov Models of Branch Prediction
by: Alpay, Faruk, et al.
Published: (2025)
by: Alpay, Faruk, et al.
Published: (2025)
On the Benefits of Traffic "Reprofiling" -- The Multiple Hops Case -- Part II
by: Qiu, Jiaming, et al.
Published: (2026)
by: Qiu, Jiaming, et al.
Published: (2026)
Cost-effective and performant virtual WANs with CORNIFER
by: Anjali, et al.
Published: (2024)
by: Anjali, et al.
Published: (2024)
On the Power Saving in High-Speed Ethernet-based Networks for Supercomputers and Data Centers
by: de la Rosa, Miguel Sánchez, et al.
Published: (2025)
by: de la Rosa, Miguel Sánchez, et al.
Published: (2025)
Hermes: A General-Purpose Proxy-Enabled Networking Architecture
by: Farkiani, Behrooz, et al.
Published: (2024)
by: Farkiani, Behrooz, et al.
Published: (2024)
Active Admission Control in a P2P Distributed Environment for Capacity Efficient Livestreaming in Mobile Wireless Networks
by: Negulescu, Andrei, et al.
Published: (2023)
by: Negulescu, Andrei, et al.
Published: (2023)
Bandwidth Efficient Livestreaming in Mobile Wireless Networks: A Peer-to-Peer ACIDE Solution
by: Negulescu, Andrei, et al.
Published: (2023)
by: Negulescu, Andrei, et al.
Published: (2023)
ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
by: Das, Tamoghno, et al.
Published: (2025)
by: Das, Tamoghno, et al.
Published: (2025)
An Integrated UVM-TLM Co-Simulation Framework for RISC-V Functional Verification and Performance Evaluation
by: Qiu, Ruizhi, et al.
Published: (2025)
by: Qiu, Ruizhi, et al.
Published: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
by: Özyılmaz, Mustafa Mert
Published: (2026)
by: Özyılmaz, Mustafa Mert
Published: (2026)
How long can you sleep? Idle Time System Inefficiencies and Opportunities
by: Antoniou, Georgia, et al.
Published: (2025)
by: Antoniou, Georgia, et al.
Published: (2025)
Design Space Exploration of Approximate Computing Techniques with a Reinforcement Learning Approach
by: Saeedi, Sepide, et al.
Published: (2023)
by: Saeedi, Sepide, et al.
Published: (2023)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Traffic-Aware Configuration of OPC UA PubSub in Industrial Automation Networks
by: Ekrad, Kasra, et al.
Published: (2026)
by: Ekrad, Kasra, et al.
Published: (2026)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
HoneyDOC: An Efficient Honeypot Architecture Enabling All-Round Design
by: Fan, Wenjun, et al.
Published: (2024)
by: Fan, Wenjun, et al.
Published: (2024)
Wavelet-Based CSI Reconstruction for Improved Wireless Security Through Channel Reciprocity
by: Basha, Nora, et al.
Published: (2025)
by: Basha, Nora, et al.
Published: (2025)
Multi-diseases detection with memristive system on chip
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
by: Tahmasebi, Faraz, et al.
Published: (2025)
by: Tahmasebi, Faraz, et al.
Published: (2025)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
by: Adhinarayanan, Vignesh, et al.
Published: (2026)
by: Adhinarayanan, Vignesh, et al.
Published: (2026)
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
by: Siracusa, Marco, et al.
Published: (2025)
by: Siracusa, Marco, et al.
Published: (2025)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026)
by: Kang, Hyunwoo
Published: (2026)
CEO-DC: Driving Decarbonization in HPC Data Centers with Actionable Insights
by: Álvarez, Rubén Rodríguez, et al.
Published: (2025)
by: Álvarez, Rubén Rodríguez, et al.
Published: (2025)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
by: Goldman, Amos, et al.
Published: (2026)
by: Goldman, Amos, et al.
Published: (2026)
Similar Items
-
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025) -
Mestra: Exploring Migration on Virtualized CGRAs
by: Kyriazis, Agamemnon, et al.
Published: (2026) -
Evaluating the Impact of Packet Scheduling and Congestion Control Algorithms on MPTCP Performance over Heterogeneous Networks
by: Dimopoulos, Dimitrios, et al.
Published: (2025) -
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025) -
DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
by: Yang, Xin, et al.
Published: (2025)