FOCUS: DLLMs Know How to Tame Their Compute Bound
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Kaihua, Tan, Xin, Zhong, An, Xu, Hong, Canini, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
by: Georgiou, Athos
Published: (2026)
by: Georgiou, Athos
Published: (2026)
REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
by: Chen, Kangqi, et al.
Published: (2025)
by: Chen, Kangqi, et al.
Published: (2025)
Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
by: Lv, Bo, et al.
Published: (2026)
by: Lv, Bo, et al.
Published: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026)
by: Ganjihal, Sanjeev Rao
Published: (2026)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
by: Lo, Yun-Chen, et al.
Published: (2024)
by: Lo, Yun-Chen, et al.
Published: (2024)
LIPPEN: A Lightweight In-Place Pointer Encryption Architecture for Pointer Integrity
by: Iravani, Erfan, et al.
Published: (2026)
by: Iravani, Erfan, et al.
Published: (2026)
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
by: Li, Richie, et al.
Published: (2025)
by: Li, Richie, et al.
Published: (2025)
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
by: Palacios, Pedro, et al.
Published: (2024)
by: Palacios, Pedro, et al.
Published: (2024)
Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
by: Jiang, Zhe, et al.
Published: (2024)
by: Jiang, Zhe, et al.
Published: (2024)
Photonic AI: A Hybrid Diffractive Holographic Neural System for Passive Optical Real-Time Image Classification
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
An Integrated UVM-TLM Co-Simulation Framework for RISC-V Functional Verification and Performance Evaluation
by: Qiu, Ruizhi, et al.
Published: (2025)
by: Qiu, Ruizhi, et al.
Published: (2025)
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
by: Chen, Mu-Chi, et al.
Published: (2026)
by: Chen, Mu-Chi, et al.
Published: (2026)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
by: Kamath, Aditya K, et al.
Published: (2026)
by: Kamath, Aditya K, et al.
Published: (2026)
RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
by: Yildirim, Muhammed, et al.
Published: (2025)
by: Yildirim, Muhammed, et al.
Published: (2025)
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024)
by: Guan, Jiapeng, et al.
Published: (2024)
SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
by: Chen, Mu-Chi, et al.
Published: (2026)
by: Chen, Mu-Chi, et al.
Published: (2026)
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
by: Li, Richie
Published: (2025)
by: Li, Richie
Published: (2025)
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
by: Aspros, Maxime Henri, et al.
Published: (2025)
by: Aspros, Maxime Henri, et al.
Published: (2025)
Multiplier-free In-Memory Vector-Matrix Multiplication Using Distributed Arithmetic
by: Zeller, Felix, et al.
Published: (2025)
by: Zeller, Felix, et al.
Published: (2025)
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
by: Cha, Jungyoub, et al.
Published: (2025)
by: Cha, Jungyoub, et al.
Published: (2025)
Machine Learning for Energy-Performance-aware Scheduling
by: Hu, Zheyuan, et al.
Published: (2026)
by: Hu, Zheyuan, et al.
Published: (2026)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026)
by: Kang, Hyunwoo
Published: (2026)
Arm DynamIQ Shared Unit and Real-Time: An Empirical Evaluation
by: Pradhan, Ashutosh, et al.
Published: (2025)
by: Pradhan, Ashutosh, et al.
Published: (2025)
Memory Access Vectors: Improving Sampling Fidelity for CPU Performance Simulations
by: Caculo, Sriyash, et al.
Published: (2025)
by: Caculo, Sriyash, et al.
Published: (2025)
Biological Intuition on Digital Hardware: An RTL Implementation of Poisson-Encoded SNNs for Static Image Classification
by: Das, Debabrata, et al.
Published: (2026)
by: Das, Debabrata, et al.
Published: (2026)
TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
by: Lin, Baizhou, et al.
Published: (2025)
by: Lin, Baizhou, et al.
Published: (2025)
Dual-Phase Federated Deep Unlearning via Weight-Aware Rollback and Reconstruction
by: Zhou, Changjun, et al.
Published: (2025)
by: Zhou, Changjun, et al.
Published: (2025)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
by: Giorgi, Roberto, et al.
Published: (2026)
by: Giorgi, Roberto, et al.
Published: (2026)
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
by: Hosseini, Peyman, et al.
Published: (2025)
by: Hosseini, Peyman, et al.
Published: (2025)
Evaluating Small Language Models for Front-Door Routing: A Harmonized Benchmark and Synthetic-Traffic Experiment
by: Johnson, Warren, et al.
Published: (2026)
by: Johnson, Warren, et al.
Published: (2026)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
by: Dudek, Piotr
Published: (2024)
by: Dudek, Piotr
Published: (2024)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted Certification
by: Andreasyan, Nick, et al.
Published: (2026)
by: Andreasyan, Nick, et al.
Published: (2026)
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs
by: Li, Haorui, et al.
Published: (2026)
by: Li, Haorui, et al.
Published: (2026)
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Similar Items
-
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
by: Yuan, Aojie, et al.
Published: (2026) -
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024) -
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026) -
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
by: Georgiou, Athos
Published: (2026)