The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jiin, Shin, Byeongjun, Chung, Jinha, Rhu, Minsoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024)
by: Yeo, Gwangoo, et al.
Published: (2024)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026)
by: Cho, Eunyeong, et al.
Published: (2026)
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
Agent-X: Full Pipeline Acceleration of On-device AI Agents
by: Chung, Jinha, et al.
Published: (2026)
by: Chung, Jinha, et al.
Published: (2026)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
by: Hyun, Bongjoon, et al.
Published: (2023)
by: Hyun, Bongjoon, et al.
Published: (2023)
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
by: Bang, Jehyeon, et al.
Published: (2023)
by: Bang, Jehyeon, et al.
Published: (2023)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
by: Lee, Dongjae, et al.
Published: (2025)
by: Lee, Dongjae, et al.
Published: (2025)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
by: Lee, Dongjae, et al.
Published: (2024)
by: Lee, Dongjae, et al.
Published: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026)
by: Kim, Hyeseong, et al.
Published: (2026)
LLM-based AI Agent for Sizing of Analog and Mixed Signal Circuit
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
by: Sharma, Amit
Published: (2025)
by: Sharma, Amit
Published: (2025)
EEsizer: LLM-Based AI Agent for Sizing of Analog and Mixed Signal Circuit
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
by: Yoon, Dongho, et al.
Published: (2025)
by: Yoon, Dongho, et al.
Published: (2025)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
by: Morisaki, Keita
Published: (2026)
by: Morisaki, Keita
Published: (2026)
Accelerating Diffusion Models for Generative AI Applications with Silicon Photonics
by: Suresh, Tharini, et al.
Published: (2026)
by: Suresh, Tharini, et al.
Published: (2026)
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
by: Park, Junyoung, et al.
Published: (2024)
by: Park, Junyoung, et al.
Published: (2024)
AiEDA: An Open-Source AI-Aided Design Library for Design-to-Vector
by: Qiu, Yihang, et al.
Published: (2025)
by: Qiu, Yihang, et al.
Published: (2025)
From LLM to Silicon: RL-Driven ASIC Architecture Exploration for On-Device AI Inference
by: Ganti, Ravindra, et al.
Published: (2026)
by: Ganti, Ravindra, et al.
Published: (2026)
A Survey of Circuit Foundation Model: Foundation AI Models for VLSI Circuit Design and EDA
by: Fang, Wenji, et al.
Published: (2025)
by: Fang, Wenji, et al.
Published: (2025)
FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AI
by: Cho, Eun-Su, et al.
Published: (2025)
by: Cho, Eun-Su, et al.
Published: (2025)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
MiCo: End-to-End Mixed Precision Neural Network Co-Exploration Framework for Edge AI
by: Jiang, Zijun, et al.
Published: (2025)
by: Jiang, Zijun, et al.
Published: (2025)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
by: Prakash, Shvetank, et al.
Published: (2025)
by: Prakash, Shvetank, et al.
Published: (2025)
GPT4AIGChip: Towards Next-Generation AI Accelerator Design Automation via Large Language Models
by: Fu, Yonggan, et al.
Published: (2023)
by: Fu, Yonggan, et al.
Published: (2023)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
by: Shao, Kunming, et al.
Published: (2026)
by: Shao, Kunming, et al.
Published: (2026)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE
by: Okubo, Ikumi, et al.
Published: (2024)
by: Okubo, Ikumi, et al.
Published: (2024)
Position Paper: From Edge AI to Adaptive Edge AI
by: Pittorino, Fabrizio, et al.
Published: (2026)
by: Pittorino, Fabrizio, et al.
Published: (2026)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
by: Wang, Run, et al.
Published: (2025)
by: Wang, Run, et al.
Published: (2025)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
by: Xie, Peichen, et al.
Published: (2025)
by: Xie, Peichen, et al.
Published: (2025)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
Autoformalizing Memory Specifications with Agents
by: Ernst, Jan Ole, et al.
Published: (2026)
by: Ernst, Jan Ole, et al.
Published: (2026)
TimingLLM: A Two-Stage Retrieval-Augmented Framework for Pre-Synthesis Timing Prediction from Verilog
by: Abdollahi, Armin, et al.
Published: (2026)
by: Abdollahi, Armin, et al.
Published: (2026)
Periodic Online Testing for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2025)
by: Peltekis, Christodoulos, et al.
Published: (2025)
On Robustness and Generalization of ML-Based Congestion Predictors to Valid and Imperceptible Perturbations
by: Holtz, Chester, et al.
Published: (2024)
by: Holtz, Chester, et al.
Published: (2024)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
by: Shin, Duckgyu, et al.
Published: (2026)
by: Shin, Duckgyu, et al.
Published: (2026)
Similar Items
-
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024) -
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026) -
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
by: Lee, Yunjae, et al.
Published: (2024) -
Agent-X: Full Pipeline Acceleration of On-device AI Agents
by: Chung, Jinha, et al.
Published: (2026) -
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
by: Hyun, Bongjoon, et al.
Published: (2023)