Optimizing Agentic Language Model Inference via Speculative Tool Calls
Fuente:
arXiv
Salvato in:
| Autori principali: | Nichols, Daniel, Singhania, Prajwal, Jekel, Charles, Bhatele, Abhinav, Menon, Harshitha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
Performance-Aligned LLMs for Generating Fast Code
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
di: Singhania, Prajwal, et al.
Pubblicazione: (2025)
di: Singhania, Prajwal, et al.
Pubblicazione: (2025)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024)
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
di: Teranishi, Keita, et al.
Pubblicazione: (2025)
di: Teranishi, Keita, et al.
Pubblicazione: (2025)
Optimizing OpenFaaS on Kubernetes: Comparative Analysis of Language Runtimes and Cluster Distributions
di: Ataie, Ehsan, et al.
Pubblicazione: (2026)
di: Ataie, Ehsan, et al.
Pubblicazione: (2026)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
di: Wang, Zirui, et al.
Pubblicazione: (2026)
di: Wang, Zirui, et al.
Pubblicazione: (2026)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
LLMs as Packagers of HPC Software
di: Melone, Caetano, et al.
Pubblicazione: (2025)
di: Melone, Caetano, et al.
Pubblicazione: (2025)
High-level Stream Processing: A Complementary Analysis of Fault Recovery
di: Vogel, Adriano, et al.
Pubblicazione: (2024)
di: Vogel, Adriano, et al.
Pubblicazione: (2024)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2024)
di: Tariq, Syed Salauddin Mohammad, et al.
Pubblicazione: (2024)
When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications
di: Henning, Sören, et al.
Pubblicazione: (2025)
di: Henning, Sören, et al.
Pubblicazione: (2025)
Where Should I Deploy My Contracts? A Practical Experience Report
di: Lazăr, Cătălina, et al.
Pubblicazione: (2025)
di: Lazăr, Cătălina, et al.
Pubblicazione: (2025)
MPI Implementation Profiling for Better Application Performance
di: Shipley, Riley, et al.
Pubblicazione: (2024)
di: Shipley, Riley, et al.
Pubblicazione: (2024)
Should I Run My Cloud Benchmark on Black Friday?
di: Henning, Sören, et al.
Pubblicazione: (2025)
di: Henning, Sören, et al.
Pubblicazione: (2025)
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024)
di: Cankur, Onur, et al.
Pubblicazione: (2024)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
di: Tuteja, Keshvi, et al.
Pubblicazione: (2025)
di: Tuteja, Keshvi, et al.
Pubblicazione: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
di: Li, Junjie
Pubblicazione: (2024)
di: Li, Junjie
Pubblicazione: (2024)
HPC-Coder: Modeling Parallel Programs using Large Language Models
di: Nichols, Daniel, et al.
Pubblicazione: (2023)
di: Nichols, Daniel, et al.
Pubblicazione: (2023)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
di: Ruiz, Christian Guzman, et al.
Pubblicazione: (2024)
di: Ruiz, Christian Guzman, et al.
Pubblicazione: (2024)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
di: Wen, Zhongzhen, et al.
Pubblicazione: (2026)
di: Wen, Zhongzhen, et al.
Pubblicazione: (2026)
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
di: Saba, Tara, et al.
Pubblicazione: (2026)
di: Saba, Tara, et al.
Pubblicazione: (2026)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
di: Costello, Ian J., et al.
Pubblicazione: (2020)
di: Costello, Ian J., et al.
Pubblicazione: (2020)
Pipit: Scripting the analysis of parallel execution traces
di: Bhatele, Abhinav, et al.
Pubblicazione: (2023)
di: Bhatele, Abhinav, et al.
Pubblicazione: (2023)
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
di: Wen, Zhongzhen, et al.
Pubblicazione: (2025)
di: Wen, Zhongzhen, et al.
Pubblicazione: (2025)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
di: Singh, Samaresh Kumar, et al.
Pubblicazione: (2026)
di: Singh, Samaresh Kumar, et al.
Pubblicazione: (2026)
Rethinking Performance Analysis for Configurable Software Systems: A Case Study from a Fitness Landscape Perspective
di: Huang, Mingyu, et al.
Pubblicazione: (2024)
di: Huang, Mingyu, et al.
Pubblicazione: (2024)
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
di: Krasnovsky, Anatoly A.
Pubblicazione: (2025)
di: Krasnovsky, Anatoly A.
Pubblicazione: (2025)
Emergence-as-Code for Self-Governing Reliable Systems
di: Krasnovsky, Anatoly A.
Pubblicazione: (2026)
di: Krasnovsky, Anatoly A.
Pubblicazione: (2026)
Evaluating Fault Tolerance and Scalability in Distributed File Systems: A Case Study of GFS, HDFS, and MinIO
di: Malhotra, Shubham, et al.
Pubblicazione: (2025)
di: Malhotra, Shubham, et al.
Pubblicazione: (2025)
Adaptive Protein Design Protocols and Middleware
di: Alsaadi, Aymen, et al.
Pubblicazione: (2025)
di: Alsaadi, Aymen, et al.
Pubblicazione: (2025)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
di: Siavashi, Mohammad, et al.
Pubblicazione: (2026)
di: Siavashi, Mohammad, et al.
Pubblicazione: (2026)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
di: Nichols, Daniel, et al.
Pubblicazione: (2026)
di: Nichols, Daniel, et al.
Pubblicazione: (2026)
CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
di: TehraniJamsaz, Ali, et al.
Pubblicazione: (2024)
di: TehraniJamsaz, Ali, et al.
Pubblicazione: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
di: Panova, Elena, et al.
Pubblicazione: (2022)
di: Panova, Elena, et al.
Pubblicazione: (2022)
Easy Acceleration with Distributed Arrays
di: Kepner, Jeremy, et al.
Pubblicazione: (2025)
di: Kepner, Jeremy, et al.
Pubblicazione: (2025)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
Do Large Language Models Understand Performance Optimization?
di: Cui, Bowen, et al.
Pubblicazione: (2025)
di: Cui, Bowen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
di: Nichols, Daniel, et al.
Pubblicazione: (2025) -
Performance-Aligned LLMs for Generating Fast Code
di: Nichols, Daniel, et al.
Pubblicazione: (2024) -
Understanding and Improving Communication Performance in Multi-node LLM Inference
di: Singhania, Prajwal, et al.
Pubblicazione: (2025) -
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024) -
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
di: Teranishi, Keita, et al.
Pubblicazione: (2025)