Enabling Performant and Flexible Model-Internal Observability for LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Nengneng, Xiong, Sixian, Zhao, Yibo, Wang, Wei, Liu, Zaoxing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance Prediction for Large Systems via Text-to-Text Regression
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
AI-Driven Resource Allocation Framework for Microservices in Hybrid Cloud Platforms
by: Barua, Biman, et al.
Published: (2024)
by: Barua, Biman, et al.
Published: (2024)
VecTrans: Enhancing Compiler Auto-Vectorization through LLM-Assisted Code Transformations
by: Zheng, Zhongchun, et al.
Published: (2025)
by: Zheng, Zhongchun, et al.
Published: (2025)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025)
by: Huang, Zhengxiang, et al.
Published: (2025)
LLM-Vectorizer: LLM-based Verified Loop Vectorizer
by: Taneja, Jubi, et al.
Published: (2024)
by: Taneja, Jubi, et al.
Published: (2024)
Do AI Models Dream of Faster Code? An Empirical Study on LLM-Proposed Performance Improvements in Real-World Software
by: Yi, Lirong, et al.
Published: (2025)
by: Yi, Lirong, et al.
Published: (2025)
What, Indeed, is an Achievable Provable Guarantee for Learning-Enabled Safety Critical Systems
by: Bensalem, Saddek, et al.
Published: (2023)
by: Bensalem, Saddek, et al.
Published: (2023)
KForge: Program Synthesis for Diverse AI Hardware Accelerators
by: Sereda, Taras, et al.
Published: (2025)
by: Sereda, Taras, et al.
Published: (2025)
Learning Performance-Improving Code Edits
by: Shypula, Alexander, et al.
Published: (2023)
by: Shypula, Alexander, et al.
Published: (2023)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
Interpreting Performance Profiles with Deep Learning
by: Liu, Zhuoran
Published: (2025)
by: Liu, Zhuoran
Published: (2025)
Predicting Configuration Performance in Multiple Environments with Sequential Meta-learning
by: Gong, Jingzhi, et al.
Published: (2024)
by: Gong, Jingzhi, et al.
Published: (2024)
AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
by: Traini, Luca, et al.
Published: (2024)
by: Traini, Luca, et al.
Published: (2024)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
by: Wu, Siyu, et al.
Published: (2026)
by: Wu, Siyu, et al.
Published: (2026)
Predicting Software Performance with Divide-and-Learn
by: Gong, Jingzhi, et al.
Published: (2023)
by: Gong, Jingzhi, et al.
Published: (2023)
Prompting for Performance: Exploring LLMs for Configuring Software
by: Spieker, Helge, et al.
Published: (2025)
by: Spieker, Helge, et al.
Published: (2025)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
Kevin: Multi-Turn RL for Generating CUDA Kernels
by: Baronio, Carlo, et al.
Published: (2025)
by: Baronio, Carlo, et al.
Published: (2025)
Lookup multivariate Kolmogorov-Arnold Networks
by: Pozdnyakov, Sergey, et al.
Published: (2025)
by: Pozdnyakov, Sergey, et al.
Published: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
Enhancing Energy-Awareness in Deep Learning through Fine-Grained Energy Measurement
by: Rajput, Saurabhsingh, et al.
Published: (2023)
by: Rajput, Saurabhsingh, et al.
Published: (2023)
Grid-Orch: An LLM-Powered Orchestrator for Distribution Grid Simulation and Analytics
by: Liu, Boming, et al.
Published: (2026)
by: Liu, Boming, et al.
Published: (2026)
On Integrating Resilience and Human Oversight into LLM-Assisted Modeling Workflows for Digital Twins
by: P, Lekshmi, et al.
Published: (2026)
by: P, Lekshmi, et al.
Published: (2026)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
by: d'Aloisio, Giordano, et al.
Published: (2024)
by: d'Aloisio, Giordano, et al.
Published: (2024)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
by: Ma, Jeffrey Jian, et al.
Published: (2025)
by: Ma, Jeffrey Jian, et al.
Published: (2025)
Root Cause Localization for Microservice Systems in Cloud-edge Collaborative Environments
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
LLM-Assisted Semantic Alignment and Integration in Collaborative Model-Based Systems Engineering Using SysML v2
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models
by: Rosas, Miguel Romero, et al.
Published: (2024)
by: Rosas, Miguel Romero, et al.
Published: (2024)
Regression Language Models for Code
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
This Is Taking Too Long -- Investigating Time as a Proxy for Energy Consumption of LLMs
by: Krupp, Lars, et al.
Published: (2026)
by: Krupp, Lars, et al.
Published: (2026)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
by: Rani, Pooja, et al.
Published: (2025)
by: Rani, Pooja, et al.
Published: (2025)
Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis
by: Kumar, Punit, et al.
Published: (2025)
by: Kumar, Punit, et al.
Published: (2025)
DiTOX: Fault Detection and Localization in the ONNX Optimizer
by: Louloudakis, Nikolaos, et al.
Published: (2025)
by: Louloudakis, Nikolaos, et al.
Published: (2025)
LADRI: LeArning-based Dynamic Risk Indicator in Automated Driving System
by: Patel, Anil Ranjitbhai, et al.
Published: (2024)
by: Patel, Anil Ranjitbhai, et al.
Published: (2024)
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
by: Saba, Tara, et al.
Published: (2026)
by: Saba, Tara, et al.
Published: (2026)
SAFLITE: Fuzzing Autonomous Systems via Large Language Models
by: Zhu, Taohong, et al.
Published: (2024)
by: Zhu, Taohong, et al.
Published: (2024)
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG
by: Chen, Pengzhou, et al.
Published: (2026)
by: Chen, Pengzhou, et al.
Published: (2026)
Worst-Case Convergence Time of ML Algorithms via Extreme Value Theory
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
Investigating Execution-Aware Language Models for Code Optimization
by: Di Menna, Federico, et al.
Published: (2025)
by: Di Menna, Federico, et al.
Published: (2025)
Similar Items
-
Performance Prediction for Large Systems via Text-to-Text Regression
by: Akhauri, Yash, et al.
Published: (2025) -
AI-Driven Resource Allocation Framework for Microservices in Hybrid Cloud Platforms
by: Barua, Biman, et al.
Published: (2024) -
VecTrans: Enhancing Compiler Auto-Vectorization through LLM-Assisted Code Transformations
by: Zheng, Zhongchun, et al.
Published: (2025) -
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025) -
LLM-Vectorizer: LLM-based Verified Loop Vectorizer
by: Taneja, Jubi, et al.
Published: (2024)