What Limits Agentic Systems Efficiency?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bian, Song, Yan, Minghao, Jayarajan, Anand, Pekhimenko, Gennady, Venkataraman, Shivaram |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Inference-Efficient Language Models
von: Bian, Song, et al.
Veröffentlicht: (2025)
von: Bian, Song, et al.
Veröffentlicht: (2025)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
von: Bian, Song, et al.
Veröffentlicht: (2025)
von: Bian, Song, et al.
Veröffentlicht: (2025)
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling
von: Gao, Yubo, et al.
Veröffentlicht: (2025)
von: Gao, Yubo, et al.
Veröffentlicht: (2025)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
Hexcute: A Compiler Framework for Automating Layout Synthesis in GPU Programs
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
When Do Multi-Agent Systems Outperform? Analysing the Learning Efficiency of Agentic Systems
von: Su, Junwei, et al.
Veröffentlicht: (2026)
von: Su, Junwei, et al.
Veröffentlicht: (2026)
PLoRA: Efficient LoRA Hyperparameter Tuning for Large Models
von: Yan, Minghao, et al.
Veröffentlicht: (2025)
von: Yan, Minghao, et al.
Veröffentlicht: (2025)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
von: Ockerman, Seth, et al.
Veröffentlicht: (2025)
von: Ockerman, Seth, et al.
Veröffentlicht: (2025)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
von: Bian, Song, et al.
Veröffentlicht: (2025)
von: Bian, Song, et al.
Veröffentlicht: (2025)
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
von: Liu, Tong, et al.
Veröffentlicht: (2026)
von: Liu, Tong, et al.
Veröffentlicht: (2026)
Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation
von: Ding, Yaoyao, et al.
Veröffentlicht: (2025)
von: Ding, Yaoyao, et al.
Veröffentlicht: (2025)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
von: Shah, Anvay, et al.
Veröffentlicht: (2026)
von: Shah, Anvay, et al.
Veröffentlicht: (2026)
Learning to Construct Practical Agentic Systems
von: Kumar, Aditya, et al.
Veröffentlicht: (2026)
von: Kumar, Aditya, et al.
Veröffentlicht: (2026)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
von: Liu, Xianyang, et al.
Veröffentlicht: (2026)
von: Liu, Xianyang, et al.
Veröffentlicht: (2026)
Incremental IVF Index Maintenance for Streaming Vector Search
von: Mohoney, Jason, et al.
Veröffentlicht: (2024)
von: Mohoney, Jason, et al.
Veröffentlicht: (2024)
Environment Scaling for Interactive Agentic Experience Collection: A Survey
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
Efficiency Robustness of Dynamic Deep Learning Systems
von: Rathnasuriya, Ravishka, et al.
Veröffentlicht: (2025)
von: Rathnasuriya, Ravishka, et al.
Veröffentlicht: (2025)
Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions
von: Gan, Guo, et al.
Veröffentlicht: (2026)
von: Gan, Guo, et al.
Veröffentlicht: (2026)
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
von: Dong, Honghua, et al.
Veröffentlicht: (2024)
von: Dong, Honghua, et al.
Veröffentlicht: (2024)
An MLCommons Scientific Benchmarks Ontology
von: Hawks, Ben, et al.
Veröffentlicht: (2025)
von: Hawks, Ben, et al.
Veröffentlicht: (2025)
Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
On the Limited Representational Power of Value Functions and its Links to Statistical (In)Efficiency
von: Cheikhi, David, et al.
Veröffentlicht: (2024)
von: Cheikhi, David, et al.
Veröffentlicht: (2024)
Re-ENACT: Reinforcement Learning for Emotional Speech Generation using Actor-Critic Strategy
von: Shankar, Ravi, et al.
Veröffentlicht: (2024)
von: Shankar, Ravi, et al.
Veröffentlicht: (2024)
RMA: an Agentic System for Research-Level Mathematical Problems
von: Zhao, Zelin, et al.
Veröffentlicht: (2026)
von: Zhao, Zelin, et al.
Veröffentlicht: (2026)
Harnessing Agentic Evolution
von: Zhang, Jiayi, et al.
Veröffentlicht: (2026)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2026)
Needles in Needle Stacks: Meaningful Clinical Information Buried in Noisy Waveform Data
von: Nagaraj, Sujay, et al.
Veröffentlicht: (2024)
von: Nagaraj, Sujay, et al.
Veröffentlicht: (2024)
Agentic Uncertainty Reveals Agentic Overconfidence
von: Kaddour, Jean, et al.
Veröffentlicht: (2026)
von: Kaddour, Jean, et al.
Veröffentlicht: (2026)
What Do Latent Action Models Actually Learn?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
ADR: An Agentic Detection System for Enterprise Agentic AI Security
von: Li, Chenning, et al.
Veröffentlicht: (2026)
von: Li, Chenning, et al.
Veröffentlicht: (2026)
AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
von: Kadekodi, Rohan, et al.
Veröffentlicht: (2025)
von: Kadekodi, Rohan, et al.
Veröffentlicht: (2025)
Verifiable Agentic Infrastructure: Proof-Derived Authorization for Sovereign AI Systems
von: He, Jun, et al.
Veröffentlicht: (2026)
von: He, Jun, et al.
Veröffentlicht: (2026)
A Plug-and-Play Fully On-the-Job Real-Time Reinforcement Learning Algorithm for a Direct-Drive Tandem-Wing Experiment Platforms Under Multiple Random Operating Conditions
von: Minghao, Zhang, et al.
Veröffentlicht: (2024)
von: Minghao, Zhang, et al.
Veröffentlicht: (2024)
Uncertainty Quantification in SVM prediction
von: Anand, Pritam
Veröffentlicht: (2025)
von: Anand, Pritam
Veröffentlicht: (2025)
GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
von: Gopalakrishnan, Anand, et al.
Veröffentlicht: (2025)
von: Gopalakrishnan, Anand, et al.
Veröffentlicht: (2025)
From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency
von: Dang, Sizhe, et al.
Veröffentlicht: (2026)
von: Dang, Sizhe, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scaling Inference-Efficient Language Models
von: Bian, Song, et al.
Veröffentlicht: (2025) -
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
von: Bian, Song, et al.
Veröffentlicht: (2025) -
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025) -
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024) -
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
von: Yan, Minghao, et al.
Veröffentlicht: (2023)