Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Lingjiao, Davis, Jared Quincy, Hanin, Boris, Bailis, Peter, Stoica, Ion, Zaharia, Matei, Zou, James |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design
by: Davis, Jared Quincy, et al.
Published: (2024)
by: Davis, Jared Quincy, et al.
Published: (2024)
Optimizing Model Selection for Compound AI Systems
by: Chen, Lingjiao, et al.
Published: (2025)
by: Chen, Lingjiao, et al.
Published: (2025)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
by: Zhu, Alan, et al.
Published: (2025)
by: Zhu, Alan, et al.
Published: (2025)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
by: Fu, Yichao, et al.
Published: (2024)
by: Fu, Yichao, et al.
Published: (2024)
AI-Driven Research for Databases
by: Cheng, Audrey, et al.
Published: (2026)
by: Cheng, Audrey, et al.
Published: (2026)
Specifications: The missing link to making the development of LLM systems an engineering discipline
by: Stoica, Ion, et al.
Published: (2024)
by: Stoica, Ion, et al.
Published: (2024)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
by: Kolasani, Sai, et al.
Published: (2025)
by: Kolasani, Sai, et al.
Published: (2025)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
by: Bai, Yuqi, et al.
Published: (2025)
by: Bai, Yuqi, et al.
Published: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
Online Speculative Decoding
by: Liu, Xiaoxuan, et al.
Published: (2023)
by: Liu, Xiaoxuan, et al.
Published: (2023)
More Agents Is All You Need
by: Li, Junyou, et al.
Published: (2024)
by: Li, Junyou, et al.
Published: (2024)
Bayesian Inference with Deep Weakly Nonlinear Networks
by: Hanin, Boris, et al.
Published: (2024)
by: Hanin, Boris, et al.
Published: (2024)
Bayesian Inference with Shaped Deep Non-linear MLPs
by: Hanin, Boris, et al.
Published: (2026)
by: Hanin, Boris, et al.
Published: (2026)
No More Adam: Learning Rate Scaling at Initialization is All You Need
by: Xu, Minghao, et al.
Published: (2024)
by: Xu, Minghao, et al.
Published: (2024)
Drowning in Documents: Consequences of Scaling Reranker Inference
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
RAFT: Adapting Language Model to Domain Specific RAG
by: Zhang, Tianjun, et al.
Published: (2024)
by: Zhang, Tianjun, et al.
Published: (2024)
Cycle is All You Need: More Is Different
by: Li, Xin
Published: (2025)
by: Li, Xin
Published: (2025)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
Some Present-Day Problems of Romanian Library Science
by: Stoica, Ion
Published: (1973)
by: Stoica, Ion
Published: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
by: Stoica, Ion
Published: (1972)
by: Stoica, Ion
Published: (1972)
Pie: Pooling CPU Memory for LLM Inference
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Implicit Bias of the JKO Scheme
by: Halmos, Peter, et al.
Published: (2025)
by: Halmos, Peter, et al.
Published: (2025)
Optimizing LLM Queries in Relational Data Analytics Workloads
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
Towards Supporting Legal Argumentation with NLP: Is More Data Really All You Need?
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
by: Cemri, Mert, et al.
Published: (2026)
by: Cemri, Mert, et al.
Published: (2026)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
Agents Are All You Need for LLM Unlearning
by: Sanyal, Debdeep, et al.
Published: (2025)
by: Sanyal, Debdeep, et al.
Published: (2025)
Is Difficulty Calibration All We Need? Towards More Practical Membership Inference Attacks
by: He, Yu, et al.
Published: (2024)
by: He, Yu, et al.
Published: (2024)
ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data
by: Patel, Liana, et al.
Published: (2024)
by: Patel, Liana, et al.
Published: (2024)
Multistep Inverse Is Not All You Need
by: Levine, Alexander, et al.
Published: (2024)
by: Levine, Alexander, et al.
Published: (2024)
Symbolic Regression Is All You Need: From Simulations to Scaling Laws in Binary Neutron Star Mergers
by: Darc, P., et al.
Published: (2025)
by: Darc, P., et al.
Published: (2025)
Delta Fair Sharing: Performance Isolation for Multi-Tenant Storage Systems
by: Griggs, Tyler, et al.
Published: (2026)
by: Griggs, Tyler, et al.
Published: (2026)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
by: Uzan, Omri, et al.
Published: (2024)
by: Uzan, Omri, et al.
Published: (2024)
Resilience Quantification and its Support for Operational Resilience
by: Matei, Ion, et al.
Published: (2026)
by: Matei, Ion, et al.
Published: (2026)
SIEVE: Sample-Efficient Parametric Learning from Natural Language
by: Asawa, Parth, et al.
Published: (2026)
by: Asawa, Parth, et al.
Published: (2026)
Similar Items
-
Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design
by: Davis, Jared Quincy, et al.
Published: (2024) -
Optimizing Model Selection for Compound AI Systems
by: Chen, Lingjiao, et al.
Published: (2025) -
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
by: Zhu, Alan, et al.
Published: (2025) -
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026) -
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
by: Fu, Yichao, et al.
Published: (2024)