Salvato in:
| Autori principali: | Li, Dacheng, Cao, Shiyi, Griggs, Tyler, Liu, Shu, Mo, Xiangxi, Tang, Eric, Hegde, Sumanth, Hakhamaneshi, Kourosh, Patil, Shishir G., Zaharia, Matei, Gonzalez, Joseph E., Stoica, Ion |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.07374 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
di: Cao, Shiyi, et al.
Pubblicazione: (2025)
di: Cao, Shiyi, et al.
Pubblicazione: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
di: Cao, Shiyi, et al.
Pubblicazione: (2024)
di: Cao, Shiyi, et al.
Pubblicazione: (2024)
Reasoning Models Can Be Effective Without Thinking
di: Ma, Wenjie, et al.
Pubblicazione: (2025)
di: Ma, Wenjie, et al.
Pubblicazione: (2025)
RAFT: Adapting Language Model to Domain Specific RAG
di: Zhang, Tianjun, et al.
Pubblicazione: (2024)
di: Zhang, Tianjun, et al.
Pubblicazione: (2024)
Optimizing LLM Queries in Relational Data Analytics Workloads
di: Liu, Shu, et al.
Pubblicazione: (2024)
di: Liu, Shu, et al.
Pubblicazione: (2024)
Pie: Pooling CPU Memory for LLM Inference
di: Xu, Yi, et al.
Pubblicazione: (2024)
di: Xu, Yi, et al.
Pubblicazione: (2024)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
di: Chen, Lingjiao, et al.
Pubblicazione: (2026)
di: Chen, Lingjiao, et al.
Pubblicazione: (2026)
Delta Fair Sharing: Performance Isolation for Multi-Tenant Storage Systems
di: Griggs, Tyler, et al.
Pubblicazione: (2026)
di: Griggs, Tyler, et al.
Pubblicazione: (2026)
HashAttention: Semantic Sparsity for Faster Inference
di: Desai, Aditya, et al.
Pubblicazione: (2024)
di: Desai, Aditya, et al.
Pubblicazione: (2024)
Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design
di: Davis, Jared Quincy, et al.
Pubblicazione: (2024)
di: Davis, Jared Quincy, et al.
Pubblicazione: (2024)
Specifications: The missing link to making the development of LLM systems an engineering discipline
di: Stoica, Ion, et al.
Pubblicazione: (2024)
di: Stoica, Ion, et al.
Pubblicazione: (2024)
RAG over Thinking Traces Can Improve Reasoning Tasks
di: Arabzadeh, Negar, et al.
Pubblicazione: (2026)
di: Arabzadeh, Negar, et al.
Pubblicazione: (2026)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
di: Patel, Liana, et al.
Pubblicazione: (2025)
di: Patel, Liana, et al.
Pubblicazione: (2025)
MemGPT: Towards LLMs as Operating Systems
di: Packer, Charles, et al.
Pubblicazione: (2023)
di: Packer, Charles, et al.
Pubblicazione: (2023)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
Optimizing Model Selection for Compound AI Systems
di: Chen, Lingjiao, et al.
Pubblicazione: (2025)
di: Chen, Lingjiao, et al.
Pubblicazione: (2025)
AI-Driven Research for Databases
di: Cheng, Audrey, et al.
Pubblicazione: (2026)
di: Cheng, Audrey, et al.
Pubblicazione: (2026)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
di: Zhu, Alan, et al.
Pubblicazione: (2025)
di: Zhu, Alan, et al.
Pubblicazione: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
di: Griggs, Tyler, et al.
Pubblicazione: (2024)
di: Griggs, Tyler, et al.
Pubblicazione: (2024)
Accelerating Direct Preference Optimization with Prefix Sharing
di: Wang, Franklin, et al.
Pubblicazione: (2024)
di: Wang, Franklin, et al.
Pubblicazione: (2024)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
di: Arabzadeh, Negar, et al.
Pubblicazione: (2026)
di: Arabzadeh, Negar, et al.
Pubblicazione: (2026)
vAttention: Verified Sparse Attention
di: Desai, Aditya, et al.
Pubblicazione: (2025)
di: Desai, Aditya, et al.
Pubblicazione: (2025)
Some Present-Day Problems of Romanian Library Science
di: Stoica, Ion
Pubblicazione: (1973)
di: Stoica, Ion
Pubblicazione: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
di: Stoica, Ion
Pubblicazione: (1972)
di: Stoica, Ion
Pubblicazione: (1972)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
di: Cao, Shiyi, et al.
Pubblicazione: (2026)
di: Cao, Shiyi, et al.
Pubblicazione: (2026)
Semi-Supervised One-Shot Imitation Learning
di: Wu, Philipp, et al.
Pubblicazione: (2024)
di: Wu, Philipp, et al.
Pubblicazione: (2024)
The Time is Here for Just-in-Time Systems: Challenges and Opportunities
di: Liu, Shu, et al.
Pubblicazione: (2026)
di: Liu, Shu, et al.
Pubblicazione: (2026)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
Fairness in Serving Large Language Models
di: Sheng, Ying, et al.
Pubblicazione: (2023)
di: Sheng, Ying, et al.
Pubblicazione: (2023)
MPC-Minimized Secure LLM Inference
di: Rathee, Deevashwer, et al.
Pubblicazione: (2024)
di: Rathee, Deevashwer, et al.
Pubblicazione: (2024)
AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
di: Cemri, Mert, et al.
Pubblicazione: (2026)
di: Cemri, Mert, et al.
Pubblicazione: (2026)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
SIEVE: Sample-Efficient Parametric Learning from Natural Language
di: Asawa, Parth, et al.
Pubblicazione: (2026)
di: Asawa, Parth, et al.
Pubblicazione: (2026)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
di: Mao, Ziming, et al.
Pubblicazione: (2024)
di: Mao, Ziming, et al.
Pubblicazione: (2024)
S*: Test Time Scaling for Code Generation
di: Li, Dacheng, et al.
Pubblicazione: (2025)
di: Li, Dacheng, et al.
Pubblicazione: (2025)
Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
di: Opsahl-Ong, Krista, et al.
Pubblicazione: (2024)
di: Opsahl-Ong, Krista, et al.
Pubblicazione: (2024)
Resilience Quantification and its Support for Operational Resilience
di: Matei, Ion, et al.
Pubblicazione: (2026)
di: Matei, Ion, et al.
Pubblicazione: (2026)
Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
di: Agarwal, Shubham, et al.
Pubblicazione: (2026)
di: Agarwal, Shubham, et al.
Pubblicazione: (2026)
LEANN: A Low-Storage Vector Index
di: Wang, Yichuan, et al.
Pubblicazione: (2025)
di: Wang, Yichuan, et al.
Pubblicazione: (2025)
ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data
di: Patel, Liana, et al.
Pubblicazione: (2024)
di: Patel, Liana, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
di: Cao, Shiyi, et al.
Pubblicazione: (2025) -
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
di: Cao, Shiyi, et al.
Pubblicazione: (2024) -
Reasoning Models Can Be Effective Without Thinking
di: Ma, Wenjie, et al.
Pubblicazione: (2025) -
RAFT: Adapting Language Model to Domain Specific RAG
di: Zhang, Tianjun, et al.
Pubblicazione: (2024) -
Optimizing LLM Queries in Relational Data Analytics Workloads
di: Liu, Shu, et al.
Pubblicazione: (2024)