Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Youpeng, LV, Jinpeng, Wu, Di, Wang, Jun, Gooley, Christopher |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
por: Zhao, Youpeng, et al.
Publicado: (2024)
por: Zhao, Youpeng, et al.
Publicado: (2024)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
por: Zhao, Youpeng, et al.
Publicado: (2024)
por: Zhao, Youpeng, et al.
Publicado: (2024)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
por: Jayakody, Shakya, et al.
Publicado: (2026)
por: Jayakody, Shakya, et al.
Publicado: (2026)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
por: Liu, Yi
Publicado: (2026)
por: Liu, Yi
Publicado: (2026)
The Race to Efficiency: A New Perspective on AI Scaling Laws
por: Lu, Chien-Ping
Publicado: (2025)
por: Lu, Chien-Ping
Publicado: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
por: Hossain, Md Arafat, et al.
Publicado: (2025)
por: Hossain, Md Arafat, et al.
Publicado: (2025)
AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
por: Traini, Luca, et al.
Publicado: (2024)
por: Traini, Luca, et al.
Publicado: (2024)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
por: Jha, Mayank
Publicado: (2026)
por: Jha, Mayank
Publicado: (2026)
Adaptive Orchestration for Large-Scale Inference on Heterogeneous Accelerator Systems Balancing Cost, Performance, and Resilience
por: Biran, Yahav, et al.
Publicado: (2025)
por: Biran, Yahav, et al.
Publicado: (2025)
DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning Workloads
por: Zhao, Qidong, et al.
Publicado: (2024)
por: Zhao, Qidong, et al.
Publicado: (2024)
Comment on paper: Position: Rethinking Post-Hoc Search-Based Neural Approaches for Solving Large-Scale Traveling Salesman Problems
por: Min, Yimeng
Publicado: (2024)
por: Min, Yimeng
Publicado: (2024)
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
por: Atinafu, Yonas, et al.
Publicado: (2026)
por: Atinafu, Yonas, et al.
Publicado: (2026)
Offloading and Quality Control for AI Generated Content Services in 6G Mobile Edge Computing Networks
por: Wang, Yitong, et al.
Publicado: (2023)
por: Wang, Yitong, et al.
Publicado: (2023)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
por: Rani, Pooja, et al.
Publicado: (2025)
por: Rani, Pooja, et al.
Publicado: (2025)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
por: Liu, Xiaoxuan, et al.
Publicado: (2024)
por: Liu, Xiaoxuan, et al.
Publicado: (2024)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
por: Liu, Jiashuo, et al.
Publicado: (2025)
por: Liu, Jiashuo, et al.
Publicado: (2025)
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
por: Ellis-Mohr, Austin R., et al.
Publicado: (2025)
por: Ellis-Mohr, Austin R., et al.
Publicado: (2025)
Root Cause Localization for Microservice Systems in Cloud-edge Collaborative Environments
por: Zhu, Yuhan, et al.
Publicado: (2024)
por: Zhu, Yuhan, et al.
Publicado: (2024)
Rapid Augmentations for Time Series (RATS): A High-Performance Library for Time Series Augmentation
por: Skaf, Wadie, et al.
Publicado: (2026)
por: Skaf, Wadie, et al.
Publicado: (2026)
Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge
por: Dan, Jonathan, et al.
Publicado: (2025)
por: Dan, Jonathan, et al.
Publicado: (2025)
XTC, A Research Platform for Optimizing AI Workload Operators
por: Hugo, Pompougnac, et al.
Publicado: (2025)
por: Hugo, Pompougnac, et al.
Publicado: (2025)
DDSA: Dual-Domain Strategic Attack for Spatial-Temporal Efficiency in Adversarial Robustness Testing
por: Hu, Jinwei, et al.
Publicado: (2026)
por: Hu, Jinwei, et al.
Publicado: (2026)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
por: Cai, Yanan, et al.
Publicado: (2025)
por: Cai, Yanan, et al.
Publicado: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
por: Zhao, Yushang, et al.
Publicado: (2025)
por: Zhao, Yushang, et al.
Publicado: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
por: Fang, Yunhua, et al.
Publicado: (2025)
por: Fang, Yunhua, et al.
Publicado: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
por: Qiao, Liang, et al.
Publicado: (2025)
por: Qiao, Liang, et al.
Publicado: (2025)
Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles
por: Gallici, Matteo, et al.
Publicado: (2025)
por: Gallici, Matteo, et al.
Publicado: (2025)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
por: Kermani, Arshia, et al.
Publicado: (2025)
por: Kermani, Arshia, et al.
Publicado: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
por: Singh, Siddharth, et al.
Publicado: (2023)
por: Singh, Siddharth, et al.
Publicado: (2023)
Energy Concerns with HPC Systems and Applications
por: Nana, Roblex, et al.
Publicado: (2023)
por: Nana, Roblex, et al.
Publicado: (2023)
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
por: Liao, Gang, et al.
Publicado: (2025)
por: Liao, Gang, et al.
Publicado: (2025)
This Is Taking Too Long -- Investigating Time as a Proxy for Energy Consumption of LLMs
por: Krupp, Lars, et al.
Publicado: (2026)
por: Krupp, Lars, et al.
Publicado: (2026)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
por: d'Aloisio, Giordano, et al.
Publicado: (2024)
por: d'Aloisio, Giordano, et al.
Publicado: (2024)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
por: Arantes, Gabriel M., et al.
Publicado: (2025)
por: Arantes, Gabriel M., et al.
Publicado: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
por: Ivanov, Andrei, et al.
Publicado: (2025)
por: Ivanov, Andrei, et al.
Publicado: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
por: Prieto, Pablo, et al.
Publicado: (2025)
por: Prieto, Pablo, et al.
Publicado: (2025)
PixLift: Accelerating Web Browsing via AI Upscaling
por: Atinafu, Yonas, et al.
Publicado: (2025)
por: Atinafu, Yonas, et al.
Publicado: (2025)
Tiny-QMoE
por: Cashman, Jack, et al.
Publicado: (2025)
por: Cashman, Jack, et al.
Publicado: (2025)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
por: Pfister, Rolf, et al.
Publicado: (2025)
por: Pfister, Rolf, et al.
Publicado: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
por: Chai, Yuji, et al.
Publicado: (2025)
por: Chai, Yuji, et al.
Publicado: (2025)
Ejemplares similares
-
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
por: Zhao, Youpeng, et al.
Publicado: (2024) -
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
por: Zhao, Youpeng, et al.
Publicado: (2024) -
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
por: Jayakody, Shakya, et al.
Publicado: (2026) -
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
por: Liu, Yi
Publicado: (2026) -
The Race to Efficiency: A New Perspective on AI Scaling Laws
por: Lu, Chien-Ping
Publicado: (2025)