A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qiyuan, Lyu, Fuyuan, Sun, Zexu, Wang, Lei, Zhang, Weixu, Hua, Wenyue, Wu, Haolun, Guo, Zhihan, Wang, Yufei, Muennighoff, Niklas, King, Irwin, Liu, Xue, Ma, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborative Performance Prediction for Large Language Models
by: Zhang, Qiyuan, et al.
Published: (2024)
by: Zhang, Qiyuan, et al.
Published: (2024)
Result Diversification in Search and Recommendation: A Survey
by: Wu, Haolun, et al.
Published: (2022)
by: Wu, Haolun, et al.
Published: (2022)
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
by: Gao, Huan-ang, et al.
Published: (2025)
by: Gao, Huan-ang, et al.
Published: (2025)
Give Users the Wheel: Towards Promptable Recommendation Paradigm
by: Lyu, Fuyuan, et al.
Published: (2026)
by: Lyu, Fuyuan, et al.
Published: (2026)
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
by: Guo, Zhihan, et al.
Published: (2025)
by: Guo, Zhihan, et al.
Published: (2025)
Test-Time Discovery via Hashing Memory
by: Lyu, Fan, et al.
Published: (2025)
by: Lyu, Fan, et al.
Published: (2025)
VoxSafeBench: Not Just What Is Said, but Who, How, and Where
by: Wang, Yuxiang, et al.
Published: (2026)
by: Wang, Yuxiang, et al.
Published: (2026)
Exploring Test-time Scaling via Prediction Merging on Large-Scale Recommendation
by: Lyu, Fuyuan, et al.
Published: (2025)
by: Lyu, Fuyuan, et al.
Published: (2025)
RubricBench: Aligning Model-Generated Rubrics with Human Standards
by: Zhang, Qiyuan, et al.
Published: (2026)
by: Zhang, Qiyuan, et al.
Published: (2026)
Beyond Message Passing: A Semantic View of Agent Communication Protocols
by: Yuan, Dun, et al.
Published: (2026)
by: Yuan, Dun, et al.
Published: (2026)
Scale Where It Matters: Training-Free Localized Scaling for Diffusion Models
by: Ren, Qin, et al.
Published: (2025)
by: Ren, Qin, et al.
Published: (2025)
NILE: Internal Consistency Alignment in Large Language Models
by: Hu, Minda, et al.
Published: (2024)
by: Hu, Minda, et al.
Published: (2024)
What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based Agents
by: Xue, Zhaoqian, et al.
Published: (2024)
by: Xue, Zhaoqian, et al.
Published: (2024)
Counterfactual Multi-player Bandits for Explainable Recommendation Diversification
by: Zhang, Yansen, et al.
Published: (2025)
by: Zhang, Yansen, et al.
Published: (2025)
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
Conformal Uncertainty Indicator for Continual Test-Time Adaptation
by: Lyu, Fan, et al.
Published: (2025)
by: Lyu, Fan, et al.
Published: (2025)
PartnerMAS: An LLM Hierarchical Multi-Agent Framework for Business Partner Selection on High-Dimensional Features
by: Li, Lingyao, et al.
Published: (2025)
by: Li, Lingyao, et al.
Published: (2025)
Variational Continual Test-Time Adaptation
by: Lyu, Fan, et al.
Published: (2024)
by: Lyu, Fan, et al.
Published: (2024)
Controllable Continual Test-Time Adaptation
by: Shi, Ziqi, et al.
Published: (2024)
by: Shi, Ziqi, et al.
Published: (2024)
Humanline: Online Alignment as Perceptual Loss
by: Liu, Sijia, et al.
Published: (2025)
by: Liu, Sijia, et al.
Published: (2025)
DecoFuse: Decomposing and Fusing the "What", "Where", and "How" for Brain-Inspired fMRI-to-Video Decoding
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
Beyond Length Scaling: Synergizing Breadth and Depth for Generative Reward Models
by: Zhang, Qiyuan, et al.
Published: (2026)
by: Zhang, Qiyuan, et al.
Published: (2026)
OptDist: Learning Optimal Distribution for Customer Lifetime Value Prediction
by: Weng, Yunpeng, et al.
Published: (2024)
by: Weng, Yunpeng, et al.
Published: (2024)
What Are the Most Important Contributors to Arctic Precipitation—When, Where, and How?
by: Melanie Lauer, et al.
Published: (2025)
by: Melanie Lauer, et al.
Published: (2025)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
by: Zhang, Qiyuan, et al.
Published: (2024)
by: Zhang, Qiyuan, et al.
Published: (2024)
PCLMix: Weakly Supervised Medical Image Segmentation via Pixel-Level Contrastive Learning and Dynamic Mix Augmentation
by: Lei, Yu, et al.
Published: (2024)
by: Lei, Yu, et al.
Published: (2024)
Crosslingual Reasoning through Test-Time Scaling
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
Less is More: Pseudo-Label Filtering for Continual Test-Time Adaptation
by: Tan, Jiayao, et al.
Published: (2024)
by: Tan, Jiayao, et al.
Published: (2024)
How Well Is Your Library Doing What It Claims To Be Doing?
by: Pearson, Richard C.
Published: (1991)
by: Pearson, Richard C.
Published: (1991)
A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?
by: Jiang, Zhihan, et al.
Published: (2023)
by: Jiang, Zhihan, et al.
Published: (2023)
Transparent AI Disclosure Obligations: Who, What, When, Where, Why, How
by: Ali, Abdallah El, et al.
Published: (2024)
by: Ali, Abdallah El, et al.
Published: (2024)
Research Methods in a Nutshell: What, Why, When, Where, Who, and How?
by: Justin Paul, et al.
Published: (2025)
by: Justin Paul, et al.
Published: (2025)
A comparison between the Avila-Gouëzel-Yoccoz norm and the Teichmüller norm
by: Su, Weixu, et al.
Published: (2022)
by: Su, Weixu, et al.
Published: (2022)
Volume of unit balls associated to quadratic differentials
by: Su, Weixu, et al.
Published: (2025)
by: Su, Weixu, et al.
Published: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
From Conflicts to Collisions: A Two-Stage Collision Scenario-Testing Approach for Autonomous Driving Systems
by: Chen, Siyuan, et al.
Published: (2026)
by: Chen, Siyuan, et al.
Published: (2026)
Quantum Concolic Testing
by: Xia, Shangzhou, et al.
Published: (2024)
by: Xia, Shangzhou, et al.
Published: (2024)
Application of Artificial Intelligence in Rock Tunnel Engineering: A Survey on Where and How
by: Xiaojie Yu, et al.
Published: (2025)
by: Xiaojie Yu, et al.
Published: (2025)
What-Meets-Where: Unified Learning of Action and Contact Localization in Images
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
OpenP5: An Open-Source Platform for Developing, Training, and Evaluating LLM-based Recommender Systems
by: Xu, Shuyuan, et al.
Published: (2023)
by: Xu, Shuyuan, et al.
Published: (2023)
Similar Items
-
Collaborative Performance Prediction for Large Language Models
by: Zhang, Qiyuan, et al.
Published: (2024) -
Result Diversification in Search and Recommendation: A Survey
by: Wu, Haolun, et al.
Published: (2022) -
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
by: Gao, Huan-ang, et al.
Published: (2025) -
Give Users the Wheel: Towards Promptable Recommendation Paradigm
by: Lyu, Fuyuan, et al.
Published: (2026) -
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
by: Guo, Zhihan, et al.
Published: (2025)