FEval-TTC: Fair Evaluation Protocol for Test-Time Compute
Fuente:
arXiv
Saved in:
| Main Authors: | Rumiantsev, Pavel, Pal, Soumyasundar, Zhang, Yingxue, Coates, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
by: Valkanas, Antonios, et al.
Published: (2025)
by: Valkanas, Antonios, et al.
Published: (2025)
Plain Transformers Can be Powerful Graph Learners
by: Ma, Liheng, et al.
Published: (2025)
by: Ma, Liheng, et al.
Published: (2025)
Refining Answer Distributions for Improved Large Language Model Reasoning
by: Pal, Soumyasundar, et al.
Published: (2024)
by: Pal, Soumyasundar, et al.
Published: (2024)
Multi-resolution Time-Series Transformer for Long-term Forecasting
by: Zhang, Yitian, et al.
Published: (2023)
by: Zhang, Yitian, et al.
Published: (2023)
Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Variation Matters: from Mitigating to Embracing Zero-Shot NAS Ranking Function Variation
by: Rumiantsev, Pavel, et al.
Published: (2025)
by: Rumiantsev, Pavel, et al.
Published: (2025)
Half Search Space is All You Need
by: Rumiantsev, Pavel, et al.
Published: (2025)
by: Rumiantsev, Pavel, et al.
Published: (2025)
GraphPPD: Posterior Predictive Modelling for Graph-Level Inference
by: Pal, Soumyasundar, et al.
Published: (2025)
by: Pal, Soumyasundar, et al.
Published: (2025)
Graph Knowledge Distillation to Mixture of Experts
by: Rumiantsev, Pavel, et al.
Published: (2024)
by: Rumiantsev, Pavel, et al.
Published: (2024)
CKGConv: General Graph Convolution with Continuous Kernels
by: Ma, Liheng, et al.
Published: (2024)
by: Ma, Liheng, et al.
Published: (2024)
InnerThoughts: Disentangling Representations and Predictions in Large Language Models
by: Chételat, Didier, et al.
Published: (2025)
by: Chételat, Didier, et al.
Published: (2025)
Sparse Decomposition of Graph Neural Networks
by: Hu, Yaochen, et al.
Published: (2024)
by: Hu, Yaochen, et al.
Published: (2024)
It Takes Two: Your GRPO Is Secretly DPO
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
Test-Time Scaling Makes Overtraining Compute-Optimal
by: Roberts, Nicholas, et al.
Published: (2026)
by: Roberts, Nicholas, et al.
Published: (2026)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
by: Amer, Hossam, et al.
Published: (2026)
by: Amer, Hossam, et al.
Published: (2026)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
by: MiniMax, et al.
Published: (2025)
by: MiniMax, et al.
Published: (2025)
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025)
by: Huang, Jenny Y., et al.
Published: (2025)
When to Ponder: Adaptive Compute Allocation for Code Generation via Test-Time Training
by: Sim, Gihyeon
Published: (2025)
by: Sim, Gihyeon
Published: (2025)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
by: Qi, Jianing, et al.
Published: (2024)
by: Qi, Jianing, et al.
Published: (2024)
Analyzing Fairness of Computer Vision and Natural Language Processing Models
by: Rashed, Ahmed, et al.
Published: (2024)
by: Rashed, Ahmed, et al.
Published: (2024)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
by: Chen, Peter Baile, et al.
Published: (2025)
by: Chen, Peter Baile, et al.
Published: (2025)
Zero-Overhead Introspection for Adaptive Test-Time Compute
by: Manvi, Rohin, et al.
Published: (2025)
by: Manvi, Rohin, et al.
Published: (2025)
TTRL: Test-Time Reinforcement Learning
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
by: Snell, Charlie, et al.
Published: (2024)
by: Snell, Charlie, et al.
Published: (2024)
Test-Time Speculation
by: Kumar, Avinash, et al.
Published: (2026)
by: Kumar, Avinash, et al.
Published: (2026)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Dynamic layer selection in decoder-only transformers
by: Glavas, Theodore, et al.
Published: (2024)
by: Glavas, Theodore, et al.
Published: (2024)
Rank1: Test-Time Compute for Reranking in Information Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
by: Ji, Yixin, et al.
Published: (2025)
by: Ji, Yixin, et al.
Published: (2025)
TimeTox: An LLM-Based Pipeline for Automated Extraction of Time Toxicity from Clinical Trial Protocols
by: Vinjamuri, Saketh, et al.
Published: (2026)
by: Vinjamuri, Saketh, et al.
Published: (2026)
Scaling Test-Time Compute for Agentic Coding
by: Kim, Joongwon, et al.
Published: (2026)
by: Kim, Joongwon, et al.
Published: (2026)
Test-Time Scaling with Reflective Generative Model
by: Wang, Zixiao, et al.
Published: (2025)
by: Wang, Zixiao, et al.
Published: (2025)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
by: Wang, Yinong Oliver, et al.
Published: (2025)
by: Wang, Yinong Oliver, et al.
Published: (2025)
Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
Similar Items
-
C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
by: Valkanas, Antonios, et al.
Published: (2025) -
Plain Transformers Can be Powerful Graph Learners
by: Ma, Liheng, et al.
Published: (2025) -
Refining Answer Distributions for Improved Large Language Model Reasoning
by: Pal, Soumyasundar, et al.
Published: (2024) -
Multi-resolution Time-Series Transformer for Long-term Forecasting
by: Zhang, Yitian, et al.
Published: (2023) -
Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
by: Zhou, Jiaming, et al.
Published: (2024)