Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Weihao, He, Keqing, Kuang, Chuqiao, Li, Xiaoguang, He, Junxian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
by: Liu, Wei, et al.
Published: (2023)
by: Liu, Wei, et al.
Published: (2023)
LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
by: Zeng, Weihao, et al.
Published: (2026)
by: Zeng, Weihao, et al.
Published: (2026)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
by: Zeng, Weihao, et al.
Published: (2025)
by: Zeng, Weihao, et al.
Published: (2025)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
by: Huang, Yuzhen, et al.
Published: (2025)
by: Huang, Yuzhen, et al.
Published: (2025)
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
by: Wan, Yuxuan, et al.
Published: (2026)
by: Wan, Yuxuan, et al.
Published: (2026)
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
by: Zhuang, Tianyi, et al.
Published: (2025)
by: Zhuang, Tianyi, et al.
Published: (2025)
UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
by: Lifshitz, Shalev, et al.
Published: (2025)
by: Lifshitz, Shalev, et al.
Published: (2025)
$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models
by: Bilal, Ahsan, et al.
Published: (2026)
by: Bilal, Ahsan, et al.
Published: (2026)
Efficient Test-Time Scaling via Temporal Reasoning Aggregation
by: Li, Jiakun, et al.
Published: (2026)
by: Li, Jiakun, et al.
Published: (2026)
Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
by: Wang, Siqi, et al.
Published: (2024)
by: Wang, Siqi, et al.
Published: (2024)
Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
by: Falcon LLM Team, et al.
Published: (2026)
by: Falcon LLM Team, et al.
Published: (2026)
Scaling Image and Video Generation via Test-Time Evolutionary Search
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
MegaMath: Pushing the Limits of Open Math Corpora
by: Zhou, Fan, et al.
Published: (2025)
by: Zhou, Fan, et al.
Published: (2025)
Test-Time Personalization: A Diagnostic Framework and Probabilistic Fix for Scaling Failures
by: Zhang, Linhai, et al.
Published: (2026)
by: Zhang, Linhai, et al.
Published: (2026)
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
by: Cui, Yingqian, et al.
Published: (2025)
by: Cui, Yingqian, et al.
Published: (2025)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
by: Chen, Jiefeng, et al.
Published: (2025)
by: Chen, Jiefeng, et al.
Published: (2025)
AgentRefine: Enhancing Agent Generalization through Refinement Tuning
by: Fu, Dayuan, et al.
Published: (2025)
by: Fu, Dayuan, et al.
Published: (2025)
Neuro-Symbolic Proof Generation for Scaling Systems Software Verification
by: He, Baoding, et al.
Published: (2026)
by: He, Baoding, et al.
Published: (2026)
AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
by: Cui, Chengming, et al.
Published: (2026)
by: Cui, Chengming, et al.
Published: (2026)
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
by: Kwok, Jacky, et al.
Published: (2025)
by: Kwok, Jacky, et al.
Published: (2025)
OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories
by: Du, Yuwen, et al.
Published: (2026)
by: Du, Yuwen, et al.
Published: (2026)
Verification Limits Code LLM Training
by: Gureja, Srishti, et al.
Published: (2025)
by: Gureja, Srishti, et al.
Published: (2025)
DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning
by: Wang, Yejie, et al.
Published: (2024)
by: Wang, Yejie, et al.
Published: (2024)
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
by: Kim, Kwanyoung, et al.
Published: (2025)
by: Kim, Kwanyoung, et al.
Published: (2025)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
by: Shao, Zhihong, et al.
Published: (2024)
by: Shao, Zhihong, et al.
Published: (2024)
Code Generation by Differential Test Time Scaling
by: He, Yifeng, et al.
Published: (2026)
by: He, Yifeng, et al.
Published: (2026)
Beyond Model Scaling: Test-Time Intervention for Efficient Deep Reasoning
by: Wang, Qianyue, et al.
Published: (2026)
by: Wang, Qianyue, et al.
Published: (2026)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
by: Yang, Siyuan, et al.
Published: (2025)
by: Yang, Siyuan, et al.
Published: (2025)
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
by: Wu, Fang, et al.
Published: (2025)
by: Wu, Fang, et al.
Published: (2025)
Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling
by: Xia, Fanzeng, et al.
Published: (2025)
by: Xia, Fanzeng, et al.
Published: (2025)
Tool Verification for Test-Time Reinforcement Learning
by: Liao, Ruotong, et al.
Published: (2026)
by: Liao, Ruotong, et al.
Published: (2026)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
by: Grigoryan, Lilit, et al.
Published: (2025)
by: Grigoryan, Lilit, et al.
Published: (2025)
G-LNS: Generative Large Neighborhood Search for LLM-Based Automatic Heuristic Design
by: Zhao, Baoyun, et al.
Published: (2026)
by: Zhao, Baoyun, et al.
Published: (2026)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
by: Yin, Yichun, et al.
Published: (2025)
by: Yin, Yichun, et al.
Published: (2025)
Pushing the Limits of BFP on Narrow Precision LLM Inference
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Pushing the Limits of Block Rotations in Post-Training Quantization
by: Sanjeet, Sai, et al.
Published: (2026)
by: Sanjeet, Sai, et al.
Published: (2026)
Similar Items
-
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
by: Liu, Wei, et al.
Published: (2023) -
LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
by: Zeng, Weihao, et al.
Published: (2026) -
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
by: Zeng, Weihao, et al.
Published: (2025) -
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
by: Huang, Yuzhen, et al.
Published: (2025) -
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
by: Wan, Yuxuan, et al.
Published: (2026)