Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Junlin, Zhu, Shang, Saad-Falcon, Jon, Athiwaratkun, Ben, Wu, Qingyang, Wang, Jue, Song, Shuaiwen Leon, Zhang, Ce, Dhingra, Bhuwan, Zou, James |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Staircase Streaming for Low-Latency Multi-Agent Inference
por: Wang, Junlin, et al.
Publicado: (2025)
por: Wang, Junlin, et al.
Publicado: (2025)
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
por: Wang, Junlin, et al.
Publicado: (2025)
por: Wang, Junlin, et al.
Publicado: (2025)
Mixture-of-Agents Enhances Large Language Model Capabilities
por: Wang, Junlin, et al.
Publicado: (2024)
por: Wang, Junlin, et al.
Publicado: (2024)
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
por: Dokmeci, Berkan, et al.
Publicado: (2025)
por: Dokmeci, Berkan, et al.
Publicado: (2025)
When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
por: Xu, Zhen, et al.
Publicado: (2025)
por: Xu, Zhen, et al.
Publicado: (2025)
Vision2Code: A Multi-Domain Benchmark for Evaluating Image-to-Code Generation
por: Periasami, Ajay Vikram, et al.
Publicado: (2026)
por: Periasami, Ajay Vikram, et al.
Publicado: (2026)
Adversarial Math Word Problem Generation
por: Xie, Roy, et al.
Publicado: (2024)
por: Xie, Roy, et al.
Publicado: (2024)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
por: Wang, Junlin, et al.
Publicado: (2024)
por: Wang, Junlin, et al.
Publicado: (2024)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
por: He, Linda, et al.
Publicado: (2025)
por: He, Linda, et al.
Publicado: (2025)
Automated Benchmark Auditing for AI Agents and Large Language Models
por: Wang, Junlin, et al.
Publicado: (2026)
por: Wang, Junlin, et al.
Publicado: (2026)
How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
por: Cai, Hongyi James, et al.
Publicado: (2025)
por: Cai, Hongyi James, et al.
Publicado: (2025)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
Atomic Consistency Preference Optimization for Long-Form Question Answering
por: Chen, Jingfeng, et al.
Publicado: (2025)
por: Chen, Jingfeng, et al.
Publicado: (2025)
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
por: Thapa, Rahul, et al.
Publicado: (2024)
por: Thapa, Rahul, et al.
Publicado: (2024)
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
por: Maheswaran, Monishwaran, et al.
Publicado: (2026)
por: Maheswaran, Monishwaran, et al.
Publicado: (2026)
Shrinking the Generation-Verification Gap with Weak Verifiers
por: Saad-Falcon, Jon, et al.
Publicado: (2025)
por: Saad-Falcon, Jon, et al.
Publicado: (2025)
A Platform for Investigating Public Health Content with Efficient Concern Classification
por: Li, Christopher, et al.
Publicado: (2025)
por: Li, Christopher, et al.
Publicado: (2025)
ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods
por: Xie, Roy, et al.
Publicado: (2024)
por: Xie, Roy, et al.
Publicado: (2024)
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
por: Thapa, Rahul, et al.
Publicado: (2025)
por: Thapa, Rahul, et al.
Publicado: (2025)
Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
por: Zhang, Muru, et al.
Publicado: (2025)
por: Zhang, Muru, et al.
Publicado: (2025)
Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding
por: Yan, Qingyang, et al.
Publicado: (2025)
por: Yan, Qingyang, et al.
Publicado: (2025)
SCI-Verifier: Scientific Verifier with Thinking
por: Zheng, Shenghe, et al.
Publicado: (2025)
por: Zheng, Shenghe, et al.
Publicado: (2025)
Real-time Factuality Assessment from Adversarial Feedback
por: Chen, Sanxing, et al.
Publicado: (2024)
por: Chen, Sanxing, et al.
Publicado: (2024)
Atomic Self-Consistency for Better Long Form Generations
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
por: Thirukovalluru, Raghuveer, et al.
Publicado: (2024)
RVPO: Risk-Sensitive Alignment via Variance Regularization
por: Montero, Ivan, et al.
Publicado: (2026)
por: Montero, Ivan, et al.
Publicado: (2026)
ChatShop: Interactive Information Seeking with Language Agents
por: Chen, Sanxing, et al.
Publicado: (2024)
por: Chen, Sanxing, et al.
Publicado: (2024)
Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff
por: Holsman, Maximilian, et al.
Publicado: (2025)
por: Holsman, Maximilian, et al.
Publicado: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
Knowing When to Stop: Efficient Context Processing via Latent Sufficiency Signals
por: Xie, Roy, et al.
Publicado: (2025)
por: Xie, Roy, et al.
Publicado: (2025)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
por: Paliotta, Daniele, et al.
Publicado: (2025)
por: Paliotta, Daniele, et al.
Publicado: (2025)
ThinkSwitcher: When to Think Hard, When to Think Fast
por: Liang, Guosheng, et al.
Publicado: (2025)
por: Liang, Guosheng, et al.
Publicado: (2025)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
por: Jia, Jinda, et al.
Publicado: (2026)
por: Jia, Jinda, et al.
Publicado: (2026)
Hierarchical Multi-Label Classification of Online Vaccine Concerns
por: Zhu, Chloe Qinyu, et al.
Publicado: (2024)
por: Zhu, Chloe Qinyu, et al.
Publicado: (2024)
Deep Think with Confidence
por: Fu, Yichao, et al.
Publicado: (2025)
por: Fu, Yichao, et al.
Publicado: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
StepFun-Prover Preview: Let's Think and Verify Step by Step
por: Shang, Shijie, et al.
Publicado: (2025)
por: Shang, Shijie, et al.
Publicado: (2025)
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
por: Wang, Junlin, et al.
Publicado: (2024)
por: Wang, Junlin, et al.
Publicado: (2024)
Bayesian Prediction-Powered Inference
por: Hofer, R. Alex, et al.
Publicado: (2024)
por: Hofer, R. Alex, et al.
Publicado: (2024)
Thinking Fast and Slow with Deep Learning and Tree Search
por: Anthony, Thomas, et al.
Publicado: (2017)
por: Anthony, Thomas, et al.
Publicado: (2017)
Ejemplares similares
-
Staircase Streaming for Low-Latency Multi-Agent Inference
por: Wang, Junlin, et al.
Publicado: (2025) -
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
por: Wang, Junlin, et al.
Publicado: (2025) -
Mixture-of-Agents Enhances Large Language Model Capabilities
por: Wang, Junlin, et al.
Publicado: (2024) -
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
por: Dokmeci, Berkan, et al.
Publicado: (2025) -
When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
por: Xu, Zhen, et al.
Publicado: (2025)