$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
Fuente:
arXiv
Guardado en:
| Autores principales: | Singh, Harman, Li, Xiuyu, Sareen, Kusha, Maheswaran, Monishwaran, Tan, Sijun, Wu, Xiaoxia, Wang, Junxiong, Ariyak, Alpay, Wu, Qingyang, Khaki, Samir, Tiwari, Rishabh, Lian, Long, Lu, Yucheng, Li, Boyi, Suhr, Alane, Athiwaratkun, Ben, Keutzer, Kurt |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
por: Maheswaran, Monishwaran, et al.
Publicado: (2026)
por: Maheswaran, Monishwaran, et al.
Publicado: (2026)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
por: Tiwari, Rishabh, et al.
Publicado: (2026)
por: Tiwari, Rishabh, et al.
Publicado: (2026)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
por: Tiwari, Rishabh, et al.
Publicado: (2026)
por: Tiwari, Rishabh, et al.
Publicado: (2026)
Learning Adaptive Parallel Reasoning with Language Models
por: Pan, Jiayi, et al.
Publicado: (2025)
por: Pan, Jiayi, et al.
Publicado: (2025)
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
por: Shao, Zelei, et al.
Publicado: (2025)
por: Shao, Zelei, et al.
Publicado: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
por: Xia, Haojun, et al.
Publicado: (2025)
por: Xia, Haojun, et al.
Publicado: (2025)
LLoCO: Learning Long Contexts Offline
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
por: Sareen, Kusha, et al.
Publicado: (2025)
por: Sareen, Kusha, et al.
Publicado: (2025)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
por: Khaki, Samir, et al.
Publicado: (2025)
por: Khaki, Samir, et al.
Publicado: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
por: Hooper, Coleman, et al.
Publicado: (2024)
por: Hooper, Coleman, et al.
Publicado: (2024)
TASER: Translation Assessment via Systematic Evaluation and Reasoning
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
CDLM: Consistency Diffusion Language Models For Faster Sampling
por: Kim, Minseo, et al.
Publicado: (2025)
por: Kim, Minseo, et al.
Publicado: (2025)
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
por: Jain, Vineet, et al.
Publicado: (2025)
por: Jain, Vineet, et al.
Publicado: (2025)
Residual Context Diffusion Language Models
por: Hu, Yuezhou, et al.
Publicado: (2026)
por: Hu, Yuezhou, et al.
Publicado: (2026)
Evaluating Model Perception of Color Illusions in Photorealistic Scenes
por: Mao, Lingjun, et al.
Publicado: (2024)
por: Mao, Lingjun, et al.
Publicado: (2024)
Grounding Language in Multi-Perspective Referential Communication
por: Tang, Zineng, et al.
Publicado: (2024)
por: Tang, Zineng, et al.
Publicado: (2024)
ETS: Efficient Tree Search for Inference-Time Scaling
por: Hooper, Coleman, et al.
Publicado: (2025)
por: Hooper, Coleman, et al.
Publicado: (2025)
Energy Loss Functions for Physical Systems
por: Kaba, Sékou-Oumar, et al.
Publicado: (2025)
por: Kaba, Sékou-Oumar, et al.
Publicado: (2025)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
por: Liu, Zhijian, et al.
Publicado: (2024)
por: Liu, Zhijian, et al.
Publicado: (2024)
Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
por: Oncescu, Costin-Andrei, et al.
Publicado: (2025)
por: Oncescu, Costin-Andrei, et al.
Publicado: (2025)
Using Language Models to Disambiguate Lexical Choices in Translation
por: Barua, Josh, et al.
Publicado: (2024)
por: Barua, Josh, et al.
Publicado: (2024)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
por: Sclar, Melanie, et al.
Publicado: (2023)
por: Sclar, Melanie, et al.
Publicado: (2023)
Long Chain-of-Thought Reasoning Across Languages
por: Barua, Josh, et al.
Publicado: (2025)
por: Barua, Josh, et al.
Publicado: (2025)
The Role of Symmetry in Optimizing Overparameterized Networks
por: Sareen, Kusha, et al.
Publicado: (2026)
por: Sareen, Kusha, et al.
Publicado: (2026)
The Need for Speed: Pruning Transformers with One Recipe
por: Khaki, Samir, et al.
Publicado: (2024)
por: Khaki, Samir, et al.
Publicado: (2024)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
por: Xi, Haocheng, et al.
Publicado: (2026)
por: Xi, Haocheng, et al.
Publicado: (2026)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
por: Lidayan, Aly, et al.
Publicado: (2025)
por: Lidayan, Aly, et al.
Publicado: (2025)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
por: Nehrdich, Sebastian, et al.
Publicado: (2026)
por: Nehrdich, Sebastian, et al.
Publicado: (2026)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
Introspective Diffusion Language Models
por: Yu, Yifan, et al.
Publicado: (2026)
por: Yu, Yifan, et al.
Publicado: (2026)
ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
por: Lian, Long, et al.
Publicado: (2025)
por: Lian, Long, et al.
Publicado: (2025)
Teaching India–Pakistan Relations
por: Anand, Kusha
Publicado: (2023)
por: Anand, Kusha
Publicado: (2023)
SciML Agents: Write the Solver, Not the Solution
por: Gaonkar, Saarth, et al.
Publicado: (2025)
por: Gaonkar, Saarth, et al.
Publicado: (2025)
SqueezeLLM: Dense-and-Sparse Quantization
por: Kim, Sehoon, et al.
Publicado: (2023)
por: Kim, Sehoon, et al.
Publicado: (2023)
When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
por: Wang, Junxiong, et al.
Publicado: (2026)
por: Wang, Junxiong, et al.
Publicado: (2026)
S*: Test Time Scaling for Code Generation
por: Li, Dacheng, et al.
Publicado: (2025)
por: Li, Dacheng, et al.
Publicado: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
por: Dokmeci, Berkan, et al.
Publicado: (2025)
por: Dokmeci, Berkan, et al.
Publicado: (2025)
Ejemplares similares
-
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
por: Maheswaran, Monishwaran, et al.
Publicado: (2026) -
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
por: Tiwari, Rishabh, et al.
Publicado: (2026) -
Learning, Fast and Slow: Towards LLMs That Adapt Continually
por: Tiwari, Rishabh, et al.
Publicado: (2026) -
Learning Adaptive Parallel Reasoning with Language Models
por: Pan, Jiayi, et al.
Publicado: (2025) -
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
por: Shao, Zelei, et al.
Publicado: (2025)