Revisiting Tree Search for LLMs: Gumbel and Sequential Halving for Budget-Scalable Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ugadiarov, Leonid, Kuratov, Yuri, Panov, Aleksandr, Skrynnik, Alexey |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Object-Centric World Models Meet Monte Carlo Tree Search
por: Vakhitov, Rodion, et al.
Publicado: (2026)
por: Vakhitov, Rodion, et al.
Publicado: (2026)
Relational Object-Centric Actor-Critic
por: Ugadiarov, Leonid, et al.
Publicado: (2023)
por: Ugadiarov, Leonid, et al.
Publicado: (2023)
CAMAR: Continuous Actions Multi-Agent Routing
por: Pshenitsyn, Artem, et al.
Publicado: (2025)
por: Pshenitsyn, Artem, et al.
Publicado: (2025)
Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding
por: Vyaltsev, Valeriy, et al.
Publicado: (2026)
por: Vyaltsev, Valeriy, et al.
Publicado: (2026)
Advancing Learnable Multi-Agent Pathfinding Solvers with Active Fine-Tuning
por: Andreychuk, Anton, et al.
Publicado: (2025)
por: Andreychuk, Anton, et al.
Publicado: (2025)
MAPF-GPT: Imitation Learning for Multi-Agent Pathfinding at Scale
por: Andreychuk, Anton, et al.
Publicado: (2024)
por: Andreychuk, Anton, et al.
Publicado: (2024)
Anytime Sequential Halving in Monte-Carlo Tree Search
por: Sagers, Dominic, et al.
Publicado: (2024)
por: Sagers, Dominic, et al.
Publicado: (2024)
POGEMA: A Benchmark Platform for Cooperative Multi-Agent Pathfinding
por: Skrynnik, Alexey, et al.
Publicado: (2024)
por: Skrynnik, Alexey, et al.
Publicado: (2024)
Recurrent Action Transformer with Memory
por: Cherepanov, Egor, et al.
Publicado: (2023)
por: Cherepanov, Egor, et al.
Publicado: (2023)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
por: Kuratov, Yuri, et al.
Publicado: (2024)
por: Kuratov, Yuri, et al.
Publicado: (2024)
Instruction Following with Goal-Conditioned Reinforcement Learning in Virtual Environments
por: Volovikova, Zoya, et al.
Publicado: (2024)
por: Volovikova, Zoya, et al.
Publicado: (2024)
ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
por: Cherepanov, Egor, et al.
Publicado: (2025)
por: Cherepanov, Egor, et al.
Publicado: (2025)
Memory Retention Is Not Enough to Master Memory Tasks in Reinforcement Learning
por: Shchendrigin, Oleg, et al.
Publicado: (2026)
por: Shchendrigin, Oleg, et al.
Publicado: (2026)
Re:Frame -- Retrieving Experience From Associative Memory
por: Zelezetsky, Daniil, et al.
Publicado: (2025)
por: Zelezetsky, Daniil, et al.
Publicado: (2025)
LookPlanGraph: Embodied Instruction Following Method with VLM Graph Augmentation
por: Onishchenko, Anatoly O., et al.
Publicado: (2025)
por: Onishchenko, Anatoly O., et al.
Publicado: (2025)
Gradual Optimization Learning for Conformational Energy Minimization
por: Tsypin, Artem, et al.
Publicado: (2023)
por: Tsypin, Artem, et al.
Publicado: (2023)
A data-driven approach to modeling brain activity using differential equations
por: Andrey, Kuratov
Publicado: (2024)
por: Andrey, Kuratov
Publicado: (2024)
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
por: Cherepanov, Egor, et al.
Publicado: (2025)
por: Cherepanov, Egor, et al.
Publicado: (2025)
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation
por: Cherepanov, Egor, et al.
Publicado: (2024)
por: Cherepanov, Egor, et al.
Publicado: (2024)
Self-Guided Plan Extraction for Instruction-Following Tasks with Goal-Conditional Reinforcement Learning
por: Volovikova, Zoya, et al.
Publicado: (2026)
por: Volovikova, Zoya, et al.
Publicado: (2026)
Symbolic Disentangled Representations for Images
por: Korchemnyi, Alexandr, et al.
Publicado: (2024)
por: Korchemnyi, Alexandr, et al.
Publicado: (2024)
CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended World
por: Volovikova, Zoya, et al.
Publicado: (2025)
por: Volovikova, Zoya, et al.
Publicado: (2025)
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
por: Chepurova, Alla, et al.
Publicado: (2025)
por: Chepurova, Alla, et al.
Publicado: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
por: Bulatov, Aydar, et al.
Publicado: (2023)
por: Bulatov, Aydar, et al.
Publicado: (2023)
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding
por: Sagirova, Alsu, et al.
Publicado: (2025)
por: Sagirova, Alsu, et al.
Publicado: (2025)
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
por: Kachaev, Nikita, et al.
Publicado: (2025)
por: Kachaev, Nikita, et al.
Publicado: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
por: Kachaev, Nikita, et al.
Publicado: (2025)
por: Kachaev, Nikita, et al.
Publicado: (2025)
Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
por: Li, Yushu, et al.
Publicado: (2026)
por: Li, Yushu, et al.
Publicado: (2026)
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
por: Cherepanov, Egor, et al.
Publicado: (2026)
por: Cherepanov, Egor, et al.
Publicado: (2026)
Object-Centric Learning with Slot Mixture Module
por: Kirilenko, Daniil, et al.
Publicado: (2023)
por: Kirilenko, Daniil, et al.
Publicado: (2023)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
por: Cinquin, Tristan, et al.
Publicado: (2025)
por: Cinquin, Tristan, et al.
Publicado: (2025)
AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment
por: Ivanova, Anastasiia, et al.
Publicado: (2025)
por: Ivanova, Anastasiia, et al.
Publicado: (2025)
HELP: Hierarchical Embodied Language Planner for Household Tasks
por: Korchemnyi, Alexandr V., et al.
Publicado: (2025)
por: Korchemnyi, Alexandr V., et al.
Publicado: (2025)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
por: Miyamoto, Sora, et al.
Publicado: (2026)
por: Miyamoto, Sora, et al.
Publicado: (2026)
Associative Recurrent Memory Transformer
por: Rodkin, Ivan, et al.
Publicado: (2024)
por: Rodkin, Ivan, et al.
Publicado: (2024)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
por: Alomrani, Mohammad Ali, et al.
Publicado: (2025)
por: Alomrani, Mohammad Ali, et al.
Publicado: (2025)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
por: Rubashevskii, Aleksandr, et al.
Publicado: (2026)
por: Rubashevskii, Aleksandr, et al.
Publicado: (2026)
Learning Successor Features with Distributed Hebbian Temporal Memory
por: Dzhivelikian, Evgenii, et al.
Publicado: (2023)
por: Dzhivelikian, Evgenii, et al.
Publicado: (2023)
Twice Sequential Monte Carlo for Tree Search
por: Oren, Yaniv, et al.
Publicado: (2025)
por: Oren, Yaniv, et al.
Publicado: (2025)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
por: Rodkin, Ivan, et al.
Publicado: (2025)
por: Rodkin, Ivan, et al.
Publicado: (2025)
Ejemplares similares
-
Object-Centric World Models Meet Monte Carlo Tree Search
por: Vakhitov, Rodion, et al.
Publicado: (2026) -
Relational Object-Centric Actor-Critic
por: Ugadiarov, Leonid, et al.
Publicado: (2023) -
CAMAR: Continuous Actions Multi-Agent Routing
por: Pshenitsyn, Artem, et al.
Publicado: (2025) -
Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding
por: Vyaltsev, Valeriy, et al.
Publicado: (2026) -
Advancing Learnable Multi-Agent Pathfinding Solvers with Active Fine-Tuning
por: Andreychuk, Anton, et al.
Publicado: (2025)