LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
Fuente:
arXiv
Salvato in:
| Autori principali: | Ziomek, Juliusz, Bankes, William, Wolf, Lorenz, Ramesh, Shyam Sundhar, Tang, Xiaohang, Bogunovic, Ilija |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robust Multi-Objective Controlled Decoding of Large Language Models
di: Son, Seongho, et al.
Pubblicazione: (2025)
di: Son, Seongho, et al.
Pubblicazione: (2025)
Mean-Field Bayesian Optimisation
di: Steinberg, Petar, et al.
Pubblicazione: (2025)
di: Steinberg, Petar, et al.
Pubblicazione: (2025)
Textual understanding boost in the WikiRace
di: Ebrahimi, Raman, et al.
Pubblicazione: (2025)
di: Ebrahimi, Raman, et al.
Pubblicazione: (2025)
This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2023)
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2023)
REDUCR: Robust Data Downsampling Using Class Priority Reweighting
di: Bankes, William, et al.
Pubblicazione: (2023)
di: Bankes, William, et al.
Pubblicazione: (2023)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
di: Ji, Xiaotong, et al.
Pubblicazione: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2026)
Group Robust Preference Optimization in Reward-free RLHF
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2024)
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2024)
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift
di: Son, Seongho, et al.
Pubblicazione: (2024)
di: Son, Seongho, et al.
Pubblicazione: (2024)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
Adversarially Robust Decision Transformer
di: Tang, Xiaohang, et al.
Pubblicazione: (2024)
di: Tang, Xiaohang, et al.
Pubblicazione: (2024)
Open-Ended Task Discovery via Bayesian Optimization
di: Adachi, Masaki, et al.
Pubblicazione: (2026)
di: Adachi, Masaki, et al.
Pubblicazione: (2026)
Time-Varying Gaussian Process Bandits with Unknown Prior
di: Ziomek, Juliusz, et al.
Pubblicazione: (2024)
di: Ziomek, Juliusz, et al.
Pubblicazione: (2024)
Bayesian Optimisation with Unknown Hyperparameters: Regret Bounds Logarithmically Closer to Optimal
di: Ziomek, Juliusz, et al.
Pubblicazione: (2024)
di: Ziomek, Juliusz, et al.
Pubblicazione: (2024)
Just One Layer Norm Guarantees Stable Extrapolation
di: Ziomek, Juliusz, et al.
Pubblicazione: (2025)
di: Ziomek, Juliusz, et al.
Pubblicazione: (2025)
Why Can Large Language Models Generate Correct Chain-of-Thoughts?
di: Tutunov, Rasul, et al.
Pubblicazione: (2023)
di: Tutunov, Rasul, et al.
Pubblicazione: (2023)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
di: Hou, Yufang, et al.
Pubblicazione: (2024)
di: Hou, Yufang, et al.
Pubblicazione: (2024)
Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation
di: Whittle, George, et al.
Pubblicazione: (2025)
di: Whittle, George, et al.
Pubblicazione: (2025)
RSPO: Regularized Self-Play Alignment of Large Language Models
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
Sample-efficient Bayesian Optimisation Using Known Invariances
di: Brown, Theodore, et al.
Pubblicazione: (2024)
di: Brown, Theodore, et al.
Pubblicazione: (2024)
Robust Bayesian Optimisation with Unbounded Corruptions
di: Ezzerg, Abdelhamid, et al.
Pubblicazione: (2025)
di: Ezzerg, Abdelhamid, et al.
Pubblicazione: (2025)
Overton Pluralistic Reinforcement Learning for Large Language Models
di: Fu, Yu, et al.
Pubblicazione: (2026)
di: Fu, Yu, et al.
Pubblicazione: (2026)
Sample-Efficient Regret-Minimizing Double Oracle in Extensive-Form Games
di: Tang, Xiaohang, et al.
Pubblicazione: (2024)
di: Tang, Xiaohang, et al.
Pubblicazione: (2024)
GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?
di: Ko, Dayoon, et al.
Pubblicazione: (2024)
di: Ko, Dayoon, et al.
Pubblicazione: (2024)
Canonical Regularisation of Wide Feature-Learning Neural Networks
di: Whittle, George, et al.
Pubblicazione: (2026)
di: Whittle, George, et al.
Pubblicazione: (2026)
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
di: Tang, Xiaohang, et al.
Pubblicazione: (2026)
di: Tang, Xiaohang, et al.
Pubblicazione: (2026)
EvoWiki: Evaluating LLMs on Evolving Knowledge
di: Tang, Wei, et al.
Pubblicazione: (2024)
di: Tang, Wei, et al.
Pubblicazione: (2024)
Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
di: Güzel, Ahmet H., et al.
Pubblicazione: (2025)
di: Güzel, Ahmet H., et al.
Pubblicazione: (2025)
PROWL: Prioritized Regret-Driven Optimization for World Model Learning
di: Güzel, Ahmet H., et al.
Pubblicazione: (2026)
di: Güzel, Ahmet H., et al.
Pubblicazione: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
di: He, Bingxiang, et al.
Pubblicazione: (2026)
di: He, Bingxiang, et al.
Pubblicazione: (2026)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
di: Li, Yifei, et al.
Pubblicazione: (2025)
di: Li, Yifei, et al.
Pubblicazione: (2025)
WiCER: Wiki-memory Compile, Evaluate, Refine Iterative Knowledge Compilation for LLM Wiki Systems
di: Huerta, Juan M.
Pubblicazione: (2026)
di: Huerta, Juan M.
Pubblicazione: (2026)
Financial Named Entity Recognition: How Far Can LLM Go?
di: Lu, Yi-Te, et al.
Pubblicazione: (2025)
di: Lu, Yi-Te, et al.
Pubblicazione: (2025)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
di: Xin, Yuan, et al.
Pubblicazione: (2025)
di: Xin, Yuan, et al.
Pubblicazione: (2025)
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
di: Yang, Luyu, et al.
Pubblicazione: (2026)
di: Yang, Luyu, et al.
Pubblicazione: (2026)
How Theory-laden are Observations of Black Holes?
di: Doboszewski, Juliusz, et al.
Pubblicazione: (2025)
di: Doboszewski, Juliusz, et al.
Pubblicazione: (2025)
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
di: Lugoloobi, William, et al.
Pubblicazione: (2026)
di: Lugoloobi, William, et al.
Pubblicazione: (2026)
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
di: Kambhampati, Subbarao, et al.
Pubblicazione: (2024)
di: Kambhampati, Subbarao, et al.
Pubblicazione: (2024)
Design and Structural Validation of a Micro-UAV with On-Board Dynamic Route Planning
di: Ravikumar, Inbazhagan, et al.
Pubblicazione: (2025)
di: Ravikumar, Inbazhagan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Robust Multi-Objective Controlled Decoding of Large Language Models
di: Son, Seongho, et al.
Pubblicazione: (2025) -
Mean-Field Bayesian Optimisation
di: Steinberg, Petar, et al.
Pubblicazione: (2025) -
Textual understanding boost in the WikiRace
di: Ebrahimi, Raman, et al.
Pubblicazione: (2025) -
This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs
di: Wolf, Lorenz, et al.
Pubblicazione: (2025) -
Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2023)