The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
Fuente:
arXiv
Guardado en:
| Autores principales: | Ezra, Elon, Weizman, Ariel, Azaria, Amos |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fool Me, Fool Me: User Attitudes Toward LLM Falsehoods
por: Nirman, Diana Bar-Or, et al.
Publicado: (2024)
por: Nirman, Diana Bar-Or, et al.
Publicado: (2024)
TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages
por: Ofer, Moshe, et al.
Publicado: (2025)
por: Ofer, Moshe, et al.
Publicado: (2025)
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
por: Yang, Yifan, et al.
Publicado: (2026)
por: Yang, Yifan, et al.
Publicado: (2026)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
por: Gehring, Jonas, et al.
Publicado: (2024)
por: Gehring, Jonas, et al.
Publicado: (2024)
Confidence Improves Self-Consistency in LLMs
por: Taubenfeld, Amir, et al.
Publicado: (2025)
por: Taubenfeld, Amir, et al.
Publicado: (2025)
ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects
por: Zhang, Jipeng, et al.
Publicado: (2025)
por: Zhang, Jipeng, et al.
Publicado: (2025)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
por: Hu, Xavier, et al.
Publicado: (2026)
por: Hu, Xavier, et al.
Publicado: (2026)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
por: Dong, Guanting, et al.
Publicado: (2024)
por: Dong, Guanting, et al.
Publicado: (2024)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
por: Tie, Guiyao, et al.
Publicado: (2025)
por: Tie, Guiyao, et al.
Publicado: (2025)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
por: Liu, Yanming, et al.
Publicado: (2026)
por: Liu, Yanming, et al.
Publicado: (2026)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
por: Yan, Weixiang, et al.
Publicado: (2023)
por: Yan, Weixiang, et al.
Publicado: (2023)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
por: Li, Junlong, et al.
Publicado: (2025)
por: Li, Junlong, et al.
Publicado: (2025)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
por: He, Muyu, et al.
Publicado: (2026)
por: He, Muyu, et al.
Publicado: (2026)
Probing the Lack of Stable Internal Beliefs in LLMs
por: Luo, Yifan, et al.
Publicado: (2026)
por: Luo, Yifan, et al.
Publicado: (2026)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
por: Chen, Yuhao, et al.
Publicado: (2026)
por: Chen, Yuhao, et al.
Publicado: (2026)
Benchmarking Real-Time Question Answering via Executable Code Workflows
por: Zhou, Wenjie, et al.
Publicado: (2026)
por: Zhou, Wenjie, et al.
Publicado: (2026)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
por: Xu, Wenda, et al.
Publicado: (2025)
por: Xu, Wenda, et al.
Publicado: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
por: Jung, Dongwon, et al.
Publicado: (2025)
por: Jung, Dongwon, et al.
Publicado: (2025)
Execution-Verified Reinforcement Learning for Optimization Modeling
por: Guan, Runda, et al.
Publicado: (2026)
por: Guan, Runda, et al.
Publicado: (2026)
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
por: Ni, Yuansheng, et al.
Publicado: (2025)
por: Ni, Yuansheng, et al.
Publicado: (2025)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
por: Schnabl, Christoph, et al.
Publicado: (2025)
por: Schnabl, Christoph, et al.
Publicado: (2025)
Gistify! Codebase-Level Understanding via Runtime Execution
por: Lee, Hyunji, et al.
Publicado: (2025)
por: Lee, Hyunji, et al.
Publicado: (2025)
GenerationPrograms: Fine-grained Attribution with Executable Programs
por: Wan, David, et al.
Publicado: (2025)
por: Wan, David, et al.
Publicado: (2025)
Executable Code Actions Elicit Better LLM Agents
por: Wang, Xingyao, et al.
Publicado: (2024)
por: Wang, Xingyao, et al.
Publicado: (2024)
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
por: Wu, Yue, et al.
Publicado: (2023)
por: Wu, Yue, et al.
Publicado: (2023)
DOCE: Finding the Sweet Spot for Execution-Based Code Generation
por: Li, Haau-Sing, et al.
Publicado: (2024)
por: Li, Haau-Sing, et al.
Publicado: (2024)
COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following
por: Bhar, Swarnadeep, et al.
Publicado: (2025)
por: Bhar, Swarnadeep, et al.
Publicado: (2025)
Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution
por: Cifani, Susanna, et al.
Publicado: (2026)
por: Cifani, Susanna, et al.
Publicado: (2026)
PatchWorld: Gradient-Free Optimization of Executable World Models
por: Bai, Jiaxin, et al.
Publicado: (2026)
por: Bai, Jiaxin, et al.
Publicado: (2026)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
por: Phute, Mansi, et al.
Publicado: (2023)
por: Phute, Mansi, et al.
Publicado: (2023)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
por: Zhang, Xiaoying, et al.
Publicado: (2024)
por: Zhang, Xiaoying, et al.
Publicado: (2024)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
por: Sivakumaran, Nithin, et al.
Publicado: (2026)
por: Sivakumaran, Nithin, et al.
Publicado: (2026)
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
por: Pang, Renning, et al.
Publicado: (2026)
por: Pang, Renning, et al.
Publicado: (2026)
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
por: Zheng, Xin, et al.
Publicado: (2024)
por: Zheng, Xin, et al.
Publicado: (2024)
Self-Evolved Reward Learning for LLMs
por: Huang, Chenghua, et al.
Publicado: (2024)
por: Huang, Chenghua, et al.
Publicado: (2024)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
por: Kumar, Adarsh, et al.
Publicado: (2025)
por: Kumar, Adarsh, et al.
Publicado: (2025)
Self-controller: Controlling LLMs with Multi-round Step-by-step Self-awareness
por: Peng, Xiao, et al.
Publicado: (2024)
por: Peng, Xiao, et al.
Publicado: (2024)
Towards Execution-Grounded Automated AI Research
por: Si, Chenglei, et al.
Publicado: (2026)
por: Si, Chenglei, et al.
Publicado: (2026)
Evolving and Executing Research Plans via Double-Loop Multi-Agent Collaboration
por: Zhang, Zhi, et al.
Publicado: (2025)
por: Zhang, Zhi, et al.
Publicado: (2025)
Beyond the Speculative Game: A Survey of Speculative Execution in Large Language Models
por: Zhang, Chen, et al.
Publicado: (2024)
por: Zhang, Chen, et al.
Publicado: (2024)
Ejemplares similares
-
Fool Me, Fool Me: User Attitudes Toward LLM Falsehoods
por: Nirman, Diana Bar-Or, et al.
Publicado: (2024) -
TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages
por: Ofer, Moshe, et al.
Publicado: (2025) -
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
por: Yang, Yifan, et al.
Publicado: (2026) -
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
por: Gehring, Jonas, et al.
Publicado: (2024) -
Confidence Improves Self-Consistency in LLMs
por: Taubenfeld, Amir, et al.
Publicado: (2025)