Evaluating Prompting and Execution-Based Methods for Deterministic Computation in LLMs
Fuente:
arXiv
Salvato in:
| Autore principale: | Yu, Hongkun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method
di: Tripathi, Satvik
Pubblicazione: (2025)
di: Tripathi, Satvik
Pubblicazione: (2025)
Protecting Context and Prompts: Deterministic Security for Non-Deterministic AI
di: Rajagopalan, Mohan, et al.
Pubblicazione: (2026)
di: Rajagopalan, Mohan, et al.
Pubblicazione: (2026)
PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
di: Ravishankara, Mayank
Pubblicazione: (2026)
di: Ravishankara, Mayank
Pubblicazione: (2026)
On the Holographic Geometry of Deterministic Computation
di: Nye, Logan
Pubblicazione: (2025)
di: Nye, Logan
Pubblicazione: (2025)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
di: Hu, Xavier, et al.
Pubblicazione: (2026)
di: Hu, Xavier, et al.
Pubblicazione: (2026)
Scalable Solution Methods for Dec-POMDPs with Deterministic Dynamics
di: You, Yang, et al.
Pubblicazione: (2025)
di: You, Yang, et al.
Pubblicazione: (2025)
LLMs as Method Actors: A Model for Prompt Engineering and Architecture
di: Doyle, Colin
Pubblicazione: (2024)
di: Doyle, Colin
Pubblicazione: (2024)
TP-UNet: Temporal Prompt Guided UNet for Medical Image Segmentation
di: Wang, Ranmin, et al.
Pubblicazione: (2024)
di: Wang, Ranmin, et al.
Pubblicazione: (2024)
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
di: Ezra, Elon, et al.
Pubblicazione: (2025)
di: Ezra, Elon, et al.
Pubblicazione: (2025)
Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
di: Gao, Xingjie, et al.
Pubblicazione: (2026)
di: Gao, Xingjie, et al.
Pubblicazione: (2026)
Deterministic or probabilistic? The psychology of LLMs as random number generators
di: Coronado-Blázquez, Javier
Pubblicazione: (2025)
di: Coronado-Blázquez, Javier
Pubblicazione: (2025)
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
di: Sinha, Akshit, et al.
Pubblicazione: (2025)
di: Sinha, Akshit, et al.
Pubblicazione: (2025)
Executable Governance for AI: Translating Policies into Rules Using LLMs
di: Datla, Gautam Varma, et al.
Pubblicazione: (2025)
di: Datla, Gautam Varma, et al.
Pubblicazione: (2025)
On LLM-generated Logic Programs and their Inference Execution Methods
di: Tarau, Paul
Pubblicazione: (2025)
di: Tarau, Paul
Pubblicazione: (2025)
Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
di: Maveli, Nickil, et al.
Pubblicazione: (2026)
di: Maveli, Nickil, et al.
Pubblicazione: (2026)
The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
di: Samiei, Mahdi, et al.
Pubblicazione: (2025)
di: Samiei, Mahdi, et al.
Pubblicazione: (2025)
How Focused Are LLMs? A Quantitative Study via Repetitive Deterministic Prediction Tasks
di: Hou, Wanda, et al.
Pubblicazione: (2025)
di: Hou, Wanda, et al.
Pubblicazione: (2025)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
di: Hua, Andong, et al.
Pubblicazione: (2025)
di: Hua, Andong, et al.
Pubblicazione: (2025)
From Agent Loops to Deterministic Graphs: Execution Lineage for Reproducible AI-Native Work
di: Rosen, Josh, et al.
Pubblicazione: (2026)
di: Rosen, Josh, et al.
Pubblicazione: (2026)
PromptKeeper: Safeguarding System Prompts for LLMs
di: Jiang, Zhifeng, et al.
Pubblicazione: (2024)
di: Jiang, Zhifeng, et al.
Pubblicazione: (2024)
Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
di: Li, Xinran, et al.
Pubblicazione: (2025)
di: Li, Xinran, et al.
Pubblicazione: (2025)
A Comparison of Prompt Engineering Techniques for Task Planning and Execution in Service Robotics
di: Bode, Jonas, et al.
Pubblicazione: (2024)
di: Bode, Jonas, et al.
Pubblicazione: (2024)
LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
di: Klem, Strahinja, et al.
Pubblicazione: (2025)
di: Klem, Strahinja, et al.
Pubblicazione: (2025)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
di: Seleznyov, Mikhail, et al.
Pubblicazione: (2025)
di: Seleznyov, Mikhail, et al.
Pubblicazione: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
Accelerated AI Inference via Dynamic Execution Methods
di: Barad, Haim, et al.
Pubblicazione: (2024)
di: Barad, Haim, et al.
Pubblicazione: (2024)
Computing the Reachability Value of Posterior-Deterministic POMDPs
di: Fijalkow, Nathanaël, et al.
Pubblicazione: (2026)
di: Fijalkow, Nathanaël, et al.
Pubblicazione: (2026)
Deterministic Computing Power Networking: Architecture, Technologies and Prospects
di: Jia, Qingmin, et al.
Pubblicazione: (2024)
di: Jia, Qingmin, et al.
Pubblicazione: (2024)
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks
di: Bi, Ting, et al.
Pubblicazione: (2025)
di: Bi, Ting, et al.
Pubblicazione: (2025)
Prompt-Based One-Shot Exact Length-Controlled Generation with LLMs
di: Xie, Juncheng, et al.
Pubblicazione: (2025)
di: Xie, Juncheng, et al.
Pubblicazione: (2025)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
di: Ning, Xuefei, et al.
Pubblicazione: (2023)
di: Ning, Xuefei, et al.
Pubblicazione: (2023)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition
di: Bumb, Mayank, et al.
Pubblicazione: (2025)
di: Bumb, Mayank, et al.
Pubblicazione: (2025)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
di: Gehring, Jonas, et al.
Pubblicazione: (2024)
di: Gehring, Jonas, et al.
Pubblicazione: (2024)
Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs
di: Sakharova, Marina, et al.
Pubblicazione: (2025)
di: Sakharova, Marina, et al.
Pubblicazione: (2025)
BadSKP: Backdoor Attacks on Knowledge Graph-Enhanced LLMs with Soft Prompts
di: Lyu, Xiaoting, et al.
Pubblicazione: (2026)
di: Lyu, Xiaoting, et al.
Pubblicazione: (2026)
Think Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured Health Text Generation
di: Cheng, Kai-Chen, et al.
Pubblicazione: (2026)
di: Cheng, Kai-Chen, et al.
Pubblicazione: (2026)
TrustGLM: Evaluating the Robustness of GraphLLMs Against Prompt, Text, and Structure Attacks
di: Zhang, Qihai, et al.
Pubblicazione: (2025)
di: Zhang, Qihai, et al.
Pubblicazione: (2025)
Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs
di: Bouchard, Dylan
Pubblicazione: (2024)
di: Bouchard, Dylan
Pubblicazione: (2024)
Evaluating Code Generation of LLMs in Advanced Computer Science Problems
di: Catir, Emir, et al.
Pubblicazione: (2025)
di: Catir, Emir, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method
di: Tripathi, Satvik
Pubblicazione: (2025) -
Protecting Context and Prompts: Deterministic Security for Non-Deterministic AI
di: Rajagopalan, Mohan, et al.
Pubblicazione: (2026) -
PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
di: Ravishankara, Mayank
Pubblicazione: (2026) -
On the Holographic Geometry of Deterministic Computation
di: Nye, Logan
Pubblicazione: (2025) -
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
di: Hu, Xavier, et al.
Pubblicazione: (2026)