EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines
Fuente:
arXiv
Guardado en:
| Autores principales: | van der Vleuten, Noah, Flores, Anthony, Mathur, Shray, Rakitin, Max, Hopkins, Thomas, Yager, Kevin G., Tsai, Esther H. R. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VISION: A Modular AI Assistant for Natural Human-Instrument Interaction at Scientific User Facilities
por: Mathur, Shray, et al.
Publicado: (2024)
por: Mathur, Shray, et al.
Publicado: (2024)
LLM Jaggedness Unlocks Scientific Creativity
por: Mathur, Shray, et al.
Publicado: (2026)
por: Mathur, Shray, et al.
Publicado: (2026)
Dr. Boot: Bootstrapping Program Synthesis Language Models to Perform Repairing
por: van der Vleuten, Noah
Publicado: (2025)
por: van der Vleuten, Noah
Publicado: (2025)
A General Bayesian Algorithm for the Autonomous Alignment of Beamlines
por: Morris, T. W., et al.
Publicado: (2024)
por: Morris, T. W., et al.
Publicado: (2024)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
por: Abdollahi, Mohammad, et al.
Publicado: (2025)
por: Abdollahi, Mohammad, et al.
Publicado: (2025)
La inspección del trabajo en la URSS
por: G. Rakitin
Publicado: (1971)
por: G. Rakitin
Publicado: (1971)
Labour inspection in the USSR
por: G. Rakitin
Publicado: (1971)
por: G. Rakitin
Publicado: (1971)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
por: Wang, Yubang, et al.
Publicado: (2026)
por: Wang, Yubang, et al.
Publicado: (2026)
Predictive Monitoring with Strong Trace Prefixes
por: Ang, Zhendong, et al.
Publicado: (2024)
por: Ang, Zhendong, et al.
Publicado: (2024)
Teaching LLMs Program Semantics via Symbolic Execution Traces
por: Bayer, Jonas, et al.
Publicado: (2026)
por: Bayer, Jonas, et al.
Publicado: (2026)
Domain-specific ChatBots for Science using Embeddings
por: Yager, Kevin G.
Publicado: (2023)
por: Yager, Kevin G.
Publicado: (2023)
Towards a Science Exocortex
por: Yager, Kevin G.
Publicado: (2024)
por: Yager, Kevin G.
Publicado: (2024)
Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Models
por: Wang, Jian, et al.
Publicado: (2025)
por: Wang, Jian, et al.
Publicado: (2025)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
por: Haque, Mirazul, et al.
Publicado: (2025)
por: Haque, Mirazul, et al.
Publicado: (2025)
Cash‐for‐Care and the Cost of Parenthood: Evidence From Same‐Sex and Adoptive Parents
por: Maaike van der Vleuten, et al.
Publicado: (2025)
por: Maaike van der Vleuten, et al.
Publicado: (2025)
X‐Ray Fluorescence Methods and Performance of the Tender‐Energy Spectroscopy Beamline at Shanghai Synchrotron Radiation Facility
por: Lingling Guo, et al.
Publicado: (2026)
por: Lingling Guo, et al.
Publicado: (2026)
Quark-Mass Dependence of Light-Nuclei Masses from Lattice QCD and Trace-Anomaly Contributions to Nuclear Bindings
por: Chakraborty, Debsubhra, et al.
Publicado: (2026)
por: Chakraborty, Debsubhra, et al.
Publicado: (2026)
Counting and Sampling Traces in Regular Languages
por: de Colnet, Alexis, et al.
Publicado: (2025)
por: de Colnet, Alexis, et al.
Publicado: (2025)
XAI for Coding Agent Failures: Transforming Raw Execution Traces into Actionable Insights
por: Joshi, Arun
Publicado: (2026)
por: Joshi, Arun
Publicado: (2026)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Traces of Social Competence in Large Language Models
por: Kouwenhoven, Tom, et al.
Publicado: (2026)
por: Kouwenhoven, Tom, et al.
Publicado: (2026)
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
por: Jiang, Xue, et al.
Publicado: (2025)
por: Jiang, Xue, et al.
Publicado: (2025)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
por: Cheng, Ching-An, et al.
Publicado: (2024)
por: Cheng, Ching-An, et al.
Publicado: (2024)
Action-Attentive Deep Reinforcement Learning for Autonomous Alignment of Beamlines
por: Wang, Siyu, et al.
Publicado: (2024)
por: Wang, Siyu, et al.
Publicado: (2024)
Tracing the Jerusalem Code
Publicado: (2021)
Publicado: (2021)
Tracing the Jerusalem Code
Publicado: (2021)
Publicado: (2021)
Divisibility of Trace Codes
por: Huang, Hexiang, et al.
Publicado: (2026)
por: Huang, Hexiang, et al.
Publicado: (2026)
Tracing the Jerusalem Code
Publicado: (2021)
Publicado: (2021)
Executable Boundary Contracts for Sound Event Traces
por: Alpay, Faruk, et al.
Publicado: (2026)
por: Alpay, Faruk, et al.
Publicado: (2026)
Simple Fault Localization using Execution Traces
por: Prenner, Julian Aron, et al.
Publicado: (2025)
por: Prenner, Julian Aron, et al.
Publicado: (2025)
Computing Alignments for Partially-ordered Traces Through Petri Net Unfoldings
por: Siddiqui, Ariba, et al.
Publicado: (2025)
por: Siddiqui, Ariba, et al.
Publicado: (2025)
TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces
por: Yang, Shu-Xun, et al.
Publicado: (2026)
por: Yang, Shu-Xun, et al.
Publicado: (2026)
The Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution Traces
por: Sahoo, Subramanyam
Publicado: (2025)
por: Sahoo, Subramanyam
Publicado: (2025)
Non-Signaling Locality Lower Bounds for Dominating Set
por: Fleming, Noah, et al.
Publicado: (2026)
por: Fleming, Noah, et al.
Publicado: (2026)
Compositionality in Coalgebraic Trace Semantics
por: Jourde, Robin, et al.
Publicado: (2026)
por: Jourde, Robin, et al.
Publicado: (2026)
The MUSE Beamline Calorimeter
por: Lin, W., et al.
Publicado: (2024)
por: Lin, W., et al.
Publicado: (2024)
Dynamic Exclusion of Low-Fidelity Data in Bayesian Optimization for Autonomous Beamline Alignment
por: Narayanan, Megha R., et al.
Publicado: (2024)
por: Narayanan, Megha R., et al.
Publicado: (2024)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
por: Lu, Yu-An, et al.
Publicado: (2026)
por: Lu, Yu-An, et al.
Publicado: (2026)
Offset Finding of Beamline Parameters on the METRIXS Beamline at BESSY II Using Machine Learning
por: Meier, David, et al.
Publicado: (2025)
por: Meier, David, et al.
Publicado: (2025)
Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
por: Bhatia, Gagan, et al.
Publicado: (2025)
por: Bhatia, Gagan, et al.
Publicado: (2025)
Ejemplares similares
-
VISION: A Modular AI Assistant for Natural Human-Instrument Interaction at Scientific User Facilities
por: Mathur, Shray, et al.
Publicado: (2024) -
LLM Jaggedness Unlocks Scientific Creativity
por: Mathur, Shray, et al.
Publicado: (2026) -
Dr. Boot: Bootstrapping Program Synthesis Language Models to Perform Repairing
por: van der Vleuten, Noah
Publicado: (2025) -
A General Bayesian Algorithm for the Autonomous Alignment of Beamlines
por: Morris, T. W., et al.
Publicado: (2024) -
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
por: Abdollahi, Mohammad, et al.
Publicado: (2025)