Code Simulation as a Proxy for High-order Tasks in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | La Malfa, Emanuele, Weinhuber, Christoph, Torre, Orazio, Lin, Fangru, Huang, X. Angelo, Marro, Samuele, Cohn, Anthony, Shadbolt, Nigel, Wooldridge, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Code Simulation Challenges for Large Language Models
by: La Malfa, Emanuele, et al.
Published: (2024)
by: La Malfa, Emanuele, et al.
Published: (2024)
A Scalable Communication Protocol for Networks of Large Language Models
by: Marro, Samuele, et al.
Published: (2024)
by: Marro, Samuele, et al.
Published: (2024)
A Notion of Complexity for Theory of Mind via Discrete World Models
by: Huang, X. Angelo, et al.
Published: (2024)
by: Huang, X. Angelo, et al.
Published: (2024)
Jailbreaking Large Language Models in Infinitely Many Ways
by: Goldstein, Oliver, et al.
Published: (2025)
by: Goldstein, Oliver, et al.
Published: (2025)
Language Models Are Implicitly Continuous
by: Marro, Samuele, et al.
Published: (2025)
by: Marro, Samuele, et al.
Published: (2025)
End-to-end PDDL Planning with Hardcoded and Dynamic Agents
by: La Malfa, Emanuele, et al.
Published: (2025)
by: La Malfa, Emanuele, et al.
Published: (2025)
Large Language Models Miss the Multi-Agent Mark
by: La Malfa, Emanuele, et al.
Published: (2025)
by: La Malfa, Emanuele, et al.
Published: (2025)
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning
by: Lin, Fangru, et al.
Published: (2024)
by: Lin, Fangru, et al.
Published: (2024)
Tacit Coordination of Large Language Models
by: Aharon, Ido, et al.
Published: (2026)
by: Aharon, Ido, et al.
Published: (2026)
Out-of-Context Reasoning in Large Language Models
by: Shaki, Jonathan, et al.
Published: (2025)
by: Shaki, Jonathan, et al.
Published: (2025)
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks
by: Lin, Fangru, et al.
Published: (2024)
by: Lin, Fangru, et al.
Published: (2024)
Fixed Point Explainability
by: La Malfa, Emanuele, et al.
Published: (2025)
by: La Malfa, Emanuele, et al.
Published: (2025)
Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
by: Huang, Xuanqiang Angelo, et al.
Published: (2026)
by: Huang, Xuanqiang Angelo, et al.
Published: (2026)
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
by: La Malfa, Gabriele, et al.
Published: (2026)
by: La Malfa, Gabriele, et al.
Published: (2026)
Deep Neural Networks via Complex Network Theory: a Perspective
by: La Malfa, Emanuele, et al.
Published: (2024)
by: La Malfa, Emanuele, et al.
Published: (2024)
LLM Agents Are the Antidote to Walled Gardens
by: Marro, Samuele, et al.
Published: (2025)
by: Marro, Samuele, et al.
Published: (2025)
Fetch.ai: An Architecture for Modern Multi-Agent Systems
by: Wooldridge, Michael J., et al.
Published: (2025)
by: Wooldridge, Michael J., et al.
Published: (2025)
Can Large Language Models Generalize Procedures Across Representations?
by: Lin, Fangru, et al.
Published: (2026)
by: Lin, Fangru, et al.
Published: (2026)
From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Profit is the Red Team: Stress-Testing Agents in Strategic Economic Interactions
by: Wang, Shouqiao, et al.
Published: (2026)
by: Wang, Shouqiao, et al.
Published: (2026)
How Subradiance Enables Nonlinearity in Weakly Driven Quantum Arrays
by: Scarlatella, Orazio, et al.
Published: (2024)
by: Scarlatella, Orazio, et al.
Published: (2024)
Fate of the Mollow triplet in strongly-coupled atomic arrays
by: Scarlatella, Orazio, et al.
Published: (2024)
by: Scarlatella, Orazio, et al.
Published: (2024)
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
The Collaboration Gap in Human-AI Work
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
"Diversity is Having the Diversity": Unpacking and Designing for Diversity in Applicant Selection
by: Natarajan, Neil, et al.
Published: (2024)
by: Natarajan, Neil, et al.
Published: (2024)
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
From Rights to Rites: Expectations Management in Smart-Home AI
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
by: Duan, Zhangqi, et al.
Published: (2026)
by: Duan, Zhangqi, et al.
Published: (2026)
Can Large Language Models Reason about the Region Connection Calculus?
by: Cohn, Anthony G, et al.
Published: (2024)
by: Cohn, Anthony G, et al.
Published: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
by: Cohn, Anthony G, et al.
Published: (2024)
by: Cohn, Anthony G, et al.
Published: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
by: Cohn, Anthony G, et al.
Published: (2025)
by: Cohn, Anthony G, et al.
Published: (2025)
Cognitive Effects in Large Language Models
by: Shaki, Jonathan, et al.
Published: (2023)
by: Shaki, Jonathan, et al.
Published: (2023)
When Claims Evolve: Evaluating and Enhancing the Robustness of Embedding Models Against Misinformation Edits
by: Magomere, Jabez, et al.
Published: (2025)
by: Magomere, Jabez, et al.
Published: (2025)
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics
by: Lin, Fangru, et al.
Published: (2024)
by: Lin, Fangru, et al.
Published: (2024)
Respectful Things: Adding Social Intelligence to 'Smart' Devices
by: Van Kleek, Max, et al.
Published: (2026)
by: Van Kleek, Max, et al.
Published: (2026)
An actionable framework for AI‐ready data
by: Neil Majithia, et al.
Published: (2026)
by: Neil Majithia, et al.
Published: (2026)
When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding
by: Yang, Xu, et al.
Published: (2026)
by: Yang, Xu, et al.
Published: (2026)
Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions
by: Abate, Alessandro, et al.
Published: (2026)
by: Abate, Alessandro, et al.
Published: (2026)
Similar Items
-
Code Simulation Challenges for Large Language Models
by: La Malfa, Emanuele, et al.
Published: (2024) -
A Scalable Communication Protocol for Networks of Large Language Models
by: Marro, Samuele, et al.
Published: (2024) -
A Notion of Complexity for Theory of Mind via Discrete World Models
by: Huang, X. Angelo, et al.
Published: (2024) -
Jailbreaking Large Language Models in Infinitely Many Ways
by: Goldstein, Oliver, et al.
Published: (2025) -
Language Models Are Implicitly Continuous
by: Marro, Samuele, et al.
Published: (2025)