Formally Specifying the High-Level Behavior of LLM-Based Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Crouse, Maxwell, Abdelaziz, Ibrahim, Astudillo, Ramon, Basu, Kinjal, Dan, Soham, Kumaravel, Sadhana, Fokoue, Achille, Kapanipathi, Pavan, Roukos, Salim, Lastras, Luis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
von: Crouse, Maxwell, et al.
Veröffentlicht: (2026)
von: Crouse, Maxwell, et al.
Veröffentlicht: (2026)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025)
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025)
R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory
von: Huang, Tenghao, et al.
Veröffentlicht: (2025)
von: Huang, Tenghao, et al.
Veröffentlicht: (2025)
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks
von: Abdelaziz, Ibrahim, et al.
Veröffentlicht: (2024)
von: Abdelaziz, Ibrahim, et al.
Veröffentlicht: (2024)
Putting It All into Context: Simplifying Agents with LCLMs
von: Jiang, Mingjian, et al.
Veröffentlicht: (2025)
von: Jiang, Mingjian, et al.
Veröffentlicht: (2025)
Few-shot Policy (de)composition in Conversational Question Answering
von: Erwin, Kyle, et al.
Veröffentlicht: (2025)
von: Erwin, Kyle, et al.
Veröffentlicht: (2025)
Compositional Program Generation for Few-Shot Systematic Generalization
von: Klinger, Tim, et al.
Veröffentlicht: (2023)
von: Klinger, Tim, et al.
Veröffentlicht: (2023)
Efficient Embedding-based Synthetic Data Generation for Complex Reasoning Tasks
von: Jayaraman, Srideepika, et al.
Veröffentlicht: (2026)
von: Jayaraman, Srideepika, et al.
Veröffentlicht: (2026)
Optimal Policy Minimum Bayesian Risk
von: Astudillo, Ramón Fernandez, et al.
Veröffentlicht: (2025)
von: Astudillo, Ramón Fernandez, et al.
Veröffentlicht: (2025)
SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Budget-Sensitive Discovery Scoring: A Formally Verified Framework for Evaluating AI-Guided Scientific Selection
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
LongFuncEval: Measuring the effectiveness of long context models for function calling
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
CLAPNQ: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
von: Rosenthal, Sara, et al.
Veröffentlicht: (2024)
von: Rosenthal, Sara, et al.
Veröffentlicht: (2024)
Self-Refinement of Language Models from External Proxy Metrics Feedback
von: Ramji, Keshav, et al.
Veröffentlicht: (2024)
von: Ramji, Keshav, et al.
Veröffentlicht: (2024)
Prism: A Minimal Compositional Metalanguage for Specifying Agent Behavior
von: Binard, Franck, et al.
Veröffentlicht: (2025)
von: Binard, Franck, et al.
Veröffentlicht: (2025)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
Transcendental Regularization of Finite Mixtures:Theoretical Guarantees and Practical Limitations
von: Fokoué, Ernest
Veröffentlicht: (2026)
von: Fokoué, Ernest
Veröffentlicht: (2026)
On Fibonacci Ensembles: An Alternative Approach to Ensemble Learning Inspired by the Timeless Architecture of the Golden Ratio
von: Fokoué, Ernest
Veröffentlicht: (2025)
von: Fokoué, Ernest
Veröffentlicht: (2025)
Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity
von: Fokoué, Ernest
Veröffentlicht: (2026)
von: Fokoué, Ernest
Veröffentlicht: (2026)
Causality as the Statistical Conscience of Artificial Intelligence: From Pearl's Ladder to Trustworthy Machines
von: Fokoué, Ernest
Veröffentlicht: (2026)
von: Fokoué, Ernest
Veröffentlicht: (2026)
Fibonacci-Driven Recursive Ensembles: Algorithms, Convergence, and Learning Dynamics
von: Fokoué, Ernest
Veröffentlicht: (2026)
von: Fokoué, Ernest
Veröffentlicht: (2026)
A General Weighting Theory for Ensemble Learning: Beyond Variance Reduction via Spectral and Geometric Structure
von: Fokoué, Ernest
Veröffentlicht: (2025)
von: Fokoué, Ernest
Veröffentlicht: (2025)
No Intelligence Without Statistics: The Invisible Backbone of Artificial Intelligence
von: Fokoué, Ernest
Veröffentlicht: (2025)
von: Fokoué, Ernest
Veröffentlicht: (2025)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
von: Jiang, Mingjian, et al.
Veröffentlicht: (2024)
von: Jiang, Mingjian, et al.
Veröffentlicht: (2024)
Tragico varamento masivo en Luisiana confirma la presencia de tortugas Kempis. / Deborah Crouse
von: Crouse, Deborah
Veröffentlicht: (1993)
von: Crouse, Deborah
Veröffentlicht: (1993)
On the Nature of Discrete Space-Time Part 2: Special Relativity in Discrete Space-Time
von: Crouse, David
Veröffentlicht: (2024)
von: Crouse, David
Veröffentlicht: (2024)
Textbooks 101: Textbook Collection at the University of Minnesota
von: Crouse, Caroline
Veröffentlicht: (2007)
von: Crouse, Caroline
Veröffentlicht: (2007)
A Formal Framework for Naturally Specifying and Verifying Sequential Algorithms
von: Yang, Chengxi, et al.
Veröffentlicht: (2025)
von: Yang, Chengxi, et al.
Veröffentlicht: (2025)
Slope instability along the northeastern Iberian and Balearic continental margins
von: G. Lastras
Veröffentlicht: (2007)
von: G. Lastras
Veröffentlicht: (2007)
Accelerated Reeds-Shepp and Under-Specified Reeds-Shepp Algorithms for Mobile Robot Path Planning
von: Ibrahim, Ibrahim, et al.
Veröffentlicht: (2025)
von: Ibrahim, Ibrahim, et al.
Veröffentlicht: (2025)
MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
von: Phan, Peter, et al.
Veröffentlicht: (2025)
von: Phan, Peter, et al.
Veröffentlicht: (2025)
An analysis of project provenance through a novel method of framework analysis
von: Larson, T., et al.
Veröffentlicht: (2025)
von: Larson, T., et al.
Veröffentlicht: (2025)
Core Curriculum Project (September 15, 1986-September 14, 1988).
von: Crouse, Joan M.
Veröffentlicht: (1988)
von: Crouse, Joan M.
Veröffentlicht: (1988)
Specifying Agent Ethics (Blue Sky Ideas)
von: Dennis, Louise A., et al.
Veröffentlicht: (2024)
von: Dennis, Louise A., et al.
Veröffentlicht: (2024)
Augmenting Intelligence: The Convergence of ML/LLMs and Statistics
von: Joaquin Carbonara, et al.
Veröffentlicht: (2025)
von: Joaquin Carbonara, et al.
Veröffentlicht: (2025)
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
von: Basu, Kinjal, et al.
Veröffentlicht: (2024) -
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
von: Crouse, Maxwell, et al.
Veröffentlicht: (2026) -
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
von: Basu, Kinjal, et al.
Veröffentlicht: (2024) -
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025) -
R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory
von: Huang, Tenghao, et al.
Veröffentlicht: (2025)