Procedural Environment Generation for Tool-Use Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sullivan, Michael, Hartmann, Mareike, Koller, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2026)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
LLM Vocabulary Compression for Low-Compute Environments
von: Vennam, Sreeram, et al.
Veröffentlicht: (2024)
von: Vennam, Sreeram, et al.
Veröffentlicht: (2024)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
The Hidden Attention of Mamba Models
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
EvilGenie: A Reward Hacking Benchmark
von: Gabor, Jonathan, et al.
Veröffentlicht: (2025)
von: Gabor, Jonathan, et al.
Veröffentlicht: (2025)
Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
von: Soroka, Emi, et al.
Veröffentlicht: (2025)
von: Soroka, Emi, et al.
Veröffentlicht: (2025)
Random Scaling of Emergent Capabilities
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
von: Dosajh, Arjun, et al.
Veröffentlicht: (2025)
von: Dosajh, Arjun, et al.
Veröffentlicht: (2025)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
On the Challenges of Creating Datasets for Analyzing Commercial Sex Advertisements to Assess Human Trafficking Risk and Organized Activity
von: Rivas, Pablo, et al.
Veröffentlicht: (2024)
von: Rivas, Pablo, et al.
Veröffentlicht: (2024)
Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
von: Forchheimer, Robert
Veröffentlicht: (2026)
von: Forchheimer, Robert
Veröffentlicht: (2026)
Spectral Clustering in Convex and Constrained Settings
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2024)
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2024)
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
von: Ahmadian, Arash, et al.
Veröffentlicht: (2024)
von: Ahmadian, Arash, et al.
Veröffentlicht: (2024)
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
von: Zhao, Rosie, et al.
Veröffentlicht: (2026)
von: Zhao, Rosie, et al.
Veröffentlicht: (2026)
Autonomous Deep Agent
von: Yu, Amy, et al.
Veröffentlicht: (2025)
von: Yu, Amy, et al.
Veröffentlicht: (2025)
Social Cooperation in Conversational AI Agents
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared Loss
von: Mulgund, Abhijeet, et al.
Veröffentlicht: (2025)
von: Mulgund, Abhijeet, et al.
Veröffentlicht: (2025)
CortexCompile: Harnessing Cortical-Inspired Architectures for Enhanced Multi-Agent NLP Code Synthesis
von: Ramachandran, Gautham, et al.
Veröffentlicht: (2024)
von: Ramachandran, Gautham, et al.
Veröffentlicht: (2024)
CRAwDAD: Causal Reasoning Augmentation with Dual-Agent Debate
von: Vamosi, Finn G., et al.
Veröffentlicht: (2025)
von: Vamosi, Finn G., et al.
Veröffentlicht: (2025)
Social Learning through Interactions with Other Agents: A Survey
von: Hillier, Dylan, et al.
Veröffentlicht: (2024)
von: Hillier, Dylan, et al.
Veröffentlicht: (2024)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
von: Saxena, Udit
Veröffentlicht: (2025)
von: Saxena, Udit
Veröffentlicht: (2025)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
von: Ye, Hua, et al.
Veröffentlicht: (2025)
von: Ye, Hua, et al.
Veröffentlicht: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
QuAnTS: Question Answering on Time Series
von: Divo, Felix, et al.
Veröffentlicht: (2025)
von: Divo, Felix, et al.
Veröffentlicht: (2025)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
von: Gao, Heyang, et al.
Veröffentlicht: (2025)
von: Gao, Heyang, et al.
Veröffentlicht: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
von: Easley, Eric, et al.
Veröffentlicht: (2026)
von: Easley, Eric, et al.
Veröffentlicht: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
von: McCann, Jordan F.
Veröffentlicht: (2026)
von: McCann, Jordan F.
Veröffentlicht: (2026)
Super Apriel: One Checkpoint, Many Speeds
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
Post Hoc Extraction of Pareto Fronts for Continuous Control
von: Thakar, Raghav, et al.
Veröffentlicht: (2026)
von: Thakar, Raghav, et al.
Veröffentlicht: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
von: Steele, Brady
Veröffentlicht: (2026)
von: Steele, Brady
Veröffentlicht: (2026)
Ähnliche Einträge
-
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026) -
Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2026) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023) -
LLM Vocabulary Compression for Low-Compute Environments
von: Vennam, Sreeram, et al.
Veröffentlicht: (2024) -
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)