TapeAgents: a Holistic Framework for Agent Development and Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bahdanau, Dzmitry, Gontier, Nicolas, Huang, Gabriel, Kamalloo, Ehsan, Pardinas, Rafael, Piché, Alex, Scholak, Torsten, Shliazhko, Oleh, Tremblay, Jordan Prince, Ghanem, Karam, Parikh, Soham, Tiwari, Mitul, Vohra, Quaizar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
von: Piché, Alexandre, et al.
Veröffentlicht: (2025)
von: Piché, Alexandre, et al.
Veröffentlicht: (2025)
LLMs can learn self-restraint through iterative self-reflection
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026)
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026)
Self-Evolving Curriculum for LLM Reasoning
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
von: Murty, Shikhar, et al.
Veröffentlicht: (2024)
von: Murty, Shikhar, et al.
Veröffentlicht: (2024)
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
Perplexed: Understanding When Large Language Models are Confused
von: Cooper, Nathan, et al.
Veröffentlicht: (2024)
von: Cooper, Nathan, et al.
Veröffentlicht: (2024)
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
Evaluating In-Context Learning of Libraries for Code Generation
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
Agentic AI Infrastructure: Platform Economics of Multi-Agent Systems
von: Ivchenko, Oleh
Veröffentlicht: (2026)
von: Ivchenko, Oleh
Veröffentlicht: (2026)
BRIDGE: Predicting Human Task Completion Time From Model Performance
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
Unifying Autoregressive and Diffusion-Based Sequence Generation
von: Fathi, Nima, et al.
Veröffentlicht: (2025)
von: Fathi, Nima, et al.
Veröffentlicht: (2025)
The Uncanny Valley: A Comprehensive Analysis of Diffusion Models
von: Ghanem, Karam, et al.
Veröffentlicht: (2024)
von: Ghanem, Karam, et al.
Veröffentlicht: (2024)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
CUBE: A Standard for Unifying Agent Benchmarks
von: Lacoste, Alexandre, et al.
Veröffentlicht: (2026)
von: Lacoste, Alexandre, et al.
Veröffentlicht: (2026)
Pure-state $N$-representability in current-spin-density-functional theory
von: Gontier, David
Veröffentlicht: (2015)
von: Gontier, David
Veröffentlicht: (2015)
$N$-representability in non-collinear spin-polarized density functional theory
von: Gontier, David
Veröffentlicht: (2013)
von: Gontier, David
Veröffentlicht: (2013)
Edge states for second order elliptic operators in a channel
von: Gontier, David
Veröffentlicht: (2021)
von: Gontier, David
Veröffentlicht: (2021)
Existence of minimizers for Kohn-Sham within the Local Spin Density Approximation
von: Gontier, David
Veröffentlicht: (2014)
von: Gontier, David
Veröffentlicht: (2014)
Rank One Hilbert Geometries
von: Islam, Mitul
Veröffentlicht: (2019)
von: Islam, Mitul
Veröffentlicht: (2019)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
Agents' Behavior and Interest Rate Model Optimization in DeFi Lending
von: Charles Bertucci, et al.
Veröffentlicht: (2025)
von: Charles Bertucci, et al.
Veröffentlicht: (2025)
emilypiche/fsmmn_qc: fsmmn_qc v1.0
von: Emily Piche
Veröffentlicht: (2026)
von: Emily Piche
Veröffentlicht: (2026)
FICHTE, SCHLEIERMACHER Y W. VON HUMBOLDT, SOBRE LA CREACIÓN DE LA UNIVERSIDAD DE BERLÍN
von: Claude Piché
Veröffentlicht: (2005)
von: Claude Piché
Veröffentlicht: (2005)
Conversational Text Extraction with Large Language Models Using Retrieval-Augmented Systems
von: Roy, Soham, et al.
Veröffentlicht: (2025)
von: Roy, Soham, et al.
Veröffentlicht: (2025)
Apriel-H1: Towards Efficient Enterprise Reasoning Models
von: Ostapenko, Oleksiy, et al.
Veröffentlicht: (2025)
von: Ostapenko, Oleksiy, et al.
Veröffentlicht: (2025)
Holistic Evaluation and Failure Diagnosis of AI Agents
von: Madvil, Netta, et al.
Veröffentlicht: (2026)
von: Madvil, Netta, et al.
Veröffentlicht: (2026)
Memoria de Trabajo y Envejecimiento
von: Jorge Gontier B.
Veröffentlicht: (2004)
von: Jorge Gontier B.
Veröffentlicht: (2004)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
Learning with Digital Agents: An Analysis based on the Activity Theory
von: Dolata, Mateusz, et al.
Veröffentlicht: (2024)
von: Dolata, Mateusz, et al.
Veröffentlicht: (2024)
AURA Score: A Metric For Holistic Audio Question Answering Evaluation
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
Improvisational Games as a Benchmark for Social Intelligence of AI Agents: The Case of Connections
von: Parikh, Gaurav Rajesh, et al.
Veröffentlicht: (2026)
von: Parikh, Gaurav Rajesh, et al.
Veröffentlicht: (2026)
How to Train Your LLM Web Agent: A Statistical Diagnosis
von: Vattikonda, Dheeraj, et al.
Veröffentlicht: (2025)
von: Vattikonda, Dheeraj, et al.
Veröffentlicht: (2025)
A General Fixed-Point Theorem for Correspondences
von: Vohra, Ranjit
Veröffentlicht: (2025)
von: Vohra, Ranjit
Veröffentlicht: (2025)
A Further Generalization of the Gale-Nikaido-Kuhn-Debreu Market Equilibrium Theorem
von: Vohra, Ranjit
Veröffentlicht: (2025)
von: Vohra, Ranjit
Veröffentlicht: (2025)
Singularity Blockchain Key Management via non-custodial key management
von: Vohra, Sumit
Veröffentlicht: (2025)
von: Vohra, Sumit
Veröffentlicht: (2025)
A Generalization of the "Brouwer-Schauder-Tychonoff" Fixed-Point Theorem
von: Vohra, Ranjit
Veröffentlicht: (2025)
von: Vohra, Ranjit
Veröffentlicht: (2025)
Ähnliche Einträge
-
PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
von: Piché, Alexandre, et al.
Veröffentlicht: (2025) -
LLMs can learn self-restraint through iterative self-reflection
von: Piché, Alexandre, et al.
Veröffentlicht: (2024) -
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026) -
Self-Evolving Curriculum for LLM Reasoning
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025) -
NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
von: Murty, Shikhar, et al.
Veröffentlicht: (2024)