AI Agents That Matter
Fuente:
arXiv
Salvato in:
| Autori principali: | Kapoor, Sayash, Stroebl, Benedikt, Siegel, Zachary S., Nadgir, Nitya, Narayanan, Arvind |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Limits of Inference Scaling Through Resampling
di: Stroebl, Benedikt, et al.
Pubblicazione: (2024)
di: Stroebl, Benedikt, et al.
Pubblicazione: (2024)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
di: Siegel, Zachary S., et al.
Pubblicazione: (2024)
di: Siegel, Zachary S., et al.
Pubblicazione: (2024)
Towards a Science of AI Agent Reliability
di: Rabanser, Stephan, et al.
Pubblicazione: (2026)
di: Rabanser, Stephan, et al.
Pubblicazione: (2026)
Log analysis is necessary for credible evaluation of AI agents
di: Kirgis, Peter, et al.
Pubblicazione: (2026)
di: Kirgis, Peter, et al.
Pubblicazione: (2026)
Promises and pitfalls of artificial intelligence for legal applications
di: Kapoor, Sayash, et al.
Pubblicazione: (2024)
di: Kapoor, Sayash, et al.
Pubblicazione: (2024)
Foundation Model Transparency Reports
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
di: Kapoor, Sayash, et al.
Pubblicazione: (2025)
di: Kapoor, Sayash, et al.
Pubblicazione: (2025)
The 2024 Foundation Model Transparency Index
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
The 2025 Foundation Model Transparency Index
di: Wan, Alexander, et al.
Pubblicazione: (2025)
di: Wan, Alexander, et al.
Pubblicazione: (2025)
Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
di: Tiwari, Nitya, et al.
Pubblicazione: (2025)
di: Tiwari, Nitya, et al.
Pubblicazione: (2025)
Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning
di: Zhang, Beining, et al.
Pubblicazione: (2025)
di: Zhang, Beining, et al.
Pubblicazione: (2025)
QDeepGR4J: Quantile-based ensemble of deep learning and GR4J hybrid rainfall-runoff models for extreme flow prediction with uncertainty quantification
di: Kapoor, Arpit, et al.
Pubblicazione: (2025)
di: Kapoor, Arpit, et al.
Pubblicazione: (2025)
The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources
di: Longpre, Shayne, et al.
Pubblicazione: (2024)
di: Longpre, Shayne, et al.
Pubblicazione: (2024)
Seven simple steps for log analysis in AI systems
di: Dubois, Magda, et al.
Pubblicazione: (2026)
di: Dubois, Magda, et al.
Pubblicazione: (2026)
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
di: Veselovsky, Veniamin, et al.
Pubblicazione: (2025)
di: Veselovsky, Veniamin, et al.
Pubblicazione: (2025)
PCS Workflow for Veridical Data Science in the Age of AI
di: Rewolinski, Zachary T., et al.
Pubblicazione: (2025)
di: Rewolinski, Zachary T., et al.
Pubblicazione: (2025)
Partner Modelling Emerges in Recurrent Agents (But Only When It Matters)
di: Mon-Williams, Ruaridh, et al.
Pubblicazione: (2025)
di: Mon-Williams, Ruaridh, et al.
Pubblicazione: (2025)
Learning to Help in Multi-Class Settings
di: Wu, Yu, et al.
Pubblicazione: (2025)
di: Wu, Yu, et al.
Pubblicazione: (2025)
On the Societal Impact of Open Foundation Models
di: Kapoor, Sayash, et al.
Pubblicazione: (2024)
di: Kapoor, Sayash, et al.
Pubblicazione: (2024)
Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization
di: Kapoor, Aditya, et al.
Pubblicazione: (2024)
di: Kapoor, Aditya, et al.
Pubblicazione: (2024)
DeepACTIF: Efficient Feature Attribution via Activation Traces in Neural Sequence Models
di: Hosp, Benedikt W.
Pubblicazione: (2025)
di: Hosp, Benedikt W.
Pubblicazione: (2025)
AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
di: Raj, Aman, et al.
Pubblicazione: (2025)
di: Raj, Aman, et al.
Pubblicazione: (2025)
Can AI autonomously build, operate, and use the entire data stack?
di: Agarwal, Arvind, et al.
Pubblicazione: (2025)
di: Agarwal, Arvind, et al.
Pubblicazione: (2025)
HiCL: Hippocampal-Inspired Continual Learning
di: Kapoor, Kushal, et al.
Pubblicazione: (2025)
di: Kapoor, Kushal, et al.
Pubblicazione: (2025)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
di: Putta, Pranav, et al.
Pubblicazione: (2024)
di: Putta, Pranav, et al.
Pubblicazione: (2024)
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
di: Luo, Xufang, et al.
Pubblicazione: (2025)
di: Luo, Xufang, et al.
Pubblicazione: (2025)
A survey on Concept-based Approaches For Model Improvement
di: Gupta, Avani, et al.
Pubblicazione: (2024)
di: Gupta, Avani, et al.
Pubblicazione: (2024)
Large Scale Constrained Clustering With Reinforcement Learning
di: Schesch, Benedikt, et al.
Pubblicazione: (2024)
di: Schesch, Benedikt, et al.
Pubblicazione: (2024)
Healthcare AI GYM for Medical Agents
di: Jeong, Minbyul
Pubblicazione: (2026)
di: Jeong, Minbyul
Pubblicazione: (2026)
AI Agents as Universal Task Solvers
di: Achille, Alessandro, et al.
Pubblicazione: (2025)
di: Achille, Alessandro, et al.
Pubblicazione: (2025)
Beyond Accuracy: EcoL2 Metric for Sustainable Neural PDE Solvers
di: Kapoor, Taniya, et al.
Pubblicazione: (2025)
di: Kapoor, Taniya, et al.
Pubblicazione: (2025)
Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
di: Kapoor, Aditya, et al.
Pubblicazione: (2025)
di: Kapoor, Aditya, et al.
Pubblicazione: (2025)
TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools
di: Gao, Shanghua, et al.
Pubblicazione: (2025)
di: Gao, Shanghua, et al.
Pubblicazione: (2025)
Development of a graph neural network surrogate for travel demand modelling
di: Makarov, Nikita, et al.
Pubblicazione: (2024)
di: Makarov, Nikita, et al.
Pubblicazione: (2024)
On Limitations of the Transformer Architecture
di: Peng, Binghui, et al.
Pubblicazione: (2024)
di: Peng, Binghui, et al.
Pubblicazione: (2024)
Learning Regularizers: Learning Optimizers that can Regularize
di: Sahoo, Suraj Kumar, et al.
Pubblicazione: (2025)
di: Sahoo, Suraj Kumar, et al.
Pubblicazione: (2025)
Inverse Reinforcement Learning from Non-Stationary Learning Agents
di: Sivakumar, Kavinayan P., et al.
Pubblicazione: (2024)
di: Sivakumar, Kavinayan P., et al.
Pubblicazione: (2024)
Aligning AI Agents via Information-Directed Sampling
di: Jeon, Hong Jun, et al.
Pubblicazione: (2024)
di: Jeon, Hong Jun, et al.
Pubblicazione: (2024)
Self-Improving AI Agents through Self-Play
di: Chojecki, Przemyslaw
Pubblicazione: (2025)
di: Chojecki, Przemyslaw
Pubblicazione: (2025)
REX: Rapid Exploration and eXploitation for AI Agents
di: Murthy, Rithesh, et al.
Pubblicazione: (2023)
di: Murthy, Rithesh, et al.
Pubblicazione: (2023)
Documenti analoghi
-
The Limits of Inference Scaling Through Resampling
di: Stroebl, Benedikt, et al.
Pubblicazione: (2024) -
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
di: Siegel, Zachary S., et al.
Pubblicazione: (2024) -
Towards a Science of AI Agent Reliability
di: Rabanser, Stephan, et al.
Pubblicazione: (2026) -
Log analysis is necessary for credible evaluation of AI agents
di: Kirgis, Peter, et al.
Pubblicazione: (2026) -
Promises and pitfalls of artificial intelligence for legal applications
di: Kapoor, Sayash, et al.
Pubblicazione: (2024)