The AI Agent Index
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Casper, Stephen, Bailey, Luke, Hunter, Rosco, Ezell, Carson, Cabalé, Emma, Gerovitch, Michael, Slocum, Stewart, Wei, Kevin, Jurkovic, Nikola, Khan, Ariba, Christoffersen, Phillip J. K., Ozisik, A. Pinar, Trivedi, Rakshit, Hadfield-Menell, Dylan, Kolt, Noam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
von: Khan, Ariba, et al.
Veröffentlicht: (2025)
von: Khan, Ariba, et al.
Veröffentlicht: (2025)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
von: Staufer, Leon, et al.
Veröffentlicht: (2026)
von: Staufer, Leon, et al.
Veröffentlicht: (2026)
Diverse Preference Learning for Capabilities and Alignment
von: Slocum, Stewart, et al.
Veröffentlicht: (2025)
von: Slocum, Stewart, et al.
Veröffentlicht: (2025)
Pitfalls of Evidence-Based AI Policy
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
von: Haupt, Andreas A., et al.
Veröffentlicht: (2022)
von: Haupt, Andreas A., et al.
Veröffentlicht: (2022)
Eight Methods to Evaluate Robust Unlearning in LLMs
von: Lynch, Aengus, et al.
Veröffentlicht: (2024)
von: Lynch, Aengus, et al.
Veröffentlicht: (2024)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Black-Box Access is Insufficient for Rigorous AI Audits
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Governing AI Agents
von: Kolt, Noam
Veröffentlicht: (2025)
von: Kolt, Noam
Veröffentlicht: (2025)
Superintelligence and Law
von: Kolt, Noam
Veröffentlicht: (2026)
von: Kolt, Noam
Veröffentlicht: (2026)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
von: Soni, Prajna, et al.
Veröffentlicht: (2025)
von: Soni, Prajna, et al.
Veröffentlicht: (2025)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
Prompt Injection as Role Confusion
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Cooperative Inverse Reinforcement Learning
von: Hadfield-Menell, Dylan, et al.
Veröffentlicht: (2016)
von: Hadfield-Menell, Dylan, et al.
Veröffentlicht: (2016)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
Layered Unlearning for Adversarial Relearning
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2024)
von: Ma, Rachel, et al.
Veröffentlicht: (2024)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
An FDA for AI? Pitfalls and Plausibility of Approval Regulation for Frontier Artificial Intelligence
von: Carpenter, Daniel, et al.
Veröffentlicht: (2024)
von: Carpenter, Daniel, et al.
Veröffentlicht: (2024)
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains
von: Khan, Emaan Bilal, et al.
Veröffentlicht: (2026)
von: Khan, Emaan Bilal, et al.
Veröffentlicht: (2026)
Build Agent Advocates, Not Platform Agents
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
Regulating AI Agents
von: Gardhouse, Kathrin, et al.
Veröffentlicht: (2026)
von: Gardhouse, Kathrin, et al.
Veröffentlicht: (2026)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Lessons from complexity theory for AI governance
von: Kolt, Noam, et al.
Veröffentlicht: (2025)
von: Kolt, Noam, et al.
Veröffentlicht: (2025)
Visibility into AI Agents
von: Chan, Alan, et al.
Veröffentlicht: (2024)
von: Chan, Alan, et al.
Veröffentlicht: (2024)
Monitoring Human Dependence On AI Systems With Reliance Drills
von: Hunter, Rosco, et al.
Veröffentlicht: (2024)
von: Hunter, Rosco, et al.
Veröffentlicht: (2024)
The importance of inland water CO2, CH4, N2O to summertime greenhouse gas exchange with the atmosphere in Arctic tundra lowlands
von: Rosco, Melanie Martyn
Veröffentlicht: (2024)
von: Rosco, Melanie Martyn
Veröffentlicht: (2024)
Practical Principles for AI Cost and Compute Accounting
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
Incident Analysis for AI Agents
von: Ezell, Carson, et al.
Veröffentlicht: (2025)
von: Ezell, Carson, et al.
Veröffentlicht: (2025)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
von: Sharma, Kartik, et al.
Veröffentlicht: (2026)
von: Sharma, Kartik, et al.
Veröffentlicht: (2026)
PRIMES STEP Experience
von: Gerovitch, Slava, et al.
Veröffentlicht: (2026)
von: Gerovitch, Slava, et al.
Veröffentlicht: (2026)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
von: Che, Zora, et al.
Veröffentlicht: (2025)
von: Che, Zora, et al.
Veröffentlicht: (2025)
La Educación Ambiental y la Educación para el Desarrollo Sostenible
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2016)
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2016)
El desarrollo a propósito del pensamiento de Rodolfo Stavenhagen
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2016)
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2016)
El desarrollo sostenible en la actividad constructiva
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2017)
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2017)
Educación para el Desarrollo Sostenible: una herramienta clave para la sostenibilidad de la actividad constructiva
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2017)
von: Elizabeth Cabalé Miranda
Veröffentlicht: (2017)
Dark Speculation: Combining Qualitative and Quantitative Understanding in Frontier AI Risk Analysis
von: Carpenter, Daniel, et al.
Veröffentlicht: (2025)
von: Carpenter, Daniel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
von: Khan, Ariba, et al.
Veröffentlicht: (2025) -
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
von: Staufer, Leon, et al.
Veröffentlicht: (2026) -
Diverse Preference Learning for Capabilities and Alignment
von: Slocum, Stewart, et al.
Veröffentlicht: (2025) -
Pitfalls of Evidence-Based AI Policy
von: Casper, Stephen, et al.
Veröffentlicht: (2025) -
Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
von: Haupt, Andreas A., et al.
Veröffentlicht: (2022)