AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Mazaheri, Parsa, Mazaheri, Kasra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026)
por: Wang, Yuchen, et al.
Publicado: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
por: Jehu-Appiah, Rodney
Publicado: (2026)
por: Jehu-Appiah, Rodney
Publicado: (2026)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
por: Costa, Rimom
Publicado: (2025)
por: Costa, Rimom
Publicado: (2025)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
por: Mazaheri, Parsa
Publicado: (2026)
por: Mazaheri, Parsa
Publicado: (2026)
Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
por: Figueiredo, Vanessa
Publicado: (2025)
por: Figueiredo, Vanessa
Publicado: (2025)
Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
por: Shafique, Muhammad Ali, et al.
Publicado: (2025)
por: Shafique, Muhammad Ali, et al.
Publicado: (2025)
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
por: Lan, Guangchen
Publicado: (2026)
por: Lan, Guangchen
Publicado: (2026)
Latent Cache Flow: Model-to-Model Communication Without Text
por: Rossi, Maximillian, et al.
Publicado: (2026)
por: Rossi, Maximillian, et al.
Publicado: (2026)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
por: Huang, Yunpeng, et al.
Publicado: (2023)
por: Huang, Yunpeng, et al.
Publicado: (2023)
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
por: Saini, Mayank, et al.
Publicado: (2025)
por: Saini, Mayank, et al.
Publicado: (2025)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
por: Kohl, Jens, et al.
Publicado: (2024)
por: Kohl, Jens, et al.
Publicado: (2024)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
por: Naik, Akshat, et al.
Publicado: (2025)
por: Naik, Akshat, et al.
Publicado: (2025)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
por: Ashrafi, Nazmus
Publicado: (2026)
por: Ashrafi, Nazmus
Publicado: (2026)
Collaborative LLM Agents for C4 Software Architecture Design Automation
por: Szczepanik, Kamil, et al.
Publicado: (2025)
por: Szczepanik, Kamil, et al.
Publicado: (2025)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
por: Ray, Aninda
Publicado: (2026)
por: Ray, Aninda
Publicado: (2026)
ContractBench: Can LLM Agents Preserve Observation Contracts?
por: Wang, Jicheng, et al.
Publicado: (2026)
por: Wang, Jicheng, et al.
Publicado: (2026)
Beyond Prompt Engineering: Neuro-Symbolic-Causal Architecture for Robust Multi-Objective AI Agents
por: Akarlar, Gokturk Aytug
Publicado: (2025)
por: Akarlar, Gokturk Aytug
Publicado: (2025)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
por: Ashley, Dylan R., et al.
Publicado: (2026)
por: Ashley, Dylan R., et al.
Publicado: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026)
por: Tang, Wenjie, et al.
Publicado: (2026)
Prompt Readiness Levels (PRL): a maturity scale and scoring framework for production grade prompt assets
por: Guinard, Sebastien
Publicado: (2026)
por: Guinard, Sebastien
Publicado: (2026)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
por: Trooskens, Geert, et al.
Publicado: (2026)
por: Trooskens, Geert, et al.
Publicado: (2026)
Applying Cognitive Design Patterns to General LLM Agents
por: Wray, Robert E., et al.
Publicado: (2025)
por: Wray, Robert E., et al.
Publicado: (2025)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
por: Wang, Zhen, et al.
Publicado: (2025)
por: Wang, Zhen, et al.
Publicado: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
por: Jia, Xiao
Publicado: (2026)
por: Jia, Xiao
Publicado: (2026)
Extending NGU to Multi-Agent RL: A Preliminary Study
por: Hernandez, Juan, et al.
Publicado: (2025)
por: Hernandez, Juan, et al.
Publicado: (2025)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
por: Tanjim, Md Mehrab, et al.
Publicado: (2026)
por: Tanjim, Md Mehrab, et al.
Publicado: (2026)
MFH: A Multi-faceted Heuristic Algorithm Selection Approach for Software Verification
por: Su, Jie, et al.
Publicado: (2025)
por: Su, Jie, et al.
Publicado: (2025)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
por: Lawand, Daniel Angelo Esteves, et al.
Publicado: (2025)
por: Lawand, Daniel Angelo Esteves, et al.
Publicado: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
por: Ferreira, Renato Cordeiro, et al.
Publicado: (2025)
por: Ferreira, Renato Cordeiro, et al.
Publicado: (2025)
FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing
por: Zhang, Mingda, et al.
Publicado: (2026)
por: Zhang, Mingda, et al.
Publicado: (2026)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
por: Alpay, Faruk, et al.
Publicado: (2025)
por: Alpay, Faruk, et al.
Publicado: (2025)
Deployment-Time Reliability of Learned Robot Policies
por: Agia, Christopher
Publicado: (2026)
por: Agia, Christopher
Publicado: (2026)
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
por: Sidik, Bronislav, et al.
Publicado: (2026)
por: Sidik, Bronislav, et al.
Publicado: (2026)
Dynamic Attentional Context Scoping: Agent-Triggered Focus Sessions for Isolated Per-Agent Steering in Multi-Agent LLM Orchestration
por: Patel, Nickson
Publicado: (2026)
por: Patel, Nickson
Publicado: (2026)
ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting
por: Chang, Jiale, et al.
Publicado: (2026)
por: Chang, Jiale, et al.
Publicado: (2026)
Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
por: Wang, Zixu, et al.
Publicado: (2026)
por: Wang, Zixu, et al.
Publicado: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
por: Lauffer, Niklas, et al.
Publicado: (2025)
por: Lauffer, Niklas, et al.
Publicado: (2025)
PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies
por: Yehia, Ahmad, et al.
Publicado: (2026)
por: Yehia, Ahmad, et al.
Publicado: (2026)
Ejemplares similares
-
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
por: Wang, Yuchen, et al.
Publicado: (2026) -
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
por: Wu, Shuai, et al.
Publicado: (2026) -
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
por: Jehu-Appiah, Rodney
Publicado: (2026) -
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
por: Costa, Rimom
Publicado: (2025) -
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
por: Mazaheri, Parsa
Publicado: (2026)