Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Mittal, Avni, Kumar, Shanu, Dandapat, Sandipan, Choudhury, Monojit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
di: Lupu, Andrei, et al.
Pubblicazione: (2025)
di: Lupu, Andrei, et al.
Pubblicazione: (2025)
Benchmarking Agentic Workflow Generation
di: Qiao, Shuofei, et al.
Pubblicazione: (2024)
di: Qiao, Shuofei, et al.
Pubblicazione: (2024)
Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting
di: Hayat, Hashim, et al.
Pubblicazione: (2025)
di: Hayat, Hashim, et al.
Pubblicazione: (2025)
Agentic AI: The Era of Semantic Decoding
di: Peyrard, Maxime, et al.
Pubblicazione: (2024)
di: Peyrard, Maxime, et al.
Pubblicazione: (2024)
FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
di: Wu, Haotian, et al.
Pubblicazione: (2025)
di: Wu, Haotian, et al.
Pubblicazione: (2025)
The Automated but Risky Game: Modeling and Benchmarking Agent-to-Agent Negotiations and Transactions in Consumer Markets
di: Zhu, Shenzhe, et al.
Pubblicazione: (2025)
di: Zhu, Shenzhe, et al.
Pubblicazione: (2025)
Logarithmic Scores, Power-Law Discoveries: Disentangling Measurement from Coverage in Agent-Based Evaluation
di: Jung, HyunJoon, et al.
Pubblicazione: (2026)
di: Jung, HyunJoon, et al.
Pubblicazione: (2026)
Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration
di: Guo, Liangxuan, et al.
Pubblicazione: (2025)
di: Guo, Liangxuan, et al.
Pubblicazione: (2025)
Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning
di: Fouilhé, Guilhem, et al.
Pubblicazione: (2026)
di: Fouilhé, Guilhem, et al.
Pubblicazione: (2026)
CompeteAI: Understanding the Competition Dynamics in Large Language Model-based Agents
di: Zhao, Qinlin, et al.
Pubblicazione: (2023)
di: Zhao, Qinlin, et al.
Pubblicazione: (2023)
TransLaw: A Large-Scale Dataset and Multi-Agent Benchmark Simulating Professional Translation of Hong Kong Case Law
di: Xuan, Xi, et al.
Pubblicazione: (2025)
di: Xuan, Xi, et al.
Pubblicazione: (2025)
LLMs Can Simulate Standardized Patients via Agent Coevolution
di: Du, Zhuoyun, et al.
Pubblicazione: (2024)
di: Du, Zhuoyun, et al.
Pubblicazione: (2024)
Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
di: Sun, Jiazheng, et al.
Pubblicazione: (2025)
di: Sun, Jiazheng, et al.
Pubblicazione: (2025)
Prune4Web: DOM Tree Pruning Programming for Web Agent
di: Zhang, Jiayuan, et al.
Pubblicazione: (2025)
di: Zhang, Jiayuan, et al.
Pubblicazione: (2025)
MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
di: Liu, Yuxuan, et al.
Pubblicazione: (2025)
di: Liu, Yuxuan, et al.
Pubblicazione: (2025)
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
di: Korgul, Karolina, et al.
Pubblicazione: (2025)
di: Korgul, Karolina, et al.
Pubblicazione: (2025)
Human-Artificial Interaction in the Age of Agentic AI: A System-Theoretical Approach
di: Borghoff, Uwe M., et al.
Pubblicazione: (2025)
di: Borghoff, Uwe M., et al.
Pubblicazione: (2025)
Algorithmic Prompt Generation for Diverse Human-like Teaming and Communication with Large Language Models
di: Srikanth, Siddharth, et al.
Pubblicazione: (2025)
di: Srikanth, Siddharth, et al.
Pubblicazione: (2025)
AgenticAD: A Specialized Multiagent System Framework for Holistic Alzheimer Disease Management
di: Bazgir, Adib, et al.
Pubblicazione: (2025)
di: Bazgir, Adib, et al.
Pubblicazione: (2025)
CultivAgents: Cultivating Relationship-Centered Multi-Agent Systems for Personalized Gardening
di: Wang, Yiyang, et al.
Pubblicazione: (2026)
di: Wang, Yiyang, et al.
Pubblicazione: (2026)
The Application of MATEC (Multi-AI Agent Team Care) Framework in Sepsis Care
di: Cho, Andrew, et al.
Pubblicazione: (2025)
di: Cho, Andrew, et al.
Pubblicazione: (2025)
Psychologically Enhanced AI Agents
di: Besta, Maciej, et al.
Pubblicazione: (2025)
di: Besta, Maciej, et al.
Pubblicazione: (2025)
Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2025)
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2025)
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
di: Zou, Henry Peng, et al.
Pubblicazione: (2025)
di: Zou, Henry Peng, et al.
Pubblicazione: (2025)
Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface
di: Hua, Wenyue, et al.
Pubblicazione: (2024)
di: Hua, Wenyue, et al.
Pubblicazione: (2024)
Plato's Cave: A Human-Centered Research Verification System
di: Maldaner, Matheus Kunzler, et al.
Pubblicazione: (2026)
di: Maldaner, Matheus Kunzler, et al.
Pubblicazione: (2026)
MAP: Evaluation and Multi-Agent Enhancement of Large Language Models for Inpatient Pathways
di: Chen, Zhen, et al.
Pubblicazione: (2025)
di: Chen, Zhen, et al.
Pubblicazione: (2025)
MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation
di: Li, Mingjin, et al.
Pubblicazione: (2025)
di: Li, Mingjin, et al.
Pubblicazione: (2025)
Interactive Debugging and Steering of Multi-Agent AI Systems
di: Epperson, Will, et al.
Pubblicazione: (2025)
di: Epperson, Will, et al.
Pubblicazione: (2025)
KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents
di: Zhu, Yuqi, et al.
Pubblicazione: (2024)
di: Zhu, Yuqi, et al.
Pubblicazione: (2024)
Interactionalism: Re-Designing Higher Learning for the Large Language Agent Era
di: Moldoveanu, Mihnea C., et al.
Pubblicazione: (2025)
di: Moldoveanu, Mihnea C., et al.
Pubblicazione: (2025)
MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation
di: Zhu, Zichen, et al.
Pubblicazione: (2024)
di: Zhu, Zichen, et al.
Pubblicazione: (2024)
Perfecting Human-AI Interaction at Clinical Scale. Turning Production Signals into Safer, More Human Conversations
di: Mukherjee, Subhabrata, et al.
Pubblicazione: (2026)
di: Mukherjee, Subhabrata, et al.
Pubblicazione: (2026)
Establishing Shared Query Understanding in an Open Multi-Agent System
di: Kondylidis, Nikolaos, et al.
Pubblicazione: (2023)
di: Kondylidis, Nikolaos, et al.
Pubblicazione: (2023)
DialogGuard: Multi-Agent Psychosocial Safety Evaluation of Sensitive LLM Responses
di: Luo, Han, et al.
Pubblicazione: (2025)
di: Luo, Han, et al.
Pubblicazione: (2025)
Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory
di: Dai, Gordon, et al.
Pubblicazione: (2024)
di: Dai, Gordon, et al.
Pubblicazione: (2024)
What Is Your AI Agent Buying? Evaluation, Biases, Model Dependence, & Emerging Implications for Agentic E-Commerce
di: Allouah, Amine, et al.
Pubblicazione: (2025)
di: Allouah, Amine, et al.
Pubblicazione: (2025)
FINER-SQL: Boosting Small Language Models for Text-to-SQL
di: Hoang, Thanh Dat, et al.
Pubblicazione: (2026)
di: Hoang, Thanh Dat, et al.
Pubblicazione: (2026)
Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy
di: Hou, Abe Bohan, et al.
Pubblicazione: (2025)
di: Hou, Abe Bohan, et al.
Pubblicazione: (2025)
Plurals: A System for Guiding LLMs Via Simulated Social Ensembles
di: Ashkinaze, Joshua, et al.
Pubblicazione: (2024)
di: Ashkinaze, Joshua, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
di: Lupu, Andrei, et al.
Pubblicazione: (2025) -
Benchmarking Agentic Workflow Generation
di: Qiao, Shuofei, et al.
Pubblicazione: (2024) -
Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting
di: Hayat, Hashim, et al.
Pubblicazione: (2025) -
Agentic AI: The Era of Semantic Decoding
di: Peyrard, Maxime, et al.
Pubblicazione: (2024) -
FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
di: Wu, Haotian, et al.
Pubblicazione: (2025)