Reasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions
Fuente:
arXiv
Saved in:
| Main Authors: | Pezeshkpour, Pouya, Kandogan, Eser, Bhutani, Nikita, Rahman, Sajjadur, Mitchell, Tom, Hruschka, Estevam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Task Solving to Robust Real-World Adaptation in LLM Agents
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
Multi-Conditional Ranking with Large Language Models
by: Pezeshkpour, Pouya, et al.
Published: (2024)
by: Pezeshkpour, Pouya, et al.
Published: (2024)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
by: Iso, Hayate, et al.
Published: (2025)
by: Iso, Hayate, et al.
Published: (2025)
Orchestrating Agents and Data for Enterprise: A Blueprint Architecture for Compound AI
by: Kandogan, Eser, et al.
Published: (2025)
by: Kandogan, Eser, et al.
Published: (2025)
Less is More for Long Document Summary Evaluation by LLMs
by: Wu, Yunshu, et al.
Published: (2023)
by: Wu, Yunshu, et al.
Published: (2023)
FactLens: Benchmarking Fine-Grained Fact Verification
by: Mitra, Kushan, et al.
Published: (2024)
by: Mitra, Kushan, et al.
Published: (2024)
Effectiveness of Prompt Optimization in NL2SQL Systems
by: Gurajada, Sairam, et al.
Published: (2025)
by: Gurajada, Sairam, et al.
Published: (2025)
From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
by: Belem, Catarina G., et al.
Published: (2024)
by: Belem, Catarina G., et al.
Published: (2024)
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
by: Bayat, Farima Fatahi, et al.
Published: (2025)
by: Bayat, Farima Fatahi, et al.
Published: (2025)
Natural Language Processing for Human Resources: A Survey
by: Otani, Naoki, et al.
Published: (2024)
by: Otani, Naoki, et al.
Published: (2024)
Towards Probabilistic Question Answering Over Tabular Data
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
Align then Train: Efficient Retrieval Adapter Learning
by: Maekawa, Seiji, et al.
Published: (2026)
by: Maekawa, Seiji, et al.
Published: (2026)
Verification-Aware Planning for Multi-Agent Systems
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
A Blueprint Architecture of Compound AI Systems for Enterprise
by: Kandogan, Eser, et al.
Published: (2024)
by: Kandogan, Eser, et al.
Published: (2024)
Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
by: Hassell, Jackson, et al.
Published: (2025)
by: Hassell, Jackson, et al.
Published: (2025)
A Dynamic Self-Evolving Extraction System
by: Amin-Naseri, Moin, et al.
Published: (2026)
by: Amin-Naseri, Moin, et al.
Published: (2026)
Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
by: Otani, Naoki, et al.
Published: (2026)
by: Otani, Naoki, et al.
Published: (2026)
Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks
by: Mishra, Aditi, et al.
Published: (2023)
by: Mishra, Aditi, et al.
Published: (2023)
MageSQL: Enhancing In-context Learning for Text-to-SQL Applications with Large Language Models
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
by: Davoodi, Arash Gholami, et al.
Published: (2024)
by: Davoodi, Arash Gholami, et al.
Published: (2024)
Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models
by: Davoodi, Arash Gholami, et al.
Published: (2026)
by: Davoodi, Arash Gholami, et al.
Published: (2026)
CMDBench: A Benchmark for Coarse-to-fine Multimodal Data Discovery in Compound AI Systems
by: Feng, Yanlin, et al.
Published: (2024)
by: Feng, Yanlin, et al.
Published: (2024)
OmniTQA: A Cost-Aware System for Hybrid Query Processing over Semi-Structured Data
by: Shahbazi, Nima, et al.
Published: (2026)
by: Shahbazi, Nima, et al.
Published: (2026)
Towards Operationalizing Heterogeneous Data Discovery
by: Wang, Jin, et al.
Published: (2025)
by: Wang, Jin, et al.
Published: (2025)
Knowledge Acquisition and Integration with Expert-in-the-loop
by: Rahman, Sajjadur, et al.
Published: (2024)
by: Rahman, Sajjadur, et al.
Published: (2024)
ClinicalAgent: Clinical Trial Multi-Agent System with Large Language Model-based Reasoning
by: Yue, Ling, et al.
Published: (2024)
by: Yue, Ling, et al.
Published: (2024)
Blue Data Intelligence Layer: Streaming Data and Agents for Multi-source Multi-modal Data-Centric Applications
by: Aminnaseri, Moin, et al.
Published: (2026)
by: Aminnaseri, Moin, et al.
Published: (2026)
Large Human Language Models: A Need and the Challenges
by: Soni, Nikita, et al.
Published: (2023)
by: Soni, Nikita, et al.
Published: (2023)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
Probing the Limits of the Lie Detector Approach to LLM Deception
by: Berger, Tom-Felix
Published: (2026)
by: Berger, Tom-Felix
Published: (2026)
ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
by: Yao, Bohan, et al.
Published: (2025)
by: Yao, Bohan, et al.
Published: (2025)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
by: Chowdhury, Nafis, et al.
Published: (2025)
by: Chowdhury, Nafis, et al.
Published: (2025)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
The Solution for The PST-KDD-2024 OAG-Challenge
by: Zhong, Shupeng, et al.
Published: (2024)
by: Zhong, Shupeng, et al.
Published: (2024)
Similar Items
-
From Task Solving to Robust Real-World Adaptation in LLM Agents
by: Pezeshkpour, Pouya, et al.
Published: (2026) -
Multi-Conditional Ranking with Large Language Models
by: Pezeshkpour, Pouya, et al.
Published: (2024) -
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026) -
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
by: Pezeshkpour, Pouya, et al.
Published: (2025) -
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
by: Pezeshkpour, Pouya, et al.
Published: (2025)