Agent Context Protocols Enhance Collective Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhardwaj, Devansh, Beniwal, Arjun, Chaudhari, Shreyas, Kalyan, Ashwin, Rajpurohit, Tanmay, Narasimhan, Karthik R., Deshpande, Ameet, Murahari, Vishvak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probing AI Safety with Source Code
von: Narayan, Ujwal, et al.
Veröffentlicht: (2025)
von: Narayan, Ujwal, et al.
Veröffentlicht: (2025)
PersonaGym: Evaluating Persona Agents and LLMs
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
GEO: Generative Engine Optimization
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023)
Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
von: Dogra, Atharvan, et al.
Veröffentlicht: (2024)
von: Dogra, Atharvan, et al.
Veröffentlicht: (2024)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
von: Dogra, Atharvan, et al.
Veröffentlicht: (2025)
von: Dogra, Atharvan, et al.
Veröffentlicht: (2025)
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
von: Gupta, Shashank, et al.
Veröffentlicht: (2023)
von: Gupta, Shashank, et al.
Veröffentlicht: (2023)
Cognitive Architectures for Language Agents
von: Sumers, Theodore R., et al.
Veröffentlicht: (2023)
von: Sumers, Theodore R., et al.
Veröffentlicht: (2023)
$τ^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
von: Barres, Victor, et al.
Veröffentlicht: (2025)
von: Barres, Victor, et al.
Veröffentlicht: (2025)
Can Language Models Solve Olympiad Programming?
von: Shi, Quan, et al.
Veröffentlicht: (2024)
von: Shi, Quan, et al.
Veröffentlicht: (2024)
$τ$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
von: Shi, Quan, et al.
Veröffentlicht: (2026)
von: Shi, Quan, et al.
Veröffentlicht: (2026)
Contextual Experience Replay for Self-Improvement of Language Agents
von: Liu, Yitao, et al.
Veröffentlicht: (2025)
von: Liu, Yitao, et al.
Veröffentlicht: (2025)
LLMs are Superior Feedback Providers: Bootstrapping Reasoning for Lie Detection with Self-Generated Feedback
von: Banerjee, Tanushree, et al.
Veröffentlicht: (2024)
von: Banerjee, Tanushree, et al.
Veröffentlicht: (2024)
Interpretable Emergent Language Using Inter-Agent Transformers
von: Bhardwaj, Mannan
Veröffentlicht: (2025)
von: Bhardwaj, Mannan
Veröffentlicht: (2025)
Regression-aware Inference with LLMs
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems
von: Gupta, Aman, et al.
Veröffentlicht: (2024)
von: Gupta, Aman, et al.
Veröffentlicht: (2024)
Language-Guided World Models: A Model-Based Approach to AI Control
von: Zhang, Alex, et al.
Veröffentlicht: (2024)
von: Zhang, Alex, et al.
Veröffentlicht: (2024)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
von: Shen, Junhong, et al.
Veröffentlicht: (2024)
von: Shen, Junhong, et al.
Veröffentlicht: (2024)
VideoGameBench: Can Vision-Language Models complete popular video games?
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts
von: Tirumala, Shreyas, et al.
Veröffentlicht: (2025)
von: Tirumala, Shreyas, et al.
Veröffentlicht: (2025)
From Context to Action: Analysis of the Impact of State Representation and Context on the Generalization of Multi-Turn Web Navigation Agents
von: Tiwary, Nalin, et al.
Veröffentlicht: (2024)
von: Tiwary, Nalin, et al.
Veröffentlicht: (2024)
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
von: Kapoor, Vansh, et al.
Veröffentlicht: (2026)
von: Kapoor, Vansh, et al.
Veröffentlicht: (2026)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
von: Agrawal, Tanmay
Veröffentlicht: (2025)
von: Agrawal, Tanmay
Veröffentlicht: (2025)
Enhancing Online Learning Efficiency Through Heterogeneous Resource Integration with a Multi-Agent RAG System
von: Srivastav, Devansh, et al.
Veröffentlicht: (2025)
von: Srivastav, Devansh, et al.
Veröffentlicht: (2025)
Cross-lingual Editing in Multilingual Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
Multilingual Information Retrieval with a Monolingual Knowledge Base
von: Zhuang, Yingying, et al.
Veröffentlicht: (2025)
von: Zhuang, Yingying, et al.
Veröffentlicht: (2025)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation
von: Sudhakar, Arjun Vaithilingam
Veröffentlicht: (2025)
von: Sudhakar, Arjun Vaithilingam
Veröffentlicht: (2025)
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
von: Wan, Guangya, et al.
Veröffentlicht: (2025)
von: Wan, Guangya, et al.
Veröffentlicht: (2025)
Refract ICL: Rethinking Example Selection in the Era of Million-Token Models
von: Akula, Arjun R., et al.
Veröffentlicht: (2025)
von: Akula, Arjun R., et al.
Veröffentlicht: (2025)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Probing AI Safety with Source Code
von: Narayan, Ujwal, et al.
Veröffentlicht: (2025) -
PersonaGym: Evaluating Persona Agents and LLMs
von: Samuel, Vinay, et al.
Veröffentlicht: (2024) -
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023) -
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024) -
GEO: Generative Engine Optimization
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023)