Improving Language Agents through BREW
Fuente:
arXiv
Saved in:
| Main Authors: | Kirtania, Shashank, Biyani, Param, Gupta, Priyanshu, Bajpai, Yasharth, Iyer, Roshni, Gulwani, Sumit, Soares, Gustavo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
by: Biyani, Param, et al.
Published: (2025)
by: Biyani, Param, et al.
Published: (2025)
Exploring Interaction Patterns for Debugging: Enhancing Conversational Capabilities of AI-assistants
by: Chopra, Bhavya, et al.
Published: (2024)
by: Chopra, Bhavya, et al.
Published: (2024)
MetaReflection: Learning Instructions for Language Agents using Past Reflections
by: Gupta, Priyanshu, et al.
Published: (2024)
by: Gupta, Priyanshu, et al.
Published: (2024)
STACKFEED: Structured Textual Actor-Critic Knowledge Base Editing with FeedBack
by: Kirtania, Shashank, et al.
Published: (2024)
by: Kirtania, Shashank, et al.
Published: (2024)
LOGIC-LM++: Multi-Step Refinement for Symbolic Formulations
by: Kirtania, Shashank, et al.
Published: (2024)
by: Kirtania, Shashank, et al.
Published: (2024)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
TableTalk: Scaffolding Spreadsheet Development with a Language Agent
by: Liang, Jenny T., et al.
Published: (2025)
by: Liang, Jenny T., et al.
Published: (2025)
Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
by: Kumar, Aayush, et al.
Published: (2025)
by: Kumar, Aayush, et al.
Published: (2025)
ConDABench: Interactive Evaluation of Language Models for Data Analysis
by: Dutta, Avik, et al.
Published: (2025)
by: Dutta, Avik, et al.
Published: (2025)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
by: Mhatre, Sanket, et al.
Published: (2025)
by: Mhatre, Sanket, et al.
Published: (2025)
Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Enhancing Creativity in Large Language Models through Associative Thinking Strategies
by: Mehrotra, Pronita, et al.
Published: (2024)
by: Mehrotra, Pronita, et al.
Published: (2024)
An Empirical Investigation of Robustness in Large Language Models under Tabular Distortions
by: Dutta, Avik, et al.
Published: (2026)
by: Dutta, Avik, et al.
Published: (2026)
Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Automating Human Tutor-Style Programming Feedback: Leveraging GPT-4 Tutor Model for Hint Generation and GPT-3.5 Student Model for Hint Validation
by: Phung, Tung, et al.
Published: (2023)
by: Phung, Tung, et al.
Published: (2023)
Do Code Models Suffer from the Dunning-Kruger Effect?
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Diffusion is a code repair operator and generator
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
by: Priyanshu, Aman, et al.
Published: (2026)
by: Priyanshu, Aman, et al.
Published: (2026)
TEN: Table Explicitization, Neurosymbolically
by: Mehrotra, Nikita, et al.
Published: (2025)
by: Mehrotra, Nikita, et al.
Published: (2025)
Tabularis Formatus: Predictive Formatting for Tables
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Auditing and Controlling AI Agent Actions in Spreadsheets
by: Sabouri, Sadra, et al.
Published: (2026)
by: Sabouri, Sadra, et al.
Published: (2026)
SMART-3D: Three-Dimensional Self-Morphing Adaptive Replanning Tree
by: Agrawal, Priyanshu, et al.
Published: (2025)
by: Agrawal, Priyanshu, et al.
Published: (2025)
Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction
by: Parab, Aishni, et al.
Published: (2025)
by: Parab, Aishni, et al.
Published: (2025)
No Transfers Required: Integrating Last Mile with Public Transit Using Opti-Mile
by: Altaf, Raashid, et al.
Published: (2023)
by: Altaf, Raashid, et al.
Published: (2023)
The Orchestration of Multi-Agent Systems: Architectures, Protocols, and Enterprise Adoption
by: Adimulam, Apoorva, et al.
Published: (2026)
by: Adimulam, Apoorva, et al.
Published: (2026)
Personalized Recommendation Systems using Multimodal, Autonomous, Multi Agent Systems
by: Thakkar, Param, et al.
Published: (2024)
by: Thakkar, Param, et al.
Published: (2024)
Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks
by: Joshi, Ratnesh Kumar, et al.
Published: (2024)
by: Joshi, Ratnesh Kumar, et al.
Published: (2024)
Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models
by: Berg, Cameron, et al.
Published: (2026)
by: Berg, Cameron, et al.
Published: (2026)
Computational Politeness in Natural Language Processing: A Survey
by: Priya, Priyanshu, et al.
Published: (2024)
by: Priya, Priyanshu, et al.
Published: (2024)
Soft-Label Training Preserves Epistemic Uncertainty
by: Singh, Agamdeep, et al.
Published: (2025)
by: Singh, Agamdeep, et al.
Published: (2025)
Improving EEG Signal Classification Accuracy Using Wasserstein Generative Adversarial Networks
by: Park, Joshua, et al.
Published: (2024)
by: Park, Joshua, et al.
Published: (2024)
Automated Analysis of Learning Outcomes and Exam Questions Based on Bloom's Taxonomy
by: Kumar, Ramya, et al.
Published: (2025)
by: Kumar, Ramya, et al.
Published: (2025)
Periodic Topological Deep Learning for Polymer Design and Discovery
by: Yadav, Yasharth, et al.
Published: (2026)
by: Yadav, Yasharth, et al.
Published: (2026)
Generative AI for Education (GAIED): Advances, Opportunities, and Challenges
by: Denny, Paul, et al.
Published: (2024)
by: Denny, Paul, et al.
Published: (2024)
Beyond Greedy Exits: Improved Early Exit Decisions for Risk Control and Reliability
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
An Empirical Study of Validating Synthetic Data for Formula Generation
by: Singh, Usneek, et al.
Published: (2024)
by: Singh, Usneek, et al.
Published: (2024)
Exploration vs. Fixation: Scaffolding Divergent and Convergent Thinking for Human-AI Co-Creation with Generative Models
by: Wen, Chao, et al.
Published: (2025)
by: Wen, Chao, et al.
Published: (2025)
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
by: Ahir, Param, et al.
Published: (2023)
by: Ahir, Param, et al.
Published: (2023)
Class-Level Code Generation from Natural Language Using Iterative, Tool-Enhanced Reasoning over Repository
by: Deshpande, Ajinkya, et al.
Published: (2024)
by: Deshpande, Ajinkya, et al.
Published: (2024)
Similar Items
-
IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
by: Biyani, Param, et al.
Published: (2025) -
Exploring Interaction Patterns for Debugging: Enhancing Conversational Capabilities of AI-assistants
by: Chopra, Bhavya, et al.
Published: (2024) -
MetaReflection: Learning Instructions for Language Agents using Past Reflections
by: Gupta, Priyanshu, et al.
Published: (2024) -
STACKFEED: Structured Textual Actor-Critic Knowledge Base Editing with FeedBack
by: Kirtania, Shashank, et al.
Published: (2024) -
LOGIC-LM++: Multi-Step Refinement for Symbolic Formulations
by: Kirtania, Shashank, et al.
Published: (2024)