COMPASS: Benchmarking Constrained Optimization in LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Tian, Bai, Felix, Hu, Ting-Yao, Vemulapalli, Raviteja, Koppula, Hema Swetha, Xu, Zhiyang, Jin, Bowen, Cemri, Mert, Lu, Jiarui, Wang, Zirui, Cao, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Self Critique and Refinement for Faithful LLM Summarization
by: Hu, Ting-Yao, et al.
Published: (2025)
by: Hu, Ting-Yao, et al.
Published: (2025)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Learning to Reason for Hallucination Span Detection
by: Su, Hsuan, et al.
Published: (2025)
by: Su, Hsuan, et al.
Published: (2025)
Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
MUSCLE: A Model Update Strategy for Compatible LLM Evolution
by: Echterhoff, Jessica, et al.
Published: (2024)
by: Echterhoff, Jessica, et al.
Published: (2024)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
by: Bai, Hao, et al.
Published: (2024)
by: Bai, Hao, et al.
Published: (2024)
ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context
by: Xiu, Zidi, et al.
Published: (2026)
by: Xiu, Zidi, et al.
Published: (2026)
Synth4Seg -- Learning Defect Data Synthesis for Defect Segmentation using Bi-level Optimization
by: Mou, Shancong, et al.
Published: (2024)
by: Mou, Shancong, et al.
Published: (2024)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
by: Li, Jeffrey, et al.
Published: (2025)
by: Li, Jeffrey, et al.
Published: (2025)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis
by: Qin, Bowen
Published: (2026)
by: Qin, Bowen
Published: (2026)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
by: Xu, Zhiyang, et al.
Published: (2026)
by: Xu, Zhiyang, et al.
Published: (2026)
AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
by: Cemri, Mert, et al.
Published: (2026)
by: Cemri, Mert, et al.
Published: (2026)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
Scarecrow‐Shaped Antenna Optimization Using Machine Learning Algorithms
by: S. Bhavani, et al.
Published: (2025)
by: S. Bhavani, et al.
Published: (2025)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
by: Vemulapalli, Raviteja, et al.
Published: (2023)
by: Vemulapalli, Raviteja, et al.
Published: (2023)
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Bounds on Successive Minima of Orders in Number Fields and Scrollar Invariants of Curves
by: Vemulapalli, Sameera
Published: (2022)
by: Vemulapalli, Sameera
Published: (2022)
The distribution of lattices arising from orders in low degree number fields
by: Vemulapalli, Sameera
Published: (2024)
by: Vemulapalli, Sameera
Published: (2024)
The Steinitz Realization Problem
by: Vemulapalli, Sameera
Published: (2024)
by: Vemulapalli, Sameera
Published: (2024)
The Role of GDPs in Oral Surgery Delivery
by: Vidwat Vemulapalli
Published: (2025)
by: Vidwat Vemulapalli
Published: (2025)
Evaluating Pulsed Field Ablation for Atrial Fibrillation: A Systematic Review and Meta‐Analysis of Procedural Outcomes
by: Poojan Prajapati, et al.
Published: (2025)
by: Poojan Prajapati, et al.
Published: (2025)
AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
Learning Safely Without Knowing the World:COMPASS-Hedge
by: Hu, Ting, et al.
Published: (2026)
by: Hu, Ting, et al.
Published: (2026)
Intelligent cloud networking: Applying ai and reinforcement learning for dynamic traffic engineering, QoS optimization and threat detection in software-defined cloud architectures
by: Guntupalli, Raviteja
Published: (2025)
by: Guntupalli, Raviteja
Published: (2025)
Evolution of Panchayati Raj System in India with Special Reference to Telangana: Issues and Challenges
by: Mallesham, Koppula, et al.
Published: (2025)
by: Mallesham, Koppula, et al.
Published: (2025)
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
by: Wan, Guangya, et al.
Published: (2025)
by: Wan, Guangya, et al.
Published: (2025)
TiC-CLIP: Continual Training of CLIP Models
by: Garg, Saurabh, et al.
Published: (2023)
by: Garg, Saurabh, et al.
Published: (2023)
CAMPHOR: Collaborative Agents for Multi-input Planning and High-Order Reasoning On Device
by: Fu, Yicheng, et al.
Published: (2024)
by: Fu, Yicheng, et al.
Published: (2024)
CatMemo at the FinLLM Challenge Task: Fine-Tuning Large Language Models using Data Fusion in Financial Applications
by: Cao, Yupeng, et al.
Published: (2024)
by: Cao, Yupeng, et al.
Published: (2024)
Synthesis of Imidazo[1,2‐ a ]pyridine‐Isoquinoline Derivatives as Potent EGFR Inhibiting Anticancer Agents
by: Mahendar Reddy Gunuguntla, et al.
Published: (2024)
by: Mahendar Reddy Gunuguntla, et al.
Published: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
by: Wei, Bowen
Published: (2025)
by: Wei, Bowen
Published: (2025)
COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents
by: Shen, Wenkai, et al.
Published: (2026)
by: Shen, Wenkai, et al.
Published: (2026)
An Accelerated Primal Dual Algorithm with Backtracking for Decentralized Constrained Optimization
by: Xu, Qiushui, et al.
Published: (2025)
by: Xu, Qiushui, et al.
Published: (2025)
Descent-Net: Learning Descent Directions for Constrained Optimization
by: Zhou, Zisheng, et al.
Published: (2025)
by: Zhou, Zisheng, et al.
Published: (2025)
LangDA: Building Context-Awareness via Language for Domain Adaptive Semantic Segmentation
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Effect of container size and types on the root phenotypic characters of Capsicum
by: M.S.V. Raviteja
Published: (2021)
by: M.S.V. Raviteja
Published: (2021)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
by: Si, Wenwen, et al.
Published: (2025)
by: Si, Wenwen, et al.
Published: (2025)
AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
by: Jiang, Tanqiu, et al.
Published: (2026)
by: Jiang, Tanqiu, et al.
Published: (2026)
Similar Items
-
Learning from Self Critique and Refinement for Faithful LLM Summarization
by: Hu, Ting-Yao, et al.
Published: (2025) -
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025) -
Learning to Reason for Hallucination Span Detection
by: Su, Hsuan, et al.
Published: (2025) -
Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
by: Jin, Bowen, et al.
Published: (2025) -
MUSCLE: A Model Update Strategy for Compatible LLM Evolution
by: Echterhoff, Jessica, et al.
Published: (2024)