Saved in:
| Main Authors: | Vishwakarma, Harsh, Agarwal, Ankush, Patil, Ojas, Devaguptapu, Chaitanya, Chandran, Mahesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.27287 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EnterpriseLab: A Full-Stack Platform for developing and deploying agents in Enterprises
by: Agarwal, Ankush, et al.
Published: (2026)
by: Agarwal, Ankush, et al.
Published: (2026)
Hybrid Graphs for Table-and-Text based Question Answering using LLMs
by: Agarwal, Ankush, et al.
Published: (2025)
by: Agarwal, Ankush, et al.
Published: (2025)
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
by: Kambhampati, Subbarao, et al.
Published: (2024)
by: Kambhampati, Subbarao, et al.
Published: (2024)
Embedded Safety-Aligned Intelligence via Differentiable Internal Alignment Embeddings
by: Rathva, Harsh, et al.
Published: (2025)
by: Rathva, Harsh, et al.
Published: (2025)
Making LLMs Work for Enterprise Data Tasks
by: Demiralp, Çağatay, et al.
Published: (2024)
by: Demiralp, Çağatay, et al.
Published: (2024)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
by: Jain, Ojas, et al.
Published: (2026)
by: Jain, Ojas, et al.
Published: (2026)
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
by: Thakkar, Parth, et al.
Published: (2025)
by: Thakkar, Parth, et al.
Published: (2025)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
by: Ni, Kangqi, et al.
Published: (2025)
by: Ni, Kangqi, et al.
Published: (2025)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
by: Mehta, Sushant, et al.
Published: (2026)
by: Mehta, Sushant, et al.
Published: (2026)
Self-Adaptive Graph Mixture of Models
by: Meena, Mohit, et al.
Published: (2025)
by: Meena, Mohit, et al.
Published: (2025)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
by: Li, Mengqi, et al.
Published: (2025)
by: Li, Mengqi, et al.
Published: (2025)
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
by: Kadu, Ankush, et al.
Published: (2025)
by: Kadu, Ankush, et al.
Published: (2025)
MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection
by: Purohit, Harsh, et al.
Published: (2025)
by: Purohit, Harsh, et al.
Published: (2025)
Can Stories Help LLMs Reason? Curating Information Space Through Narrative
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
by: Shastry, KN Ajay, et al.
Published: (2026)
by: Shastry, KN Ajay, et al.
Published: (2026)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Fuzzy Rule based Intelligent Cardiovascular Disease Prediction using Complex Event Processing
by: Kumar, Shashi Shekhar, et al.
Published: (2024)
by: Kumar, Shashi Shekhar, et al.
Published: (2024)
Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents
by: K, Deeraj S, et al.
Published: (2026)
by: K, Deeraj S, et al.
Published: (2026)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
by: Lee, Hayeong, et al.
Published: (2026)
by: Lee, Hayeong, et al.
Published: (2026)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026)
by: Yuan, Danlong, et al.
Published: (2026)
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
by: He, Yanheng, et al.
Published: (2024)
by: He, Yanheng, et al.
Published: (2024)
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
by: Chen, Wentao, et al.
Published: (2025)
by: Chen, Wentao, et al.
Published: (2025)
AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery
by: Kleczek, Darek, et al.
Published: (2026)
by: Kleczek, Darek, et al.
Published: (2026)
Optimal partition of feature using Bayesian classifier
by: Vishwakarma, Sanjay, et al.
Published: (2023)
by: Vishwakarma, Sanjay, et al.
Published: (2023)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Your Code Agent Can Grow Alongside You with Structured Memory
by: Deng, Yi-Xuan, et al.
Published: (2026)
by: Deng, Yi-Xuan, et al.
Published: (2026)
Adaptive LLM Routing under Budget Constraints
by: Panda, Pranoy, et al.
Published: (2025)
by: Panda, Pranoy, et al.
Published: (2025)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
FinBloom: Knowledge Grounding Large Language Model with Real-time Financial Data
by: Sinha, Ankur, et al.
Published: (2025)
by: Sinha, Ankur, et al.
Published: (2025)
Procedural Memory Is Not All You Need: Bridging Cognitive Gaps in LLM-Based Agents
by: Wheeler, Schaun, et al.
Published: (2025)
by: Wheeler, Schaun, et al.
Published: (2025)
What Can You Do When You Have Zero Rewards During RL?
by: Prakash, Jatin, et al.
Published: (2025)
by: Prakash, Jatin, et al.
Published: (2025)
Momentum Boosted Episodic Memory for Improving Learning in Long-Tailed RL Environments
by: Fernandes, Dolton, et al.
Published: (2025)
by: Fernandes, Dolton, et al.
Published: (2025)
HDL-GPT: High-Quality HDL is All You Need
by: Kumar, Bhuvnesh, et al.
Published: (2024)
by: Kumar, Bhuvnesh, et al.
Published: (2024)
LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents
by: Xiao, Chang, et al.
Published: (2024)
by: Xiao, Chang, et al.
Published: (2024)
Transformers are Graph Neural Networks
by: Joshi, Chaitanya K.
Published: (2025)
by: Joshi, Chaitanya K.
Published: (2025)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
Similar Items
-
EnterpriseLab: A Full-Stack Platform for developing and deploying agents in Enterprises
by: Agarwal, Ankush, et al.
Published: (2026) -
Hybrid Graphs for Table-and-Text based Question Answering using LLMs
by: Agarwal, Ankush, et al.
Published: (2025) -
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
by: Kambhampati, Subbarao, et al.
Published: (2024) -
Embedded Safety-Aligned Intelligence via Differentiable Internal Alignment Embeddings
by: Rathva, Harsh, et al.
Published: (2025) -
Making LLMs Work for Enterprise Data Tasks
by: Demiralp, Çağatay, et al.
Published: (2024)