Benchmarking Agents in Insurance Underwriting Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Dsouza, Amanda, Ramakrishnan, Ramya, Dickens, Charles, Pohani, Bhavishya, Glaze, Christopher M |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
by: Dsouza, Amanda, et al.
Published: (2024)
by: Dsouza, Amanda, et al.
Published: (2024)
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique
by: Roy, Joyjit, et al.
Published: (2026)
by: Roy, Joyjit, et al.
Published: (2026)
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
by: Qi, Zhengyang, et al.
Published: (2026)
by: Qi, Zhengyang, et al.
Published: (2026)
Architecture for Simulating Behavior Mode Changes in Norm-Aware Autonomous Agents
by: Glaze, Sean, et al.
Published: (2025)
by: Glaze, Sean, et al.
Published: (2025)
Multi-Step Dialogue Workflow Action Prediction
by: Ramakrishnan, Ramya, et al.
Published: (2023)
by: Ramakrishnan, Ramya, et al.
Published: (2023)
Experience as a Compass: Multi-agent RAG with Evolving Orchestration and Agent Prompts
by: Li, Sha, et al.
Published: (2026)
by: Li, Sha, et al.
Published: (2026)
Debiasing Alternative Data for Credit Underwriting Using Causal Inference
by: Lam, Chris
Published: (2024)
by: Lam, Chris
Published: (2024)
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
by: Sun, Jingwei, et al.
Published: (2026)
by: Sun, Jingwei, et al.
Published: (2026)
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
by: Rawles, Christopher, et al.
Published: (2024)
by: Rawles, Christopher, et al.
Published: (2024)
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
by: Froger, Romain, et al.
Published: (2026)
by: Froger, Romain, et al.
Published: (2026)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
To hire or not to hire? Underwriter staffing analysis for AMERICAN Insurance Company
by: Kevin Pan, et al.
Published: (2025)
by: Kevin Pan, et al.
Published: (2025)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
by: Shi, Wentao, et al.
Published: (2026)
by: Shi, Wentao, et al.
Published: (2026)
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
by: Mo, Ying, et al.
Published: (2026)
by: Mo, Ying, et al.
Published: (2026)
Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
by: Jia, Zheng, et al.
Published: (2025)
by: Jia, Zheng, et al.
Published: (2025)
TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents
by: Wang, Dawei, et al.
Published: (2026)
by: Wang, Dawei, et al.
Published: (2026)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
by: Li, Kuan, et al.
Published: (2026)
by: Li, Kuan, et al.
Published: (2026)
BenchMARL: Benchmarking Multi-Agent Reinforcement Learning
by: Bettini, Matteo, et al.
Published: (2023)
by: Bettini, Matteo, et al.
Published: (2023)
InsurAgent: A Large Language Model-Empowered Agent for Simulating Individual Behavior in Purchasing Flood Insurance
by: Geng, Ziheng, et al.
Published: (2025)
by: Geng, Ziheng, et al.
Published: (2025)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
by: Wang, Wenhao, et al.
Published: (2026)
by: Wang, Wenhao, et al.
Published: (2026)
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
by: Han, Yi, et al.
Published: (2026)
by: Han, Yi, et al.
Published: (2026)
PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments
by: Liu, Shuochen, et al.
Published: (2026)
by: Liu, Shuochen, et al.
Published: (2026)
Scaling Multi-Agent Environment Co-Design with Diffusion Models
by: Li, Hao Xiang, et al.
Published: (2025)
by: Li, Hao Xiang, et al.
Published: (2025)
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
by: Gao, Chen, et al.
Published: (2024)
by: Gao, Chen, et al.
Published: (2024)
InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models
by: Ding, Jing, et al.
Published: (2025)
by: Ding, Jing, et al.
Published: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024)
by: Xie, Tianbao, et al.
Published: (2024)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
by: Ji, Haonian, et al.
Published: (2026)
by: Ji, Haonian, et al.
Published: (2026)
A Quantum Approach to Stochastic Optimization in Insurance Underwriting
by: Bordelon, Mitchell, et al.
Published: (2026)
by: Bordelon, Mitchell, et al.
Published: (2026)
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
by: Seo, Gyuhyeon, et al.
Published: (2025)
by: Seo, Gyuhyeon, et al.
Published: (2025)
Classical AI vs. LLMs for Decision-Maker Alignment in Health Insurance Choices
by: Mainali, Mallika, et al.
Published: (2025)
by: Mainali, Mallika, et al.
Published: (2025)
Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment
by: Xu, Binxia, et al.
Published: (2026)
by: Xu, Binxia, et al.
Published: (2026)
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
by: Jiang, Yixing, et al.
Published: (2025)
by: Jiang, Yixing, et al.
Published: (2025)
MDGYM: Benchmarking AI Agents on Molecular Simulations
by: Kumar, Vinay, et al.
Published: (2026)
by: Kumar, Vinay, et al.
Published: (2026)
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
by: Bhalerao, Parth, et al.
Published: (2026)
by: Bhalerao, Parth, et al.
Published: (2026)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
by: Liu, Xinge, et al.
Published: (2026)
by: Liu, Xinge, et al.
Published: (2026)
Similar Items
-
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
by: Dsouza, Amanda, et al.
Published: (2024) -
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique
by: Roy, Joyjit, et al.
Published: (2026) -
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
by: Qi, Zhengyang, et al.
Published: (2026) -
Architecture for Simulating Behavior Mode Changes in Norm-Aware Autonomous Agents
by: Glaze, Sean, et al.
Published: (2025) -
Multi-Step Dialogue Workflow Action Prediction
by: Ramakrishnan, Ramya, et al.
Published: (2023)