OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Opsahl-Ong, Krista, Singhvi, Arnav, Collins, Jasmine, Zhou, Ivan, Wang, Cindy, Baheti, Ashutosh, Oertell, Owen, Portes, Jacob, Havens, Sam, Elsen, Erich, Bendersky, Michael, Zaharia, Matei, Chen, Xing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KARL: Knowledge Agents via Reinforcement Learning
by: Chang, Jonathan D., et al.
Published: (2026)
by: Chang, Jonathan D., et al.
Published: (2026)
Long Context RAG Performance of Large Language Models
by: Leng, Quinn, et al.
Published: (2024)
by: Leng, Quinn, et al.
Published: (2024)
Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
by: Opsahl-Ong, Krista, et al.
Published: (2024)
by: Opsahl-Ong, Krista, et al.
Published: (2024)
LangProBe: a Language Programs Benchmark
by: Tan, Shangyin, et al.
Published: (2025)
by: Tan, Shangyin, et al.
Published: (2025)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
by: Singhvi, Arnav, et al.
Published: (2023)
by: Singhvi, Arnav, et al.
Published: (2023)
Automating the Enterprise with Foundation Models
by: Wornow, Michael, et al.
Published: (2024)
by: Wornow, Michael, et al.
Published: (2024)
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization
by: Zamani, Hamed, et al.
Published: (2024)
by: Zamani, Hamed, et al.
Published: (2024)
A State-of-the-Art SQL Reasoning Model using RLVR
by: Ali, Alnur, et al.
Published: (2025)
by: Ali, Alnur, et al.
Published: (2025)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
by: Agrawal, Lakshya A, et al.
Published: (2025)
by: Agrawal, Lakshya A, et al.
Published: (2025)
The End of Modernism
by: Collins Donahue, William
Published: (2020)
by: Collins Donahue, William
Published: (2020)
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Routing End User Queries to Enterprise Databases
by: Sudarshan, Saikrishna, et al.
Published: (2026)
by: Sudarshan, Saikrishna, et al.
Published: (2026)
SIEVE: Sample-Efficient Parametric Learning from Natural Language
by: Asawa, Parth, et al.
Published: (2026)
by: Asawa, Parth, et al.
Published: (2026)
Convergence Of Consistency Model With Multistep Sampling Under General Data Assumptions
by: Chen, Yiding, et al.
Published: (2025)
by: Chen, Yiding, et al.
Published: (2025)
ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
by: Velampalli, Sirisha, et al.
Published: (2025)
by: Velampalli, Sirisha, et al.
Published: (2025)
ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data
by: Patel, Liana, et al.
Published: (2024)
by: Patel, Liana, et al.
Published: (2024)
World Model on Million-Length Video And Language With Blockwise RingAttention
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
End-to-End Learning of Correlated Operating Reserve Requirements in Security-Constrained Economic Dispatch
by: Shen, Owen, et al.
Published: (2026)
by: Shen, Owen, et al.
Published: (2026)
CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
by: Xie, Jingchao, et al.
Published: (2025)
by: Xie, Jingchao, et al.
Published: (2025)
Prompt Triage: Structured Optimization Enhances Vision-Language Model Performance on Medical Imaging Benchmarks
by: Singhvi, Arnav, et al.
Published: (2025)
by: Singhvi, Arnav, et al.
Published: (2025)
Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals
by: Opsahl, Tobias A.
Published: (2024)
by: Opsahl, Tobias A.
Published: (2024)
Optimally Decoding Two-Dimensional Reed-Solomon Codes Against Deletion Errors
by: Singhvi, Shubhransh
Published: (2024)
by: Singhvi, Shubhransh
Published: (2024)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
TurboHopp: Accelerated Molecule Scaffold Hopping with Consistency Models
by: Yoo, Kiwoong, et al.
Published: (2024)
by: Yoo, Kiwoong, et al.
Published: (2024)
The Epidemic of Men Medically Unfit for Transurethral Resection of the Prostate Has Come to an End
by: Henry H. Woo, et al.
Published: (2025)
by: Henry H. Woo, et al.
Published: (2025)
"Making Money is not an End in Itself": Creating Meaningfulness among Employees of Social Enterprises
by: Christian Franklin Svensson
Published: (2014)
by: Christian Franklin Svensson
Published: (2014)
Black Television Travels
by: Havens, Timothy
Published: (2024)
by: Havens, Timothy
Published: (2024)
MIRVs and Money
by: Havens, Shirley
Published: (1969)
by: Havens, Shirley
Published: (1969)
Libraries: Drawn to Knowledge.
by: Havens, Kevin
Published: (2003)
by: Havens, Kevin
Published: (2003)
The Case for Aging: A Bibliographic Essay.
by: Havens, Carolyn
Published: (1988)
by: Havens, Carolyn
Published: (1988)
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
by: Li, Mohan, et al.
Published: (2024)
by: Li, Mohan, et al.
Published: (2024)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
by: Oertell, Owen, et al.
Published: (2024)
by: Oertell, Owen, et al.
Published: (2024)
EndToEndML: An Open-Source End-to-End Pipeline for Machine Learning Applications
by: Pillai, Nisha, et al.
Published: (2024)
by: Pillai, Nisha, et al.
Published: (2024)
Evidence gaps in orthognathic surgery, a Delphi study protocol
by: Josefina Bendersky
Published: (2023)
by: Josefina Bendersky
Published: (2023)
End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction
by: Li, Xi, et al.
Published: (2025)
by: Li, Xi, et al.
Published: (2025)
The Impact of Green Finance on Enterprise Environmental Strategies: Source Prevention or End‐Of‐Pipe Treatment?
by: Lin Wang, et al.
Published: (2025)
by: Lin Wang, et al.
Published: (2025)
Similar Items
-
KARL: Knowledge Agents via Reinforcement Learning
by: Chang, Jonathan D., et al.
Published: (2026) -
Long Context RAG Performance of Large Language Models
by: Leng, Quinn, et al.
Published: (2024) -
Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
by: Opsahl-Ong, Krista, et al.
Published: (2024) -
LangProBe: a Language Programs Benchmark
by: Tan, Shangyin, et al.
Published: (2025) -
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)