Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Choong, Yee-Yin, Greene, Kristen, Qian, Alice, Marasli, Meryem, Yang, Ziqi, Chen, Sophia, Dabbish, Laura, Rao, Anand, Shen, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Worker Discretion Advised: Co-designing Risk Disclosure in Crowdsourced Responsible AI (RAI) Content Work
by: Qian, Alice, et al.
Published: (2025)
by: Qian, Alice, et al.
Published: (2025)
Two Ways to Set Up Wireless Hotspot: Comparing Apples and Oranges
by: Mutch, Andrew, et al.
Published: (2006)
by: Mutch, Andrew, et al.
Published: (2006)
In Comparison: Apples and Oranges and Lemons? Online Elementary Periodical Indices.
by: Levetan, Janice
Published: (1999)
by: Levetan, Janice
Published: (1999)
Locating Risk: Task Designers and the Challenge of Risk Disclosure in RAI Content Work
by: Qian, Alice, et al.
Published: (2025)
by: Qian, Alice, et al.
Published: (2025)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026)
by: Puyin, Li, et al.
Published: (2026)
Understanding Image2Video Domain Shift in Food Segmentation: An Instance-level Analysis on Apples
by: Park, Keonvin, et al.
Published: (2026)
by: Park, Keonvin, et al.
Published: (2026)
A New Simple Vision Algorithm for Detecting the Enzymic Browning Defects in Golden Delicious Apples
by: Balanji, Hamid Majidi
Published: (2021)
by: Balanji, Hamid Majidi
Published: (2021)
Comparing Apples to Oranges: A Taxonomy for Navigating the Global Landscape of AI Regulation
by: Alanoca, Sacha, et al.
Published: (2025)
by: Alanoca, Sacha, et al.
Published: (2025)
Apples in the Apple Library--How One Library Took a Byte.
by: Ertel, Monica
Published: (1983)
by: Ertel, Monica
Published: (1983)
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes
by: Loakman, Tyler, et al.
Published: (2025)
by: Loakman, Tyler, et al.
Published: (2025)
Apples and Oranges and ARL Statistics.
by: Stubbs, Kendon
Published: (1988)
by: Stubbs, Kendon
Published: (1988)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
REALM: A Dataset of Real-World LLM Use Cases
by: Cheng, Jingwen, et al.
Published: (2025)
by: Cheng, Jingwen, et al.
Published: (2025)
Apples to Apples in Jet Quenching: robustness of Machine Learning classification of quenched jets to Underlying Event contamination
by: Gonçalves, João Arruda, et al.
Published: (2025)
by: Gonçalves, João Arruda, et al.
Published: (2025)
Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task
by: Ali, Hassan, et al.
Published: (2024)
by: Ali, Hassan, et al.
Published: (2024)
Apples: Journal of Applied Language Studies
Published: (2011)
Published: (2011)
Robotic Pollination of Apples in Commercial Orchards
by: Sapkota, Ranjan, et al.
Published: (2023)
by: Sapkota, Ranjan, et al.
Published: (2023)
Comparing Apples with Apples: Robust Detection Limits for Exoplanet High-Contrast Imaging in the Presence of non-Gaussian Noise
by: Bonse, Markus J., et al.
Published: (2023)
by: Bonse, Markus J., et al.
Published: (2023)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
by: Hou, Yutao, et al.
Published: (2026)
by: Hou, Yutao, et al.
Published: (2026)
QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
by: Wu, Yao, et al.
Published: (2026)
by: Wu, Yao, et al.
Published: (2026)
Critical Theory: From Michael Apple’s Perspective (Review)
by: Sandra Vega Carrero
Published: (2016)
by: Sandra Vega Carrero
Published: (2016)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
by: Ochieng, Millicent, et al.
Published: (2024)
by: Ochieng, Millicent, et al.
Published: (2024)
Remote Photoplethysmography in Real-World and Extreme Lighting Scenarios
by: Shao, Hang, et al.
Published: (2025)
by: Shao, Hang, et al.
Published: (2025)
Postharvest Preservation of Red Apples Using Edible Coatings and Packaging
by: Mohsen Azadbakht, et al.
Published: (2026)
by: Mohsen Azadbakht, et al.
Published: (2026)
Still "Choosing Our Futures": How Many Apples in the Seed?
by: Neal, James G.
Published: (2015)
by: Neal, James G.
Published: (2015)
Literature Frameworks--From Apples to Zoos. Professional Growth Series.
by: McElmeel, Sharron L.
Published: (1997)
by: McElmeel, Sharron L.
Published: (1997)
Ultra‐Fast Infiltration Behavior of Vacuum Freeze‐Dried Apples
by: Kui Suo, et al.
Published: (2026)
by: Kui Suo, et al.
Published: (2026)
Apples and oranges: Conceptual review as task analysis method1
by: Annemarie van Stee
Published: (2025)
by: Annemarie van Stee
Published: (2025)
WorldEval: World Model as Real-World Robot Policies Evaluator
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
by: Ma, Ziqi, et al.
Published: (2026)
by: Ma, Ziqi, et al.
Published: (2026)
We Should Evaluate Real-World Impact
by: Reiter, Ehud
Published: (2025)
by: Reiter, Ehud
Published: (2025)
Evaluating Memory Capability in Continuous Lifelog Scenario
by: Zheng, Jianjie, et al.
Published: (2026)
by: Zheng, Jianjie, et al.
Published: (2026)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
by: Meng, Jinxiang, et al.
Published: (2026)
by: Meng, Jinxiang, et al.
Published: (2026)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Copiloting Diagnosis of Autism in Real Clinical Scenarios via LLMs
by: Jiang, Yi, et al.
Published: (2024)
by: Jiang, Yi, et al.
Published: (2024)
WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World Scenario
by: Zhang, Jiebin, et al.
Published: (2024)
by: Zhang, Jiebin, et al.
Published: (2024)
UniTalk: Towards Universal Active Speaker Detection in Real World Scenarios
by: Nguyen, Le Thien Phuc, et al.
Published: (2025)
by: Nguyen, Le Thien Phuc, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Similar Items
-
Worker Discretion Advised: Co-designing Risk Disclosure in Crowdsourced Responsible AI (RAI) Content Work
by: Qian, Alice, et al.
Published: (2025) -
Two Ways to Set Up Wireless Hotspot: Comparing Apples and Oranges
by: Mutch, Andrew, et al.
Published: (2006) -
In Comparison: Apples and Oranges and Lemons? Online Elementary Periodical Indices.
by: Levetan, Janice
Published: (1999) -
Locating Risk: Task Designers and the Challenge of Risk Disclosure in RAI Content Work
by: Qian, Alice, et al.
Published: (2025) -
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026)