Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Kangyu, He, Hongliang, Liu, Lin, Liang, Ruiqi, Lan, Zhenzhong, Li, Jianguo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Expressions of Depression and Anxiety in Chinese Psycho-counseling: Usage of First-person Singular Pronoun and Negative Emotional Words
by: Ma, Lizhi, et al.
Published: (2025)
by: Ma, Lizhi, et al.
Published: (2025)
PsyChat: A Client-Centric Dialogue System for Mental Health Support
by: Qiu, Huachuan, et al.
Published: (2023)
by: Qiu, Huachuan, et al.
Published: (2023)
Toward Automated Qualitative Analysis: Leveraging Large Language Models for Tutoring Dialogue Evaluation
by: Gu, Megan, et al.
Published: (2025)
by: Gu, Megan, et al.
Published: (2025)
Exploring Automated Keyword Mnemonics Generation with Large Language Models via Overgenerate-and-Rank
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
Advancing STT for Low-Resource Real-World Speech
by: D'Intino, Flavio, et al.
Published: (2025)
by: D'Intino, Flavio, et al.
Published: (2025)
Exploring the Potential of Large Language Models for Estimating the Reading Comprehension Question Difficulty
by: Jain, Yoshee, et al.
Published: (2025)
by: Jain, Yoshee, et al.
Published: (2025)
Your Co-Workers Matter: Evaluating Collaborative Capabilities of Language Models in Blocks World
by: Wu, Guande, et al.
Published: (2024)
by: Wu, Guande, et al.
Published: (2024)
Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools
by: Hao, Yilun, et al.
Published: (2024)
by: Hao, Yilun, et al.
Published: (2024)
Exploring the Ethical Concerns in User Reviews of Mental Health Apps using Topic Modeling and Sentiment Analysis
by: Rahman, Mohammad Masudur, et al.
Published: (2026)
by: Rahman, Mohammad Masudur, et al.
Published: (2026)
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Will the Real Linda Please Stand up...to Large Language Models? Examining the Representativeness Heuristic in LLMs
by: Wang, Pengda, et al.
Published: (2024)
by: Wang, Pengda, et al.
Published: (2024)
Open-vocabulary Auditory Neural Decoding Using fMRI-prompted LLM
by: Chen, Xiaoyu, et al.
Published: (2024)
by: Chen, Xiaoyu, et al.
Published: (2024)
Evaluating the Usage of African-American Vernacular English in Large Language Models
by: Dunlap, Deja, et al.
Published: (2026)
by: Dunlap, Deja, et al.
Published: (2026)
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
Leveraging Large Language Models for Identifying Knowledge Components
by: Wang, Canwen, et al.
Published: (2025)
by: Wang, Canwen, et al.
Published: (2025)
A Survey on Human-AI Collaboration with Large Foundation Models
by: Vats, Vanshika, et al.
Published: (2024)
by: Vats, Vanshika, et al.
Published: (2024)
TaleFrame: An Interactive Story Generation System with Fine-Grained Control and Large Language Models
by: Wang, Yunchao, et al.
Published: (2025)
by: Wang, Yunchao, et al.
Published: (2025)
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
by: Martin-Boyle, Anna, et al.
Published: (2026)
by: Martin-Boyle, Anna, et al.
Published: (2026)
Large Language Model-based Human-Agent Collaboration for Complex Task Solving
by: Feng, Xueyang, et al.
Published: (2024)
by: Feng, Xueyang, et al.
Published: (2024)
EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal World
by: Ye, Jing, et al.
Published: (2026)
by: Ye, Jing, et al.
Published: (2026)
Evaluating Telugu Proficiency in Large Language Models_ A Comparative Analysis of ChatGPT and Gemini
by: Kishore, Katikela Sreeharsha, et al.
Published: (2024)
by: Kishore, Katikela Sreeharsha, et al.
Published: (2024)
"Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews
by: Wan, Ruyuan, et al.
Published: (2026)
by: Wan, Ruyuan, et al.
Published: (2026)
Think, Act, and Ask: Open-World Interactive Personalized Robot Navigation
by: Dai, Yinpei, et al.
Published: (2023)
by: Dai, Yinpei, et al.
Published: (2023)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
Evaluation Of P300 Speller Performance Using Large Language Models Along With Cross-Subject Training
by: Parthasarathy, Nithin, et al.
Published: (2024)
by: Parthasarathy, Nithin, et al.
Published: (2024)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
by: Ghosh, Himel, et al.
Published: (2026)
by: Ghosh, Himel, et al.
Published: (2026)
Evaluating Large Language Models in Theory of Mind Tasks
by: Kosinski, Michal
Published: (2023)
by: Kosinski, Michal
Published: (2023)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
Do Text-to-Vis Benchmarks Test Real Use of Visualisations?
by: Nguyen, Hy, et al.
Published: (2024)
by: Nguyen, Hy, et al.
Published: (2024)
LLMR: Real-time Prompting of Interactive Worlds using Large Language Models
by: De La Torre, Fernanda, et al.
Published: (2023)
by: De La Torre, Fernanda, et al.
Published: (2023)
LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
A Systematic Review on Prompt Engineering in Large Language Models for K-12 STEM Education
by: Chen, Eason, et al.
Published: (2024)
by: Chen, Eason, et al.
Published: (2024)
Aptly: Making Mobile Apps from Natural Language
by: Patton, Evan W., et al.
Published: (2024)
by: Patton, Evan W., et al.
Published: (2024)
Low-code LLM: Graphical User Interface over Large Language Models
by: Cai, Yuzhe, et al.
Published: (2023)
by: Cai, Yuzhe, et al.
Published: (2023)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
by: Ibrahim, Lujain, et al.
Published: (2025)
by: Ibrahim, Lujain, et al.
Published: (2025)
Interview AI-ssistant: Designing for Real-Time Human-AI Collaboration in Interview Preparation and Execution
by: Liu, Zhe
Published: (2025)
by: Liu, Zhe
Published: (2025)
VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs
by: Xiang, Yurui, et al.
Published: (2026)
by: Xiang, Yurui, et al.
Published: (2026)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
Are Humans as Brittle as Large Language Models?
by: Li, Jiahui, et al.
Published: (2025)
by: Li, Jiahui, et al.
Published: (2025)
Similar Items
-
The Expressions of Depression and Anxiety in Chinese Psycho-counseling: Usage of First-person Singular Pronoun and Negative Emotional Words
by: Ma, Lizhi, et al.
Published: (2025) -
PsyChat: A Client-Centric Dialogue System for Mental Health Support
by: Qiu, Huachuan, et al.
Published: (2023) -
Toward Automated Qualitative Analysis: Leveraging Large Language Models for Tutoring Dialogue Evaluation
by: Gu, Megan, et al.
Published: (2025) -
Exploring Automated Keyword Mnemonics Generation with Large Language Models via Overgenerate-and-Rank
by: Lee, Jaewook, et al.
Published: (2024) -
Advancing STT for Low-Resource Real-World Speech
by: D'Intino, Flavio, et al.
Published: (2025)