DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shrivastava, Aryan, Aoyagui, Paula Akemi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Subjectivity for more Human-Centric Assessment of Social Biases in Large Language Models
by: Aoyagui, Paula Akemi, et al.
Published: (2024)
by: Aoyagui, Paula Akemi, et al.
Published: (2024)
Just Like Me: The Role of Opinions and Personal Experiences in The Perception of Explanations in Subjective Decision-Making
by: Ferguson, Sharon, et al.
Published: (2024)
by: Ferguson, Sharon, et al.
Published: (2024)
A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
by: Aoyagui, Paula Akemi, et al.
Published: (2025)
by: Aoyagui, Paula Akemi, et al.
Published: (2025)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025)
by: He, Keyu, et al.
Published: (2025)
Large Language Lovers: Lived Experiences of Negotiating Agency and Platform Control in AI Companionship
by: Lee, Patrick Yung Kang, et al.
Published: (2026)
by: Lee, Patrick Yung Kang, et al.
Published: (2026)
The Art of Saying No: Contextual Noncompliance in Language Models
by: Brahman, Faeze, et al.
Published: (2024)
by: Brahman, Faeze, et al.
Published: (2024)
A Scalable Framework for Evaluating Health Language Models
by: Mallinar, Neil, et al.
Published: (2025)
by: Mallinar, Neil, et al.
Published: (2025)
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
by: Tang, Zixin, et al.
Published: (2025)
by: Tang, Zixin, et al.
Published: (2025)
Evaluating Adaptive Personalization of Educational Readings with Simulated Learners
by: Woo, Ryan T., et al.
Published: (2026)
by: Woo, Ryan T., et al.
Published: (2026)
Hybrid EEG--Driven Brain--Computer Interface: A Large Language Model Framework for Personalized Language Rehabilitation
by: Hossain, Ismail, et al.
Published: (2025)
by: Hossain, Ismail, et al.
Published: (2025)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
by: Yang, Jackie Junrui, et al.
Published: (2023)
by: Yang, Jackie Junrui, et al.
Published: (2023)
Evaluating the Usage of African-American Vernacular English in Large Language Models
by: Dunlap, Deja, et al.
Published: (2026)
by: Dunlap, Deja, et al.
Published: (2026)
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control
by: Chittem, Adithya, et al.
Published: (2025)
by: Chittem, Adithya, et al.
Published: (2025)
Evaluating Telugu Proficiency in Large Language Models_ A Comparative Analysis of ChatGPT and Gemini
by: Kishore, Katikela Sreeharsha, et al.
Published: (2024)
by: Kishore, Katikela Sreeharsha, et al.
Published: (2024)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
by: Yu, Yeyong, et al.
Published: (2024)
by: Yu, Yeyong, et al.
Published: (2024)
Beyond Models: A Framework for Contextual and Cultural Intelligence in African AI Deployment
by: Ndlovu, Qness
Published: (2025)
by: Ndlovu, Qness
Published: (2025)
Toward Automated Qualitative Analysis: Leveraging Large Language Models for Tutoring Dialogue Evaluation
by: Gu, Megan, et al.
Published: (2025)
by: Gu, Megan, et al.
Published: (2025)
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
by: Martin-Boyle, Anna, et al.
Published: (2026)
by: Martin-Boyle, Anna, et al.
Published: (2026)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
by: Sun, Chongyan, et al.
Published: (2024)
by: Sun, Chongyan, et al.
Published: (2024)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
by: Merchant, Zain, et al.
Published: (2024)
by: Merchant, Zain, et al.
Published: (2024)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Evaluation Of P300 Speller Performance Using Large Language Models Along With Cross-Subject Training
by: Parthasarathy, Nithin, et al.
Published: (2024)
by: Parthasarathy, Nithin, et al.
Published: (2024)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
by: Afzoon, Saleh, et al.
Published: (2026)
by: Afzoon, Saleh, et al.
Published: (2026)
Evaluating Large Language Models in Theory of Mind Tasks
by: Kosinski, Michal
Published: (2023)
by: Kosinski, Michal
Published: (2023)
A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
by: Sarıtaş, Karahan, et al.
Published: (2025)
by: Sarıtaş, Karahan, et al.
Published: (2025)
Evaluation of a Sign Language Avatar on Comprehensibility, User Experience \& Acceptability
by: Wasserroth, Fenya, et al.
Published: (2025)
by: Wasserroth, Fenya, et al.
Published: (2025)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
by: Ibrahim, Lujain, et al.
Published: (2025)
by: Ibrahim, Lujain, et al.
Published: (2025)
Are Humans as Brittle as Large Language Models?
by: Li, Jiahui, et al.
Published: (2025)
by: Li, Jiahui, et al.
Published: (2025)
Aligning Language Models with Demonstrated Feedback
by: Shaikh, Omar, et al.
Published: (2024)
by: Shaikh, Omar, et al.
Published: (2024)
Grounding Gaps in Language Model Generations
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models
by: Tran, Son Quoc, et al.
Published: (2025)
by: Tran, Son Quoc, et al.
Published: (2025)
Large Language Models for Depression Recognition in Spoken Language Integrating Psychological Knowledge
by: Li, Yupei, et al.
Published: (2025)
by: Li, Yupei, et al.
Published: (2025)
An Evaluation of Estimative Uncertainty in Large Language Models
by: Tang, Zhisheng, et al.
Published: (2024)
by: Tang, Zhisheng, et al.
Published: (2024)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Similar Items
-
Exploring Subjectivity for more Human-Centric Assessment of Social Biases in Large Language Models
by: Aoyagui, Paula Akemi, et al.
Published: (2024) -
Just Like Me: The Role of Opinions and Personal Experiences in The Perception of Explanations in Subjective Decision-Making
by: Ferguson, Sharon, et al.
Published: (2024) -
A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
by: Aoyagui, Paula Akemi, et al.
Published: (2025) -
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025) -
Large Language Lovers: Lived Experiences of Negotiating Agency and Platform Control in AI Companionship
by: Lee, Patrick Yung Kang, et al.
Published: (2026)