Guardado en:
| Autores principales: | Sun, Chongyan, Lin, Ken, Wang, Shiwei, Wu, Hulong, Fu, Chengfei, Wang, Zhen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2408.13338 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
por: Ashktorab, Zahra, et al.
Publicado: (2024)
por: Ashktorab, Zahra, et al.
Publicado: (2024)
QueryGenie: Making LLM-Based Database Querying Transparent and Controllable
por: Chen, Longfei, et al.
Publicado: (2025)
por: Chen, Longfei, et al.
Publicado: (2025)
Holistic Specification of the Human Digital Twin: Stakeholders, Users, Functionalities, and Applications
por: Mandischer, Nils, et al.
Publicado: (2025)
por: Mandischer, Nils, et al.
Publicado: (2025)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
por: Mozannar, Hussein, et al.
Publicado: (2024)
por: Mozannar, Hussein, et al.
Publicado: (2024)
SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care
por: Peng, Dongshen, et al.
Publicado: (2026)
por: Peng, Dongshen, et al.
Publicado: (2026)
Human-Computer Interaction and Visualization in Natural Language Generation Models: Applications, Challenges, and Opportunities
por: Wang, Yunchao, et al.
Publicado: (2024)
por: Wang, Yunchao, et al.
Publicado: (2024)
MI 2 MI: Training Dyad with Collaborative Brain-Computer Interface and Cooperative Motor Imagery Tasks for Better BCI Performance
por: Cheng, Shiwei, et al.
Publicado: (2024)
por: Cheng, Shiwei, et al.
Publicado: (2024)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
por: Kim, Tae Soo, et al.
Publicado: (2023)
por: Kim, Tae Soo, et al.
Publicado: (2023)
VisEval: A Benchmark for Data Visualization in the Era of Large Language Models
por: Chen, Nan, et al.
Publicado: (2024)
por: Chen, Nan, et al.
Publicado: (2024)
ComViewer: An Interactive Visual Tool to Help Viewers Seek Social Support in Online Mental Health Communities
por: Wu, Shiwei, et al.
Publicado: (2024)
por: Wu, Shiwei, et al.
Publicado: (2024)
Tangible Scenography as a Holistic Design Method for Human-Robot Interaction
por: Koike, Amy, et al.
Publicado: (2024)
por: Koike, Amy, et al.
Publicado: (2024)
Raiven: LLM-Based Visualization Authoring via Domain-Specific Language Mediation
por: Irger, Alexandra, et al.
Publicado: (2026)
por: Irger, Alexandra, et al.
Publicado: (2026)
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
por: Ashktorab, Zahra, et al.
Publicado: (2025)
por: Ashktorab, Zahra, et al.
Publicado: (2025)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
por: Lee, Suhyun, et al.
Publicado: (2026)
por: Lee, Suhyun, et al.
Publicado: (2026)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
por: Liu, Tianjian, et al.
Publicado: (2025)
por: Liu, Tianjian, et al.
Publicado: (2025)
SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation
por: Zhou, Joyce, et al.
Publicado: (2026)
por: Zhou, Joyce, et al.
Publicado: (2026)
Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition
por: Feng, Kehua, et al.
Publicado: (2024)
por: Feng, Kehua, et al.
Publicado: (2024)
Visualization Generation with Large Language Models: An Evaluation
por: Wang, Xinyu, et al.
Publicado: (2024)
por: Wang, Xinyu, et al.
Publicado: (2024)
Facilitating Holistic Evaluations with LLMs: Insights from Scenario-Based Experiments
por: Ishida, Toru, et al.
Publicado: (2024)
por: Ishida, Toru, et al.
Publicado: (2024)
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
por: Yao, Bingsheng, et al.
Publicado: (2026)
por: Yao, Bingsheng, et al.
Publicado: (2026)
FlowEval: Reference-based Evaluation of Generated User Interfaces
por: Wu, Jason, et al.
Publicado: (2026)
por: Wu, Jason, et al.
Publicado: (2026)
Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework
por: Chakraborty, Mohna, et al.
Publicado: (2025)
por: Chakraborty, Mohna, et al.
Publicado: (2025)
Towards Holistic Prompt Craft
por: Lindley, Joseph, et al.
Publicado: (2025)
por: Lindley, Joseph, et al.
Publicado: (2025)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
por: Chiang, Charles, et al.
Publicado: (2026)
por: Chiang, Charles, et al.
Publicado: (2026)
TaleFrame: An Interactive Story Generation System with Fine-Grained Control and Large Language Models
por: Wang, Yunchao, et al.
Publicado: (2025)
por: Wang, Yunchao, et al.
Publicado: (2025)
GlyphWeaver: Unlocking Glyph Design Creativity with Uniform Glyph DSL and AI
por: Liu, Can, et al.
Publicado: (2025)
por: Liu, Can, et al.
Publicado: (2025)
Exploring the Usage of Generative AI for Group Project-Based Offline Art Courses in Elementary Schools
por: Wang, Zhiqing, et al.
Publicado: (2025)
por: Wang, Zhiqing, et al.
Publicado: (2025)
Leveraging Large Language Models for Identifying Knowledge Components
por: Wang, Canwen, et al.
Publicado: (2025)
por: Wang, Canwen, et al.
Publicado: (2025)
Human and LLM-Based Voice Assistant Interaction: An Analytical Framework for User Verbal and Nonverbal Behaviors
por: Chan, Szeyi, et al.
Publicado: (2024)
por: Chan, Szeyi, et al.
Publicado: (2024)
Are Humans as Brittle as Large Language Models?
por: Li, Jiahui, et al.
Publicado: (2025)
por: Li, Jiahui, et al.
Publicado: (2025)
Generating Analytic Specifications for Data Visualization from Natural Language Queries using Large Language Models
por: Sah, Subham, et al.
Publicado: (2024)
por: Sah, Subham, et al.
Publicado: (2024)
Leveraging Large Language Models for Generating Mobile Sensing Strategies in Human Behavior Modeling
por: Gao, Nan, et al.
Publicado: (2023)
por: Gao, Nan, et al.
Publicado: (2023)
EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
por: Lian, Zheng, et al.
Publicado: (2025)
por: Lian, Zheng, et al.
Publicado: (2025)
Understanding User Experience in Large Language Model Interactions
por: Wang, Jiayin, et al.
Publicado: (2024)
por: Wang, Jiayin, et al.
Publicado: (2024)
Evaluation of Large Language Model-Driven AutoML in Data and Model Management from Human-Centered Perspective
por: Yao, Jiapeng, et al.
Publicado: (2025)
por: Yao, Jiapeng, et al.
Publicado: (2025)
Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model
por: Jayakumar, Eswari, et al.
Publicado: (2025)
por: Jayakumar, Eswari, et al.
Publicado: (2025)
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models
por: Wang, Zhiyuan, et al.
Publicado: (2024)
por: Wang, Zhiyuan, et al.
Publicado: (2024)
MindChat: Enhancing BCI Spelling with Large Language Models in Realistic Scenarios
por: Wang, JIaheng, et al.
Publicado: (2025)
por: Wang, JIaheng, et al.
Publicado: (2025)
Large Language Model-based Human-Agent Collaboration for Complex Task Solving
por: Feng, Xueyang, et al.
Publicado: (2024)
por: Feng, Xueyang, et al.
Publicado: (2024)
Beyond Text: Probing K-12 Educators' Perspectives and Ideas for Learning Opportunities Leveraging Multimodal Large Language Models
por: Tseng, Tiffany, et al.
Publicado: (2025)
por: Tseng, Tiffany, et al.
Publicado: (2025)
Ejemplares similares
-
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
por: Ashktorab, Zahra, et al.
Publicado: (2024) -
QueryGenie: Making LLM-Based Database Querying Transparent and Controllable
por: Chen, Longfei, et al.
Publicado: (2025) -
Holistic Specification of the Human Digital Twin: Stakeholders, Users, Functionalities, and Applications
por: Mandischer, Nils, et al.
Publicado: (2025) -
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
por: Mozannar, Hussein, et al.
Publicado: (2024) -
SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care
por: Peng, Dongshen, et al.
Publicado: (2026)