LLM Prompt Evaluation for Educational Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Holmes, Langdon, Coscia, Adam, Crossley, Scott, Choi, Joon Suh, Morris, Wesley |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries
by: Coscia, Adam, et al.
Published: (2024)
by: Coscia, Adam, et al.
Published: (2024)
Improving Domain-Specific ASR with LLM-Generated Contextual Descriptions
by: Suh, Jiwon, et al.
Published: (2024)
by: Suh, Jiwon, et al.
Published: (2024)
KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts
by: Coscia, Adam, et al.
Published: (2024)
by: Coscia, Adam, et al.
Published: (2024)
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
by: Pei, Aihua, et al.
Published: (2024)
by: Pei, Aihua, et al.
Published: (2024)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
by: Cho, Yousang, et al.
Published: (2025)
by: Cho, Yousang, et al.
Published: (2025)
When "Better" Prompts Hurt: Evaluation-Driven Iteration for LLM Applications
by: Commey, Daniel
Published: (2026)
by: Commey, Daniel
Published: (2026)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
by: Abeysinghe, Bhashithe, et al.
Published: (2024)
by: Abeysinghe, Bhashithe, et al.
Published: (2024)
JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
by: Lin, Leonard, et al.
Published: (2026)
by: Lin, Leonard, et al.
Published: (2026)
StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
by: Jang, Geonhui, et al.
Published: (2026)
by: Jang, Geonhui, et al.
Published: (2026)
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
by: Soh, Yun Joon, et al.
Published: (2024)
by: Soh, Yun Joon, et al.
Published: (2024)
Prompt Sentiment: The Catalyst for LLM Change
by: Gandhi, Vishal, et al.
Published: (2025)
by: Gandhi, Vishal, et al.
Published: (2025)
Enhancing LLM Agent Safety via Causal Influence Prompting
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
COPAL: Continual Pruning in Large Language Generative Models
by: Malla, Srikanth, et al.
Published: (2024)
by: Malla, Srikanth, et al.
Published: (2024)
Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation
by: Yoo, YoungJoon, et al.
Published: (2023)
by: Yoo, YoungJoon, et al.
Published: (2023)
ADEPT: Adaptive Dynamic Early-Exit Process for Transformers
by: Yoo, Sangmin, et al.
Published: (2026)
by: Yoo, Sangmin, et al.
Published: (2026)
Parse Trees Guided LLM Prompt Compression
by: Mao, Wenhao, et al.
Published: (2024)
by: Mao, Wenhao, et al.
Published: (2024)
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
by: Tsai, Yu-Che, et al.
Published: (2026)
by: Tsai, Yu-Che, et al.
Published: (2026)
Prompt Recursive Search: A Living Framework with Adaptive Growth in LLM Auto-Prompting
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
Generative Prompt Internalization
by: Shin, Haebin, et al.
Published: (2024)
by: Shin, Haebin, et al.
Published: (2024)
LLM Agents for Education: Advances and Applications
by: Chu, Zhendong, et al.
Published: (2025)
by: Chu, Zhendong, et al.
Published: (2025)
LinkQ: An LLM-Assisted Visual Interface for Knowledge Graph Question-Answering
by: Li, Harry, et al.
Published: (2024)
by: Li, Harry, et al.
Published: (2024)
Multi-LLM Thematic Analysis with Dual Reliability Metrics: Combining Cohen's Kappa and Semantic Similarity for Qualitative Research Validation
by: Jain, Nilesh, et al.
Published: (2025)
by: Jain, Nilesh, et al.
Published: (2025)
Conveying Imagistic Thinking in Traditional Chinese Medicine Translation: A Prompt Engineering and LLM-Based Evaluation Framework
by: Han, Jiatong
Published: (2025)
by: Han, Jiatong
Published: (2025)
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
by: Kanithi, Praveenkumar, et al.
Published: (2024)
by: Kanithi, Praveenkumar, et al.
Published: (2024)
SCOPE: A Generative Approach for LLM Prompt Compression
by: Zhang, Tinghui, et al.
Published: (2025)
by: Zhang, Tinghui, et al.
Published: (2025)
Efficient Prompting for LLM-based Generative Internet of Things
by: Xiao, Bin, et al.
Published: (2024)
by: Xiao, Bin, et al.
Published: (2024)
Trustworthy LLM-Mediated Communication: Evaluating Information Fidelity in LLM as a Communicator (LAAC) Framework in Multiple Application Domains
by: Rafi, Mohammed Musthafa, et al.
Published: (2025)
by: Rafi, Mohammed Musthafa, et al.
Published: (2025)
Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
by: Choi, Eunseong, et al.
Published: (2025)
by: Choi, Eunseong, et al.
Published: (2025)
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator
by: Sordo, Alessio, et al.
Published: (2026)
by: Sordo, Alessio, et al.
Published: (2026)
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
by: Bhattacharyya, Chaitali, et al.
Published: (2025)
by: Bhattacharyya, Chaitali, et al.
Published: (2025)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization
by: Chung, John Joon Young, et al.
Published: (2024)
by: Chung, John Joon Young, et al.
Published: (2024)
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
by: Arnaudo, Anna, et al.
Published: (2026)
by: Arnaudo, Anna, et al.
Published: (2026)
Can Separators Improve Chain-of-Thought Prompting?
by: Park, Yoonjeong, et al.
Published: (2024)
by: Park, Yoonjeong, et al.
Published: (2024)
Prompt Injection attack against LLM-integrated Applications
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications
by: Kogan, David, et al.
Published: (2025)
by: Kogan, David, et al.
Published: (2025)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Improving LLM Abilities in Idiomatic Translation
by: Donthi, Sundesh, et al.
Published: (2024)
by: Donthi, Sundesh, et al.
Published: (2024)
Submodular Evaluation Subset Selection in Automatic Prompt Optimization
by: Nian, Jinming, et al.
Published: (2026)
by: Nian, Jinming, et al.
Published: (2026)
Similar Items
-
iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries
by: Coscia, Adam, et al.
Published: (2024) -
Improving Domain-Specific ASR with LLM-Generated Contextual Descriptions
by: Suh, Jiwon, et al.
Published: (2024) -
KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts
by: Coscia, Adam, et al.
Published: (2024) -
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
by: Pei, Aihua, et al.
Published: (2024) -
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)