How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Karizaki, Mahsa Sheikhi, Gnesdilow, Dana, Puntambekar, Sadhana, Passonneau, Rebecca J. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VerAs: Verify then Assess STEM Lab Reports
par: Atil, Berk, et autres
Publié: (2024)
par: Atil, Berk, et autres
Publié: (2024)
Middle School Students' Application of Science Learning From Physical Versus Virtual Labs to New Contexts
par: Dana Gnesdilow, et autres
Publié: (2025)
par: Dana Gnesdilow, et autres
Publié: (2025)
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
par: Knipper, R. Alexander, et autres
Publié: (2025)
par: Knipper, R. Alexander, et autres
Publié: (2025)
Joint Training for Selective Prediction
par: Li, Zhaohui, et autres
Publié: (2024)
par: Li, Zhaohui, et autres
Publié: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
par: Atil, Berk, et autres
Publié: (2025)
par: Atil, Berk, et autres
Publié: (2025)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
par: Atil, Berk, et autres
Publié: (2026)
par: Atil, Berk, et autres
Publié: (2026)
Model Unlearning Objectives Vary for Distinct Language Functions
par: Atil, Berk, et autres
Publié: (2026)
par: Atil, Berk, et autres
Publié: (2026)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
par: Atil, Berk, et autres
Publié: (2025)
par: Atil, Berk, et autres
Publié: (2025)
Something Just Like TRuST : Toxicity Recognition of Span and Target
par: Atil, Berk, et autres
Publié: (2025)
par: Atil, Berk, et autres
Publié: (2025)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
par: Gupta, Vipul, et autres
Publié: (2023)
par: Gupta, Vipul, et autres
Publié: (2023)
Sociodemographic Bias in Language Models: A Survey and Forward Path
par: Gupta, Vipul, et autres
Publié: (2023)
par: Gupta, Vipul, et autres
Publié: (2023)
Chain of Simulation: A Dual-Mode Reasoning Framework for Large Language Models with Dynamic Problem Routing
par: Sheikhi, Saeid
Publié: (2026)
par: Sheikhi, Saeid
Publié: (2026)
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
par: Yang, Sohee, et autres
Publié: (2025)
par: Yang, Sohee, et autres
Publié: (2025)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
par: Ge, Huaizhi, et autres
Publié: (2024)
par: Ge, Huaizhi, et autres
Publié: (2024)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
par: Bhaskar, Adithya, et autres
Publié: (2025)
par: Bhaskar, Adithya, et autres
Publié: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
par: Gupta, Vipul, et autres
Publié: (2024)
par: Gupta, Vipul, et autres
Publié: (2024)
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
par: Sahoo, Subramanyam, et autres
Publié: (2025)
par: Sahoo, Subramanyam, et autres
Publié: (2025)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
par: Ferrer, Robinson, et autres
Publié: (2026)
par: Ferrer, Robinson, et autres
Publié: (2026)
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
par: Schopf, Tim, et autres
Publié: (2026)
par: Schopf, Tim, et autres
Publié: (2026)
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
par: Bianchi, Federico, et autres
Publié: (2024)
par: Bianchi, Federico, et autres
Publié: (2024)
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO
par: Ng, Man Tik, et autres
Publié: (2024)
par: Ng, Man Tik, et autres
Publié: (2024)
You Can't Fight in Here! This is BBS!
par: Futrell, Richard, et autres
Publié: (2026)
par: Futrell, Richard, et autres
Publié: (2026)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
par: Pletenev, Sergey, et autres
Publié: (2025)
par: Pletenev, Sergey, et autres
Publié: (2025)
How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse
par: Imai, Saki, et autres
Publié: (2025)
par: Imai, Saki, et autres
Publié: (2025)
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
par: Khallaf, Nouran, et autres
Publié: (2026)
par: Khallaf, Nouran, et autres
Publié: (2026)
How Well Can We Decode Vowels from Auditory EEG -- A Rigorous Cross-Subject Benchmark with Honest Assessment
par: Li, Xiaoyang
Publié: (2026)
par: Li, Xiaoyang
Publié: (2026)
How Many Bytes Can You Take Out Of Brain-To-Text Decoding?
par: Antonello, Richard, et autres
Publié: (2024)
par: Antonello, Richard, et autres
Publié: (2024)
Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
par: Sheikhi, Hadi, et autres
Publié: (2025)
par: Sheikhi, Hadi, et autres
Publié: (2025)
The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
par: Retkowski, Fabian, et autres
Publié: (2025)
par: Retkowski, Fabian, et autres
Publié: (2025)
Concept-based Rubrics Improve LLM Formative Assessment and Data Synthesis
par: Wei, Yuchen, et autres
Publié: (2025)
par: Wei, Yuchen, et autres
Publié: (2025)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
par: Li, Yuxuan, et autres
Publié: (2026)
par: Li, Yuxuan, et autres
Publié: (2026)
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
par: Baumler, Connor, et autres
Publié: (2026)
par: Baumler, Connor, et autres
Publié: (2026)
Can ChatGPT Read Who You Are?
par: Derner, Erik, et autres
Publié: (2023)
par: Derner, Erik, et autres
Publié: (2023)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
par: Hwang, Yerin, et autres
Publié: (2025)
par: Hwang, Yerin, et autres
Publié: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
par: Chopra, Muskaan, et autres
Publié: (2025)
par: Chopra, Muskaan, et autres
Publié: (2025)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
par: Ashihara, Takanori, et autres
Publié: (2023)
par: Ashihara, Takanori, et autres
Publié: (2023)
SAEs Are Good for Steering -- If You Select the Right Features
par: Arad, Dana, et autres
Publié: (2025)
par: Arad, Dana, et autres
Publié: (2025)
Automating Date Format Detection for Data Visualization
par: Liang, Zixuan
Publié: (2025)
par: Liang, Zixuan
Publié: (2025)
sebis at ArchEHR-QA 2026: How Much Can You Do Locally? Evaluating Grounded EHR QA on a Single Notebook
par: Yurt, Ibrahim Ebrar, et autres
Publié: (2026)
par: Yurt, Ibrahim Ebrar, et autres
Publié: (2026)
"I know myself better, but not really greatly": How Well Can LLMs Detect and Explain LLM-Generated Texts?
par: Ji, Jiazhou, et autres
Publié: (2025)
par: Ji, Jiazhou, et autres
Publié: (2025)
Documents similaires
-
VerAs: Verify then Assess STEM Lab Reports
par: Atil, Berk, et autres
Publié: (2024) -
Middle School Students' Application of Science Learning From Physical Versus Virtual Labs to New Contexts
par: Dana Gnesdilow, et autres
Publié: (2025) -
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
par: Knipper, R. Alexander, et autres
Publié: (2025) -
Joint Training for Selective Prediction
par: Li, Zhaohui, et autres
Publié: (2024) -
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
par: Atil, Berk, et autres
Publié: (2025)