How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
Fuente:
arXiv
Guardado en:
| Autores principales: | Karizaki, Mahsa Sheikhi, Gnesdilow, Dana, Puntambekar, Sadhana, Passonneau, Rebecca J. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VerAs: Verify then Assess STEM Lab Reports
por: Atil, Berk, et al.
Publicado: (2024)
por: Atil, Berk, et al.
Publicado: (2024)
Middle School Students' Application of Science Learning From Physical Versus Virtual Labs to New Contexts
por: Dana Gnesdilow, et al.
Publicado: (2025)
por: Dana Gnesdilow, et al.
Publicado: (2025)
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
por: Knipper, R. Alexander, et al.
Publicado: (2025)
por: Knipper, R. Alexander, et al.
Publicado: (2025)
Joint Training for Selective Prediction
por: Li, Zhaohui, et al.
Publicado: (2024)
por: Li, Zhaohui, et al.
Publicado: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Model Unlearning Objectives Vary for Distinct Language Functions
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
Something Just Like TRuST : Toxicity Recognition of Span and Target
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
por: Gupta, Vipul, et al.
Publicado: (2023)
por: Gupta, Vipul, et al.
Publicado: (2023)
Sociodemographic Bias in Language Models: A Survey and Forward Path
por: Gupta, Vipul, et al.
Publicado: (2023)
por: Gupta, Vipul, et al.
Publicado: (2023)
Chain of Simulation: A Dual-Mode Reasoning Framework for Large Language Models with Dynamic Problem Routing
por: Sheikhi, Saeid
Publicado: (2026)
por: Sheikhi, Saeid
Publicado: (2026)
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
por: Yang, Sohee, et al.
Publicado: (2025)
por: Yang, Sohee, et al.
Publicado: (2025)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
por: Ge, Huaizhi, et al.
Publicado: (2024)
por: Ge, Huaizhi, et al.
Publicado: (2024)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
por: Bhaskar, Adithya, et al.
Publicado: (2025)
por: Bhaskar, Adithya, et al.
Publicado: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
por: Gupta, Vipul, et al.
Publicado: (2024)
por: Gupta, Vipul, et al.
Publicado: (2024)
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
por: Sahoo, Subramanyam, et al.
Publicado: (2025)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
por: Ferrer, Robinson, et al.
Publicado: (2026)
por: Ferrer, Robinson, et al.
Publicado: (2026)
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
por: Schopf, Tim, et al.
Publicado: (2026)
por: Schopf, Tim, et al.
Publicado: (2026)
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
por: Bianchi, Federico, et al.
Publicado: (2024)
por: Bianchi, Federico, et al.
Publicado: (2024)
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO
por: Ng, Man Tik, et al.
Publicado: (2024)
por: Ng, Man Tik, et al.
Publicado: (2024)
You Can't Fight in Here! This is BBS!
por: Futrell, Richard, et al.
Publicado: (2026)
por: Futrell, Richard, et al.
Publicado: (2026)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
por: Pletenev, Sergey, et al.
Publicado: (2025)
por: Pletenev, Sergey, et al.
Publicado: (2025)
How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse
por: Imai, Saki, et al.
Publicado: (2025)
por: Imai, Saki, et al.
Publicado: (2025)
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
por: Khallaf, Nouran, et al.
Publicado: (2026)
por: Khallaf, Nouran, et al.
Publicado: (2026)
How Well Can We Decode Vowels from Auditory EEG -- A Rigorous Cross-Subject Benchmark with Honest Assessment
por: Li, Xiaoyang
Publicado: (2026)
por: Li, Xiaoyang
Publicado: (2026)
How Many Bytes Can You Take Out Of Brain-To-Text Decoding?
por: Antonello, Richard, et al.
Publicado: (2024)
por: Antonello, Richard, et al.
Publicado: (2024)
Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
por: Sheikhi, Hadi, et al.
Publicado: (2025)
por: Sheikhi, Hadi, et al.
Publicado: (2025)
The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
por: Retkowski, Fabian, et al.
Publicado: (2025)
por: Retkowski, Fabian, et al.
Publicado: (2025)
Concept-based Rubrics Improve LLM Formative Assessment and Data Synthesis
por: Wei, Yuchen, et al.
Publicado: (2025)
por: Wei, Yuchen, et al.
Publicado: (2025)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
por: Li, Yuxuan, et al.
Publicado: (2026)
por: Li, Yuxuan, et al.
Publicado: (2026)
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
por: Baumler, Connor, et al.
Publicado: (2026)
por: Baumler, Connor, et al.
Publicado: (2026)
Can ChatGPT Read Who You Are?
por: Derner, Erik, et al.
Publicado: (2023)
por: Derner, Erik, et al.
Publicado: (2023)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
por: Hwang, Yerin, et al.
Publicado: (2025)
por: Hwang, Yerin, et al.
Publicado: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
por: Chopra, Muskaan, et al.
Publicado: (2025)
por: Chopra, Muskaan, et al.
Publicado: (2025)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
por: Ashihara, Takanori, et al.
Publicado: (2023)
por: Ashihara, Takanori, et al.
Publicado: (2023)
SAEs Are Good for Steering -- If You Select the Right Features
por: Arad, Dana, et al.
Publicado: (2025)
por: Arad, Dana, et al.
Publicado: (2025)
Automating Date Format Detection for Data Visualization
por: Liang, Zixuan
Publicado: (2025)
por: Liang, Zixuan
Publicado: (2025)
sebis at ArchEHR-QA 2026: How Much Can You Do Locally? Evaluating Grounded EHR QA on a Single Notebook
por: Yurt, Ibrahim Ebrar, et al.
Publicado: (2026)
por: Yurt, Ibrahim Ebrar, et al.
Publicado: (2026)
"I know myself better, but not really greatly": How Well Can LLMs Detect and Explain LLM-Generated Texts?
por: Ji, Jiazhou, et al.
Publicado: (2025)
por: Ji, Jiazhou, et al.
Publicado: (2025)
Ejemplares similares
-
VerAs: Verify then Assess STEM Lab Reports
por: Atil, Berk, et al.
Publicado: (2024) -
Middle School Students' Application of Science Learning From Physical Versus Virtual Labs to New Contexts
por: Dana Gnesdilow, et al.
Publicado: (2025) -
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
por: Knipper, R. Alexander, et al.
Publicado: (2025) -
Joint Training for Selective Prediction
por: Li, Zhaohui, et al.
Publicado: (2024) -
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025)