KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
Fuente:
arXiv
Salvato in:
| Autori principali: | Shu, Matthew, Balepur, Nishant, Feng, Shi, Boyd-Graber, Jordan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
di: Balepur, Nishant, et al.
Pubblicazione: (2023)
di: Balepur, Nishant, et al.
Pubblicazione: (2023)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
di: Srikanth, Neha, et al.
Pubblicazione: (2026)
di: Srikanth, Neha, et al.
Pubblicazione: (2026)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
Labeled Interactive Topic Models
di: Seelman, Kyle, et al.
Pubblicazione: (2023)
di: Seelman, Kyle, et al.
Pubblicazione: (2023)
CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
di: Hoyle, Alexander, et al.
Pubblicazione: (2025)
di: Hoyle, Alexander, et al.
Pubblicazione: (2025)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
di: Gao, Cheng, et al.
Pubblicazione: (2026)
di: Gao, Cheng, et al.
Pubblicazione: (2026)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
di: Li, Zongxia, et al.
Pubblicazione: (2024)
di: Li, Zongxia, et al.
Pubblicazione: (2024)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
di: Srikanth, Neha, et al.
Pubblicazione: (2023)
di: Srikanth, Neha, et al.
Pubblicazione: (2023)
Personalized Help for Optimizing Low-Skilled Users' Strategy
di: Gu, Feng, et al.
Pubblicazione: (2024)
di: Gu, Feng, et al.
Pubblicazione: (2024)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
di: Si, Chenglei, et al.
Pubblicazione: (2023)
di: Si, Chenglei, et al.
Pubblicazione: (2023)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
di: Li, Zongxia, et al.
Pubblicazione: (2024)
di: Li, Zongxia, et al.
Pubblicazione: (2024)
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
di: Mondal, Ishani, et al.
Pubblicazione: (2025)
di: Mondal, Ishani, et al.
Pubblicazione: (2025)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
di: Gu, Feng, et al.
Pubblicazione: (2025)
di: Gu, Feng, et al.
Pubblicazione: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
di: Palta, Shramay, et al.
Pubblicazione: (2024)
di: Palta, Shramay, et al.
Pubblicazione: (2024)
Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2025)
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2025)
AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA
di: Kabir, Tasnim, et al.
Pubblicazione: (2026)
di: Kabir, Tasnim, et al.
Pubblicazione: (2026)
Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
Active Continual Learning: On Balancing Knowledge Retention and Learnability
di: Vu, Thuy-Trang, et al.
Pubblicazione: (2023)
di: Vu, Thuy-Trang, et al.
Pubblicazione: (2023)
More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2024)
di: Wongkamjan, Wichayaporn, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
di: Balepur, Nishant, et al.
Pubblicazione: (2024) -
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025) -
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
di: Balepur, Nishant, et al.
Pubblicazione: (2024) -
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
di: Balepur, Nishant, et al.
Pubblicazione: (2025) -
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
di: Balepur, Nishant, et al.
Pubblicazione: (2025)