Saved in:
| Main Authors: | Nguyen, Bang, Yu, Mengxia, Huang, Yun, Jiang, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.12242 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
by: Nguyen, Bang, et al.
Published: (2025)
by: Nguyen, Bang, et al.
Published: (2025)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
by: Deng, Yihe, et al.
Published: (2023)
by: Deng, Yihe, et al.
Published: (2023)
Context Selection and Rewriting for Video-based Educational Question Generation
by: Yu, Mengxia, et al.
Published: (2025)
by: Yu, Mengxia, et al.
Published: (2025)
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
by: Dadfar, Zachary Pedram
Published: (2026)
by: Dadfar, Zachary Pedram
Published: (2026)
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
by: Zelikman, Eric, et al.
Published: (2024)
by: Zelikman, Eric, et al.
Published: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
by: Wang, Qianli, et al.
Published: (2026)
by: Wang, Qianli, et al.
Published: (2026)
QOG:Question and Options Generation based on Language Model
by: Zhou, Jincheng
Published: (2024)
by: Zhou, Jincheng
Published: (2024)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
by: Jiang, Mingjian, et al.
Published: (2024)
by: Jiang, Mingjian, et al.
Published: (2024)
Synthetic Multimodal Question Generation
by: Wu, Ian, et al.
Published: (2024)
by: Wu, Ian, et al.
Published: (2024)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
by: Dua, Radhika, et al.
Published: (2025)
by: Dua, Radhika, et al.
Published: (2025)
Beyond Independent Passages: Adaptive Passage Combination Retrieval for Retrieval Augmented Open-Domain Question Answering
by: Ko, Ting-Wen, et al.
Published: (2025)
by: Ko, Ting-Wen, et al.
Published: (2025)
Towards Verifiable Text Generation with Symbolic References
by: Hennigen, Lucas Torroba, et al.
Published: (2023)
by: Hennigen, Lucas Torroba, et al.
Published: (2023)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
Rhetorical Questions in LLM Representations: A Linear Probing Study
by: Yao, Louie Hong, et al.
Published: (2026)
by: Yao, Louie Hong, et al.
Published: (2026)
Machine Unlearning in Generative AI: A Survey
by: Liu, Zheyuan, et al.
Published: (2024)
by: Liu, Zheyuan, et al.
Published: (2024)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
by: Arasteh, Soroosh Tayebi, et al.
Published: (2024)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2024)
Towards Enriched Controllability for Educational Question Generation
by: Leite, Bernardo, et al.
Published: (2023)
by: Leite, Bernardo, et al.
Published: (2023)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
by: Kulkarni, Atharva, et al.
Published: (2025)
by: Kulkarni, Atharva, et al.
Published: (2025)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
by: Yadav, Vikas, et al.
Published: (2024)
by: Yadav, Vikas, et al.
Published: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
by: Lee, Yooseop, et al.
Published: (2025)
by: Lee, Yooseop, et al.
Published: (2025)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
by: Haussmann, Aden
Published: (2025)
by: Haussmann, Aden
Published: (2025)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering
by: Ji, An-Yang, et al.
Published: (2026)
by: Ji, An-Yang, et al.
Published: (2026)
ReALM: Reference Resolution As Language Modeling
by: Moniz, Joel Ruben Antony, et al.
Published: (2024)
by: Moniz, Joel Ruben Antony, et al.
Published: (2024)
A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task
by: Fujisaki, Yuya, et al.
Published: (2024)
by: Fujisaki, Yuya, et al.
Published: (2024)
Large Language Models in Fire Engineering: An Examination of Technical Questions Against Domain Knowledge
by: Hostetter, Haley, et al.
Published: (2024)
by: Hostetter, Haley, et al.
Published: (2024)
Multi-hop Question Answering under Temporal Knowledge Editing
by: Cheng, Keyuan, et al.
Published: (2024)
by: Cheng, Keyuan, et al.
Published: (2024)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Deep Learning Approaches for Improving Question Answering Systems in Hepatocellular Carcinoma Research
by: Huo, Shuning, et al.
Published: (2024)
by: Huo, Shuning, et al.
Published: (2024)
MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters
by: Dada, Amin, et al.
Published: (2025)
by: Dada, Amin, et al.
Published: (2025)
Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
by: Thulke, David, et al.
Published: (2025)
by: Thulke, David, et al.
Published: (2025)
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
by: Singh, Janvijay, et al.
Published: (2025)
by: Singh, Janvijay, et al.
Published: (2025)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
by: Zhu, Hao, et al.
Published: (2025)
by: Zhu, Hao, et al.
Published: (2025)
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
by: Saha, Binita, et al.
Published: (2025)
by: Saha, Binita, et al.
Published: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Similar Items
-
QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
by: Nguyen, Bang, et al.
Published: (2025) -
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
by: Deng, Yihe, et al.
Published: (2023) -
Context Selection and Rewriting for Video-based Educational Question Generation
by: Yu, Mengxia, et al.
Published: (2025) -
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
by: Dadfar, Zachary Pedram
Published: (2026) -
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
by: Zelikman, Eric, et al.
Published: (2024)