Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiazheng, Xu, Hainiu, Sun, Zhaoyue, Zhou, Yuxiang, West, David, Aloisi, Cesare, He, Yulan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Automated Explainable Educational Assessment System Built on LLMs
by: Li, Jiazheng, et al.
Published: (2024)
by: Li, Jiazheng, et al.
Published: (2024)
EnigmaToM: Improve LLMs' Theory-of-Mind Reasoning Capabilities with Neural Knowledge Base of Entity States
by: Xu, Hainiu, et al.
Published: (2025)
by: Xu, Hainiu, et al.
Published: (2025)
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment
by: Li, Jiazheng, et al.
Published: (2024)
by: Li, Jiazheng, et al.
Published: (2024)
ExDDI: Explaining Drug-Drug Interaction Predictions with Natural Language
by: Sun, Zhaoyue, et al.
Published: (2024)
by: Sun, Zhaoyue, et al.
Published: (2024)
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives
by: Zhao, Runcong, et al.
Published: (2024)
by: Zhao, Runcong, et al.
Published: (2024)
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
Modeling Subjectivity in Cognitive Appraisal with Language Models
by: Zhou, Yuxiang, et al.
Published: (2025)
by: Zhou, Yuxiang, et al.
Published: (2025)
Towards Unified Task Embeddings Across Multiple Models: Bridging the Gap for Prompt-Based Large Language Models and Beyond
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
by: Sun, Zhaoyue, et al.
Published: (2026)
by: Sun, Zhaoyue, et al.
Published: (2026)
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
by: Lu, Junru, et al.
Published: (2024)
by: Lu, Junru, et al.
Published: (2024)
When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
by: Xu, Hainiu, et al.
Published: (2024)
by: Xu, Hainiu, et al.
Published: (2024)
Leveraging ChatGPT in Pharmacovigilance Event Extraction: An Empirical Study
by: Sun, Zhaoyue, et al.
Published: (2024)
by: Sun, Zhaoyue, et al.
Published: (2024)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
by: Chu, SeongYeub, et al.
Published: (2024)
by: Chu, SeongYeub, et al.
Published: (2024)
The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis
by: Zhou, Yuxiang, et al.
Published: (2023)
by: Zhou, Yuxiang, et al.
Published: (2023)
EvolvTrip: Enhancing Literary Character Understanding with Temporal Theory-of-Mind Graphs
by: Yang, Bohao, et al.
Published: (2025)
by: Yang, Bohao, et al.
Published: (2025)
Cascading Large Language Models for Salient Event Graph Generation
by: Tan, Xingwei, et al.
Published: (2024)
by: Tan, Xingwei, et al.
Published: (2024)
Set-Aligning Framework for Auto-Regressive Event Temporal Graph Generation
by: Tan, Xingwei, et al.
Published: (2024)
by: Tan, Xingwei, et al.
Published: (2024)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Enhancing Event Causality Identification with Rationale and Structure-Aware Causal Question Answering
by: Zhang, Baiyan, et al.
Published: (2024)
by: Zhang, Baiyan, et al.
Published: (2024)
Rethinking Human Preference Evaluation of LLM Rationales
by: Li, Ziang, et al.
Published: (2025)
by: Li, Ziang, et al.
Published: (2025)
FIPO: Free-form Instruction-oriented Prompt Optimization with Preference Dataset and Modular Fine-tuning Schema
by: Lu, Junru, et al.
Published: (2024)
by: Lu, Junru, et al.
Published: (2024)
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering
by: Sohn, Jiwoong, et al.
Published: (2024)
by: Sohn, Jiwoong, et al.
Published: (2024)
Are NLP Models Good at Tracing Thoughts: An Overview of Narrative Understanding
by: Zhu, Lixing, et al.
Published: (2023)
by: Zhu, Lixing, et al.
Published: (2023)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
RGAR: Recurrence Generation-augmented Retrieval for Factual-aware Medical Question Answering
by: Liang, Sichu, et al.
Published: (2025)
by: Liang, Sichu, et al.
Published: (2025)
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
by: Mazzaccara, Davide, et al.
Published: (2024)
by: Mazzaccara, Davide, et al.
Published: (2024)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
From Prediction to Justification: Aligning Sentiment Reasoning with Human Rationale via Reinforcement Learning
by: Zhang, Shihao, et al.
Published: (2026)
by: Zhang, Shihao, et al.
Published: (2026)
Deductive Beam Search: Decoding Deducible Rationale for Chain-of-Thought Reasoning
by: Zhu, Tinghui, et al.
Published: (2024)
by: Zhu, Tinghui, et al.
Published: (2024)
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
by: Zhu, Chiwei, et al.
Published: (2025)
by: Zhu, Chiwei, et al.
Published: (2025)
RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following
by: Lu, Junru, et al.
Published: (2025)
by: Lu, Junru, et al.
Published: (2025)
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
Teach-to-Reason with Scoring: Self-Explainable Rationale-Driven Multi-Trait Essay Scoring
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
by: Karim, Ahmed, et al.
Published: (2025)
by: Karim, Ahmed, et al.
Published: (2025)
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization
by: Deng, Mengyi, et al.
Published: (2026)
by: Deng, Mengyi, et al.
Published: (2026)
Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Similar Items
-
An Automated Explainable Educational Assessment System Built on LLMs
by: Li, Jiazheng, et al.
Published: (2024) -
EnigmaToM: Improve LLMs' Theory-of-Mind Reasoning Capabilities with Neural Knowledge Base of Entity States
by: Xu, Hainiu, et al.
Published: (2025) -
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
by: Li, Jiazheng, et al.
Published: (2025) -
AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment
by: Li, Jiazheng, et al.
Published: (2024) -
ExDDI: Explaining Drug-Drug Interaction Predictions with Natural Language
by: Sun, Zhaoyue, et al.
Published: (2024)