Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Zhonghao, Qiu, Tianyi, Shirado, Hirokazu, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Improvement as Coherence Optimization: A Theoretical Account
von: Qiu, Tianyi, et al.
Veröffentlicht: (2026)
von: Qiu, Tianyi, et al.
Veröffentlicht: (2026)
Spontaneous Giving and Calculated Greed in Language Models
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
The Lock-in Hypothesis: Stagnation by Algorithm
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2025)
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2025)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
Temporal Consistency for LLM Reasoning Process Error Identification
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
von: Hallaç, İbrahim Rıza, et al.
Veröffentlicht: (2026)
von: Hallaç, İbrahim Rıza, et al.
Veröffentlicht: (2026)
PredictaBoard: Benchmarking LLM Score Predictability
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2025)
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2025)
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
von: Jesson, Andrew, et al.
Veröffentlicht: (2024)
von: Jesson, Andrew, et al.
Veröffentlicht: (2024)
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
von: Gisler, Isaia, et al.
Veröffentlicht: (2026)
von: Gisler, Isaia, et al.
Veröffentlicht: (2026)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
Fluid Language Model Benchmarking
von: Hofmann, Valentin, et al.
Veröffentlicht: (2025)
von: Hofmann, Valentin, et al.
Veröffentlicht: (2025)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
On Designing Effective RL Reward at Training Time for LLM Reasoning
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
von: Guo, Yiju, et al.
Veröffentlicht: (2026)
von: Guo, Yiju, et al.
Veröffentlicht: (2026)
What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
BALAR : A Bayesian Agentic Loop for Active Reasoning
von: Echarghaoui, Aymen, et al.
Veröffentlicht: (2026)
von: Echarghaoui, Aymen, et al.
Veröffentlicht: (2026)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
von: Park, Jin Hyun, et al.
Veröffentlicht: (2025)
von: Park, Jin Hyun, et al.
Veröffentlicht: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
von: Shi, Jiajun, et al.
Veröffentlicht: (2025)
von: Shi, Jiajun, et al.
Veröffentlicht: (2025)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
von: Landesberg, Eddie
Veröffentlicht: (2026)
von: Landesberg, Eddie
Veröffentlicht: (2026)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
Explainable LLM Unlearning Through Reasoning
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
Token-Budget-Aware LLM Reasoning
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
Deep Language Geometry: Constructing a Metric Space from LLM Weights
von: Shamrai, Maksym, et al.
Veröffentlicht: (2025)
von: Shamrai, Maksym, et al.
Veröffentlicht: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Improvement as Coherence Optimization: A Theoretical Account
von: Qiu, Tianyi, et al.
Veröffentlicht: (2026) -
Spontaneous Giving and Calculated Greed in Language Models
von: Li, Yuxuan, et al.
Veröffentlicht: (2025) -
Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs
von: Li, Yuxuan, et al.
Veröffentlicht: (2025) -
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025) -
The Lock-in Hypothesis: Stagnation by Algorithm
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2025)