Reasoning Models Better Express Their Confidence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yoon, Dongkeun, Kim, Seungone, Yang, Sohee, Kim, Sunkyoung, Kim, Soyeon, Kim, Yongil, Choi, Eunbi, Kim, Yireun, Seo, Minjoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
von: Kim, Sunkyoung, et al.
Veröffentlicht: (2023)
von: Kim, Sunkyoung, et al.
Veröffentlicht: (2023)
LangBridge: Multilingual Reasoning Without Multilingual Supervision
von: Yoon, Dongkeun, et al.
Veröffentlicht: (2024)
von: Yoon, Dongkeun, et al.
Veröffentlicht: (2024)
BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
Latent Reasoning via Sentence Embedding Prediction
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2025)
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2025)
EXAONE Deep: Reasoning Enhanced Language Models
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2023)
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2023)
Generative Prompt Internalization
von: Shin, Haebin, et al.
Veröffentlicht: (2024)
von: Shin, Haebin, et al.
Veröffentlicht: (2024)
Rethinking the Role of Proxy Rewards in Language Model Alignment
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
von: Hong, Seokhee, et al.
Veröffentlicht: (2025)
von: Hong, Seokhee, et al.
Veröffentlicht: (2025)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2024)
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2024)
Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models
von: Kim, Jiyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jiyeon, et al.
Veröffentlicht: (2026)
K-EXAONE Technical Report
von: Choi, Eunbi, et al.
Veröffentlicht: (2026)
von: Choi, Eunbi, et al.
Veröffentlicht: (2026)
The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models
von: Kim, Dueun, et al.
Veröffentlicht: (2026)
von: Kim, Dueun, et al.
Veröffentlicht: (2026)
EXAONE 3.0 7.8B Instruction Tuned Language Model
von: An, Soyoung, et al.
Veröffentlicht: (2024)
von: An, Soyoung, et al.
Veröffentlicht: (2024)
M-Prometheus: A Suite of Open Multilingual LLM Judges
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
von: Lee, Seongyun, et al.
Veröffentlicht: (2025)
von: Lee, Seongyun, et al.
Veröffentlicht: (2025)
How Well Do Large Language Models Truly Ground?
von: Lee, Hyunji, et al.
Veröffentlicht: (2023)
von: Lee, Hyunji, et al.
Veröffentlicht: (2023)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Better Instruction-Following Through Minimum Bayes Risk
von: Wu, Ian, et al.
Veröffentlicht: (2024)
von: Wu, Ian, et al.
Veröffentlicht: (2024)
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Active Merchant Non-player Characters
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
Latent Preference Modeling for Cross-Session Personalized Tool Calling
von: Yoon, Yejin, et al.
Veröffentlicht: (2026)
von: Yoon, Yejin, et al.
Veröffentlicht: (2026)
FREESON: Retriever-Free Retrieval-Augmented Reasoning via Corpus-Traversing MCTS
von: Kim, Chaeeun, et al.
Veröffentlicht: (2025)
von: Kim, Chaeeun, et al.
Veröffentlicht: (2025)
DeFrame: Debiasing Large Language Models Against Framing Effects
von: Lim, Kahee, et al.
Veröffentlicht: (2026)
von: Lim, Kahee, et al.
Veröffentlicht: (2026)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
von: Jang, Chaeyun, et al.
Veröffentlicht: (2025)
von: Jang, Chaeyun, et al.
Veröffentlicht: (2025)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
von: Lee, Hojae, et al.
Veröffentlicht: (2024)
von: Lee, Hojae, et al.
Veröffentlicht: (2024)
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
Aligning to Thousands of Preferences via System Message Generalization
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
DEBATE: Devil's Advocate-Based Assessment and Text Evaluation
von: Kim, Alex, et al.
Veröffentlicht: (2024)
von: Kim, Alex, et al.
Veröffentlicht: (2024)
SAAS: Solving Ability Amplification Strategy for Enhanced Mathematical Reasoning in Large Language Models
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
von: Kim, Jun Seo, et al.
Veröffentlicht: (2025)
von: Kim, Jun Seo, et al.
Veröffentlicht: (2025)
DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in Large Language Models
von: Kim, Olivia
Veröffentlicht: (2025)
von: Kim, Olivia
Veröffentlicht: (2025)
Confidence-guided Refinement Reasoning for Zero-shot Question Answering
von: Jang, Youwon, et al.
Veröffentlicht: (2025)
von: Jang, Youwon, et al.
Veröffentlicht: (2025)
CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
von: Kim, Jin Young, et al.
Veröffentlicht: (2025)
von: Kim, Jin Young, et al.
Veröffentlicht: (2025)
Self-Training Elicits Concise Reasoning in Large Language Models
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
von: Munkhbat, Tergel, et al.
Veröffentlicht: (2025)
Can Language Models Evaluate Human Written Text? Case Study on Korean Student Writing for Education
von: Kim, Seungyoon, et al.
Veröffentlicht: (2024)
von: Kim, Seungyoon, et al.
Veröffentlicht: (2024)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
von: Lee, Seongyun, et al.
Veröffentlicht: (2026)
von: Lee, Seongyun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
von: Kim, Sunkyoung, et al.
Veröffentlicht: (2023) -
LangBridge: Multilingual Reasoning Without Multilingual Supervision
von: Yoon, Dongkeun, et al.
Veröffentlicht: (2024) -
BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026) -
Latent Reasoning via Sentence Embedding Prediction
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2025) -
EXAONE Deep: Reasoning Enhanced Language Models
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)