Do LLMs Know When to Flip a Coin? Strategic Randomization through Reasoning and Experience
Fuente:
arXiv
Salvato in:
| Autore principale: | Yang, Lingyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Random is Random? Evaluating the Randomness and Humaness of LLMs' Coin Flips
di: Van Koevering, Katherine, et al.
Pubblicazione: (2024)
di: Van Koevering, Katherine, et al.
Pubblicazione: (2024)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
di: Zhang, Xing, et al.
Pubblicazione: (2026)
di: Zhang, Xing, et al.
Pubblicazione: (2026)
Enough Coin Flips Can Make LLMs Act Bayesian
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
di: Gupta, Ritwik, et al.
Pubblicazione: (2025)
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
di: Dong, Zihan, et al.
Pubblicazione: (2026)
di: Dong, Zihan, et al.
Pubblicazione: (2026)
When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
di: Hao, Chenjie, et al.
Pubblicazione: (2026)
di: Hao, Chenjie, et al.
Pubblicazione: (2026)
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
di: Herrmann, Nils A., et al.
Pubblicazione: (2026)
di: Herrmann, Nils A., et al.
Pubblicazione: (2026)
Do Retrieval Augmented Language Models Know When They Don't Know?
di: Zhou, Youchao, et al.
Pubblicazione: (2025)
di: Zhou, Youchao, et al.
Pubblicazione: (2025)
Do Persona-Infused LLMs Affect Performance in a Strategic Reasoning Game?
di: Licato, John, et al.
Pubblicazione: (2025)
di: Licato, John, et al.
Pubblicazione: (2025)
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
di: Yao, Jianzhu, et al.
Pubblicazione: (2025)
di: Yao, Jianzhu, et al.
Pubblicazione: (2025)
Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning
di: Stoisser, Josefa Lia, et al.
Pubblicazione: (2025)
di: Stoisser, Josefa Lia, et al.
Pubblicazione: (2025)
Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know'
di: Manchingal, Shireen Kudukkil, et al.
Pubblicazione: (2025)
di: Manchingal, Shireen Kudukkil, et al.
Pubblicazione: (2025)
Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
di: McGovern, Hope, et al.
Pubblicazione: (2026)
di: McGovern, Hope, et al.
Pubblicazione: (2026)
Epistemic Deep Learning: Enabling Machine Learning Models to Know When They Do Not Know
di: Manchingal, Shireen Kudukkil
Pubblicazione: (2025)
di: Manchingal, Shireen Kudukkil
Pubblicazione: (2025)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
di: Schwinn, Leo, et al.
Pubblicazione: (2026)
di: Schwinn, Leo, et al.
Pubblicazione: (2026)
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
di: He, Yidong, et al.
Pubblicazione: (2026)
di: He, Yidong, et al.
Pubblicazione: (2026)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
di: Abdaljalil, Samir, et al.
Pubblicazione: (2026)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2026)
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
FlipAttack: Jailbreak LLMs via Flipping
di: Liu, Yue, et al.
Pubblicazione: (2024)
di: Liu, Yue, et al.
Pubblicazione: (2024)
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
di: Machcha, Sravanthi, et al.
Pubblicazione: (2026)
di: Machcha, Sravanthi, et al.
Pubblicazione: (2026)
Do Language Models Know When They're Hallucinating References?
di: Agrawal, Ayush, et al.
Pubblicazione: (2023)
di: Agrawal, Ayush, et al.
Pubblicazione: (2023)
Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs
di: Unlu, Eren
Pubblicazione: (2026)
di: Unlu, Eren
Pubblicazione: (2026)
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
di: Trinh, Tu, et al.
Pubblicazione: (2026)
di: Trinh, Tu, et al.
Pubblicazione: (2026)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
di: Wang, Kai, et al.
Pubblicazione: (2025)
di: Wang, Kai, et al.
Pubblicazione: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
di: Zhang, Yue, et al.
Pubblicazione: (2026)
di: Zhang, Yue, et al.
Pubblicazione: (2026)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
di: Bui, Anh Thu Maria, et al.
Pubblicazione: (2024)
di: Bui, Anh Thu Maria, et al.
Pubblicazione: (2024)
Base Models Know How to Reason, Thinking Models Learn When
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
di: Yang, Diji, et al.
Pubblicazione: (2025)
di: Yang, Diji, et al.
Pubblicazione: (2025)
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning
di: Ding, Chenlu, et al.
Pubblicazione: (2026)
di: Ding, Chenlu, et al.
Pubblicazione: (2026)
Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning
di: Zhao, Qiannian, et al.
Pubblicazione: (2026)
di: Zhao, Qiannian, et al.
Pubblicazione: (2026)
When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
di: Xu, Yifan, et al.
Pubblicazione: (2025)
di: Xu, Yifan, et al.
Pubblicazione: (2025)
Strategic Chain-of-Thought: Guiding Accurate Reasoning in LLMs through Strategy Elicitation
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
When Truthful Representations Flip Under Deceptive Instructions?
di: Long, Xianxuan, et al.
Pubblicazione: (2025)
di: Long, Xianxuan, et al.
Pubblicazione: (2025)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
di: Queipo-de-Llano, Enrique, et al.
Pubblicazione: (2025)
di: Queipo-de-Llano, Enrique, et al.
Pubblicazione: (2025)
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
Capability Self-Assessment: Teaching LLMs to Know Their Limits
di: Yang, Haoyan, et al.
Pubblicazione: (2026)
di: Yang, Haoyan, et al.
Pubblicazione: (2026)
FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning
di: Khalil, Khurram, et al.
Pubblicazione: (2025)
di: Khalil, Khurram, et al.
Pubblicazione: (2025)
Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations
di: Liu, Yilong, et al.
Pubblicazione: (2026)
di: Liu, Yilong, et al.
Pubblicazione: (2026)
When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
di: He, Zhiyuan, et al.
Pubblicazione: (2025)
di: He, Zhiyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
How Random is Random? Evaluating the Randomness and Humaness of LLMs' Coin Flips
di: Van Koevering, Katherine, et al.
Pubblicazione: (2024) -
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
di: Mei, Zhiting, et al.
Pubblicazione: (2025) -
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
di: Zhang, Xing, et al.
Pubblicazione: (2026) -
Enough Coin Flips Can Make LLMs Act Bayesian
di: Gupta, Ritwik, et al.
Pubblicazione: (2025) -
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
di: Dong, Zihan, et al.
Pubblicazione: (2026)