Keep Guessing? When Considering Inference Scaling, Mind the Baselines
Fuente:
arXiv
Salvato in:
| Autori principali: | Yona, Gal, Honovich, Or, Levy, Omer, Aharoni, Roee |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
di: Yona, Gal, et al.
Pubblicazione: (2024)
di: Yona, Gal, et al.
Pubblicazione: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
di: Yona, Gal, et al.
Pubblicazione: (2024)
di: Yona, Gal, et al.
Pubblicazione: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
di: Shaham, Uri, et al.
Pubblicazione: (2024)
di: Shaham, Uri, et al.
Pubblicazione: (2024)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
di: Calderon, Nitay, et al.
Pubblicazione: (2026)
di: Calderon, Nitay, et al.
Pubblicazione: (2026)
Confidence Improves Self-Consistency in LLMs
di: Taubenfeld, Amir, et al.
Pubblicazione: (2025)
di: Taubenfeld, Amir, et al.
Pubblicazione: (2025)
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
di: Cattan, Arie, et al.
Pubblicazione: (2025)
di: Cattan, Arie, et al.
Pubblicazione: (2025)
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
di: Khairi, Ammar, et al.
Pubblicazione: (2025)
di: Khairi, Ammar, et al.
Pubblicazione: (2025)
Asterisk*: Keep it Simple
di: Semenov, Andrew
Pubblicazione: (2024)
di: Semenov, Andrew
Pubblicazione: (2024)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
di: Yona, Itay, et al.
Pubblicazione: (2026)
di: Yona, Itay, et al.
Pubblicazione: (2026)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2025)
di: Wu, Cheng-Kuang, et al.
Pubblicazione: (2025)
GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
di: Hutson, Dylan, et al.
Pubblicazione: (2025)
di: Hutson, Dylan, et al.
Pubblicazione: (2025)
LiveMind: Low-latency Large Language Models with Simultaneous Inference
di: Chen, Chuangtao, et al.
Pubblicazione: (2024)
di: Chen, Chuangtao, et al.
Pubblicazione: (2024)
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
When Benchmarks Leak: Inference-Time Decontamination for LLMs
di: Chai, Jianzhe, et al.
Pubblicazione: (2026)
di: Chai, Jianzhe, et al.
Pubblicazione: (2026)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
di: Wu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Wu, Xiaoyuan, et al.
Pubblicazione: (2025)
Keep It Private: Unsupervised Privatization of Online Text
di: Bao, Calvin, et al.
Pubblicazione: (2024)
di: Bao, Calvin, et al.
Pubblicazione: (2024)
Extrapolation Merging: Keep Improving With Extrapolation and Merging
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
di: Lin, Yiguan, et al.
Pubblicazione: (2025)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
di: Cheng, Fenghua, et al.
Pubblicazione: (2025)
di: Cheng, Fenghua, et al.
Pubblicazione: (2025)
Mind2: Mind-to-Mind Emotional Support System with Bidirectional Cognitive Discourse Analysis
di: Hong, Shi Yin, et al.
Pubblicazione: (2025)
di: Hong, Shi Yin, et al.
Pubblicazione: (2025)
Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
di: Nguyen, Luong N.
Pubblicazione: (2026)
di: Nguyen, Luong N.
Pubblicazione: (2026)
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
di: Liang, Fangzhou, et al.
Pubblicazione: (2025)
di: Liang, Fangzhou, et al.
Pubblicazione: (2025)
MindSET: Advancing Mental Health Benchmarking through Large-Scale Social Media Data
di: Mankarious, Saad, et al.
Pubblicazione: (2025)
di: Mankarious, Saad, et al.
Pubblicazione: (2025)
Inverse Scaling: When Bigger Isn't Better
di: McKenzie, Ian R., et al.
Pubblicazione: (2023)
di: McKenzie, Ian R., et al.
Pubblicazione: (2023)
CRoCoDiL: Continuous and Robust Conditioned Diffusion for Language
di: Uziel, Roy, et al.
Pubblicazione: (2026)
di: Uziel, Roy, et al.
Pubblicazione: (2026)
`Keep it Together': Enforcing Cohesion in Extractive Summaries by Simulating Human Memory
di: Cardenas, Ronald, et al.
Pubblicazione: (2024)
di: Cardenas, Ronald, et al.
Pubblicazione: (2024)
Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
di: Huang, Wenke, et al.
Pubblicazione: (2025)
di: Huang, Wenke, et al.
Pubblicazione: (2025)
Scaling LLM Inference with Optimized Sample Compute Allocation
di: Zhang, Kexun, et al.
Pubblicazione: (2024)
di: Zhang, Kexun, et al.
Pubblicazione: (2024)
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning
di: Wagner, Eitan, et al.
Pubblicazione: (2024)
di: Wagner, Eitan, et al.
Pubblicazione: (2024)
MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
di: Chen, Zehui, et al.
Pubblicazione: (2024)
di: Chen, Zehui, et al.
Pubblicazione: (2024)
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
di: Xu, Haoming, et al.
Pubblicazione: (2026)
di: Xu, Haoming, et al.
Pubblicazione: (2026)
UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
di: Qian, Cheng, et al.
Pubblicazione: (2026)
di: Qian, Cheng, et al.
Pubblicazione: (2026)
MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural Learning
di: Lică, Mircea, et al.
Pubblicazione: (2024)
di: Lică, Mircea, et al.
Pubblicazione: (2024)
UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling
di: Huang, Kaiyu, et al.
Pubblicazione: (2026)
di: Huang, Kaiyu, et al.
Pubblicazione: (2026)
Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration
di: Saad, Fardin, et al.
Pubblicazione: (2025)
di: Saad, Fardin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
di: Yona, Gal, et al.
Pubblicazione: (2024) -
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
di: Yona, Gal, et al.
Pubblicazione: (2024) -
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
di: Ifergan, Maxim, et al.
Pubblicazione: (2024) -
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
di: Shaham, Uri, et al.
Pubblicazione: (2024) -
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)