Going All-In on LLM Accuracy: Fake Prediction Markets, Real Confidence Signals
Fuente:
arXiv
Guardado en:
| Autor principal: | Todasco, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
LLMs for Game Theory: Entropy-Guided In-Context Learning and Adaptive CoT Reasoning
por: Banfi, Tommaso Felice, et al.
Publicado: (2026)
por: Banfi, Tommaso Felice, et al.
Publicado: (2026)
Node-Level Uncertainty Estimation in LLM-Generated SQL
por: Hasson, Hilaf, et al.
Publicado: (2025)
por: Hasson, Hilaf, et al.
Publicado: (2025)
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
por: Eisenstadt, Roy, et al.
Publicado: (2025)
por: Eisenstadt, Roy, et al.
Publicado: (2025)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
por: Yuan, Xin, et al.
Publicado: (2025)
por: Yuan, Xin, et al.
Publicado: (2025)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
por: Panuganti, Rajkiran
Publicado: (2026)
por: Panuganti, Rajkiran
Publicado: (2026)
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
por: Kadu, Ankush, et al.
Publicado: (2025)
por: Kadu, Ankush, et al.
Publicado: (2025)
Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
por: Xu, Shuyao, et al.
Publicado: (2025)
por: Xu, Shuyao, et al.
Publicado: (2025)
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
por: Zhao, Zhilong, et al.
Publicado: (2025)
por: Zhao, Zhilong, et al.
Publicado: (2025)
Social Cooperation in Conversational AI Agents
por: Çelikok, Mustafa Mert, et al.
Publicado: (2025)
por: Çelikok, Mustafa Mert, et al.
Publicado: (2025)
How Does Unfaithful Reasoning Emerge from Autoregressive Training? A Study of Synthetic Experiments
por: Wang, Fuxin, et al.
Publicado: (2026)
por: Wang, Fuxin, et al.
Publicado: (2026)
Extreme AutoML: Analysis of Classification, Regression, and NLP Performance
por: Ratner, Edward, et al.
Publicado: (2024)
por: Ratner, Edward, et al.
Publicado: (2024)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
por: Ding, Ruiyi, et al.
Publicado: (2026)
por: Ding, Ruiyi, et al.
Publicado: (2026)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
por: Deshmukh, Pratik, et al.
Publicado: (2026)
por: Deshmukh, Pratik, et al.
Publicado: (2026)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
OFMU: Optimization-Driven Framework for Machine Unlearning
por: Asif, Sadia, et al.
Publicado: (2025)
por: Asif, Sadia, et al.
Publicado: (2025)
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
por: Ustaomeroglu, Muhammed, et al.
Publicado: (2026)
por: Ustaomeroglu, Muhammed, et al.
Publicado: (2026)
Persona Features Control Emergent Misalignment
por: Wang, Miles, et al.
Publicado: (2025)
por: Wang, Miles, et al.
Publicado: (2025)
Autonomous Deep Agent
por: Yu, Amy, et al.
Publicado: (2025)
por: Yu, Amy, et al.
Publicado: (2025)
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
por: Adapala, Sai Teja Reddy
Publicado: (2025)
por: Adapala, Sai Teja Reddy
Publicado: (2025)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
por: Zhang, Yizhou, et al.
Publicado: (2025)
por: Zhang, Yizhou, et al.
Publicado: (2025)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
por: Avinash, Mynampati Sri Ranganadha
Publicado: (2026)
por: Avinash, Mynampati Sri Ranganadha
Publicado: (2026)
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
por: Zhao, Xingyu, et al.
Publicado: (2026)
por: Zhao, Xingyu, et al.
Publicado: (2026)
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
por: Heyman, Alex, et al.
Publicado: (2025)
por: Heyman, Alex, et al.
Publicado: (2025)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
por: Anderson, Samuel Cyrenius
Publicado: (2026)
por: Anderson, Samuel Cyrenius
Publicado: (2026)
Forget Attention: Importance-Aware Attention Is All You Need
por: Shin, Soohyeong, et al.
Publicado: (2026)
por: Shin, Soohyeong, et al.
Publicado: (2026)
Weakly Supervised Distillation of Hallucination Signals into Transformer Representations
por: Salehmohamed, Shoaib Sadiq, et al.
Publicado: (2026)
por: Salehmohamed, Shoaib Sadiq, et al.
Publicado: (2026)
Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo
por: Gadzhiev, Artem, et al.
Publicado: (2026)
por: Gadzhiev, Artem, et al.
Publicado: (2026)
AMEL: Accumulated Message Effects on LLM Judgments
por: Temkit, Sid-Ali
Publicado: (2026)
por: Temkit, Sid-Ali
Publicado: (2026)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
por: Sun, Kangkang, et al.
Publicado: (2026)
por: Sun, Kangkang, et al.
Publicado: (2026)
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
por: Pan, Muyu, et al.
Publicado: (2026)
por: Pan, Muyu, et al.
Publicado: (2026)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
por: Abramov, Roman, et al.
Publicado: (2025)
por: Abramov, Roman, et al.
Publicado: (2025)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2025)
por: Cho, Seonglae, et al.
Publicado: (2025)
The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
por: Larsen, Erik
Publicado: (2025)
por: Larsen, Erik
Publicado: (2025)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
por: Fang, Qitong, et al.
Publicado: (2026)
por: Fang, Qitong, et al.
Publicado: (2026)
On the Limits of Learned Importance Scoring for KV Cache Compression
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
por: Steele, Brady, et al.
Publicado: (2026)
por: Steele, Brady, et al.
Publicado: (2026)
Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
por: Xu, Jiexi
Publicado: (2025)
por: Xu, Jiexi
Publicado: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
por: Schesch, Benedikt, et al.
Publicado: (2026)
por: Schesch, Benedikt, et al.
Publicado: (2026)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
Ejemplares similares
-
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025) -
LLMs for Game Theory: Entropy-Guided In-Context Learning and Adaptive CoT Reasoning
por: Banfi, Tommaso Felice, et al.
Publicado: (2026) -
Node-Level Uncertainty Estimation in LLM-Generated SQL
por: Hasson, Hilaf, et al.
Publicado: (2025) -
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
por: Eisenstadt, Roy, et al.
Publicado: (2025) -
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
por: Yuan, Xin, et al.
Publicado: (2025)