Salvato in:
| Autori principali: | Renda, Alana, Ross, Jillian, Cafarella, Michael, Andreas, Jacob |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2510.15096 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
CONCUR: A Framework for Continual Constrained and Unconstrained Routing
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
di: Kim, Junu, et al.
Pubblicazione: (2025)
di: Kim, Junu, et al.
Pubblicazione: (2025)
GAPO: Robust Advantage Estimation for Real-World Code LLMs
di: Zhang, Jianqing, et al.
Pubblicazione: (2025)
di: Zhang, Jianqing, et al.
Pubblicazione: (2025)
DAOpt: Modeling and Evaluation of Data-Driven Optimization under Uncertainty with LLMs
di: Zhu, WenZhuo, et al.
Pubblicazione: (2025)
di: Zhu, WenZhuo, et al.
Pubblicazione: (2025)
Toward In-Context Teaching: Adapting Examples to Students' Misconceptions
di: Ross, Alexis, et al.
Pubblicazione: (2024)
di: Ross, Alexis, et al.
Pubblicazione: (2024)
Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
di: Lu, Janna
Pubblicazione: (2025)
di: Lu, Janna
Pubblicazione: (2025)
Learning to Compile Programs to Neural Networks
di: Weber, Logan, et al.
Pubblicazione: (2024)
di: Weber, Logan, et al.
Pubblicazione: (2024)
Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models
di: Bassett, Bruce A., et al.
Pubblicazione: (2026)
di: Bassett, Bruce A., et al.
Pubblicazione: (2026)
Learning To Explore With Predictive World Model Via Self-Supervised Learning
di: Santana, Alana, et al.
Pubblicazione: (2025)
di: Santana, Alana, et al.
Pubblicazione: (2025)
Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots
di: Li, Chenhao, et al.
Pubblicazione: (2025)
di: Li, Chenhao, et al.
Pubblicazione: (2025)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
di: Damani, Mehul, et al.
Pubblicazione: (2025)
di: Damani, Mehul, et al.
Pubblicazione: (2025)
Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
di: Wallace, Eric, et al.
Pubblicazione: (2025)
di: Wallace, Eric, et al.
Pubblicazione: (2025)
When Counterfactual Reasoning Fails: Chaos and Real-World Complexity
di: Aalaila, Yahya, et al.
Pubblicazione: (2025)
di: Aalaila, Yahya, et al.
Pubblicazione: (2025)
A Production Scheduling Framework for Reinforcement Learning Under Real-World Constraints
di: Hoss, Jonathan, et al.
Pubblicazione: (2025)
di: Hoss, Jonathan, et al.
Pubblicazione: (2025)
How Uncertainty Estimation Scales with Sampling in Reasoning Models
di: Del, Maksym, et al.
Pubblicazione: (2026)
di: Del, Maksym, et al.
Pubblicazione: (2026)
Operationalizing Longitudinal Causal Discovery Under Real-World Workflow Constraints
di: Okuda, Tadahisa, et al.
Pubblicazione: (2026)
di: Okuda, Tadahisa, et al.
Pubblicazione: (2026)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
di: Hou, Yufang, et al.
Pubblicazione: (2024)
di: Hou, Yufang, et al.
Pubblicazione: (2024)
The Path to Open Innovation: Peer-Review Under Fire (PRUF)
di: Billions, Ava, et al.
Pubblicazione: (2025)
di: Billions, Ava, et al.
Pubblicazione: (2025)
Mars: Situated Inductive Reasoning in an Open-World Environment
di: Tang, Xiaojuan, et al.
Pubblicazione: (2024)
di: Tang, Xiaojuan, et al.
Pubblicazione: (2024)
Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
di: Ielanskyi, Mykyta, et al.
Pubblicazione: (2025)
di: Ielanskyi, Mykyta, et al.
Pubblicazione: (2025)
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
di: Camposampiero, Giacomo, et al.
Pubblicazione: (2025)
di: Camposampiero, Giacomo, et al.
Pubblicazione: (2025)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
Holistic Uncertainty Estimation For Open-Set Recognition
di: Erlygin, Leonid, et al.
Pubblicazione: (2024)
di: Erlygin, Leonid, et al.
Pubblicazione: (2024)
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation
di: Levy, Jacob, et al.
Pubblicazione: (2026)
di: Levy, Jacob, et al.
Pubblicazione: (2026)
REPEAT: Improving Uncertainty Estimation in Representation Learning Explainability
di: Wickstrøm, Kristoffer K., et al.
Pubblicazione: (2024)
di: Wickstrøm, Kristoffer K., et al.
Pubblicazione: (2024)
Towards Explainable, Safe Autonomous Driving with Language Embeddings for Novelty Identification and Active Learning: Framework and Experimental Analysis with Real-World Data Sets
di: Greer, Ross, et al.
Pubblicazione: (2024)
di: Greer, Ross, et al.
Pubblicazione: (2024)
Between the Layers Lies the Truth: Uncertainty Estimation in LLMs Using Intra-Layer Local Information Scores
di: Badash, Zvi N., et al.
Pubblicazione: (2026)
di: Badash, Zvi N., et al.
Pubblicazione: (2026)
Preservation of Feature Stability in Machine Learning Under Data Uncertainty for Decision Support in Critical Domains
di: Capała, Karol, et al.
Pubblicazione: (2024)
di: Capała, Karol, et al.
Pubblicazione: (2024)
Can LLMs Leverage Observational Data? Towards Data-Driven Causal Discovery with LLMs
di: Susanti, Yuni, et al.
Pubblicazione: (2025)
di: Susanti, Yuni, et al.
Pubblicazione: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
Beyond the Answer: Decoding the Behavior of LLMs as Scientific Reasoners
di: Pandey, Rohan, et al.
Pubblicazione: (2026)
di: Pandey, Rohan, et al.
Pubblicazione: (2026)
High-Fidelity Longitudinal Patient Simulation Using Real-World Data
di: Akagi, Yu, et al.
Pubblicazione: (2026)
di: Akagi, Yu, et al.
Pubblicazione: (2026)
Boosting Automatic Exercise Evaluation Through Musculoskeletal Simulation-Based IMU Data Augmentation
di: Spilz, Andreas, et al.
Pubblicazione: (2025)
di: Spilz, Andreas, et al.
Pubblicazione: (2025)
Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
di: Wang, Haohui, et al.
Pubblicazione: (2025)
di: Wang, Haohui, et al.
Pubblicazione: (2025)
Are LLMs Ready for Real-World Materials Discovery?
di: Miret, Santiago, et al.
Pubblicazione: (2024)
di: Miret, Santiago, et al.
Pubblicazione: (2024)
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
di: He, Yu, et al.
Pubblicazione: (2025)
di: He, Yu, et al.
Pubblicazione: (2025)
HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMs
di: Zeng, Xinyue, et al.
Pubblicazione: (2026)
di: Zeng, Xinyue, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
di: Chen, Peter Baile, et al.
Pubblicazione: (2025) -
CONCUR: A Framework for Continual Constrained and Unconstrained Routing
di: Chen, Peter Baile, et al.
Pubblicazione: (2025) -
Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
di: Kim, Junu, et al.
Pubblicazione: (2025) -
GAPO: Robust Advantage Estimation for Real-World Code LLMs
di: Zhang, Jianqing, et al.
Pubblicazione: (2025) -
DAOpt: Modeling and Evaluation of Data-Driven Optimization under Uncertainty with LLMs
di: Zhu, WenZhuo, et al.
Pubblicazione: (2025)