Foundations of Global Consistency Checking with Noisy LLM Oracles
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Paul, Kirschbaum, Elke, Kasiviswanathan, Shiva |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples?
por: Hiremath, Sujai, et al.
Publicado: (2025)
por: Hiremath, Sujai, et al.
Publicado: (2025)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
por: Jia, Sheng, et al.
Publicado: (2025)
por: Jia, Sheng, et al.
Publicado: (2025)
A Quantitative Characterization of Forgetting in Post-Training
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026)
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
por: Park, Junhyung, et al.
Publicado: (2024)
por: Park, Junhyung, et al.
Publicado: (2024)
A Classical View on Benign Overfitting: The Role of Sample Size
por: Park, Junhyung, et al.
Publicado: (2025)
por: Park, Junhyung, et al.
Publicado: (2025)
Debiasing Reward Models by Representation Learning with Guarantees
por: Ng, Ignavier, et al.
Publicado: (2025)
por: Ng, Ignavier, et al.
Publicado: (2025)
What Causes Postoperative Aspiration?
por: Nagesh, Supriya, et al.
Publicado: (2025)
por: Nagesh, Supriya, et al.
Publicado: (2025)
The PetShop Dataset -- Finding Causes of Performance Issues across Microservices
por: Hardt, Michaela, et al.
Publicado: (2023)
por: Hardt, Michaela, et al.
Publicado: (2023)
QA-Calibration of Language Model Confidence Scores
por: Manggala, Putra, et al.
Publicado: (2024)
por: Manggala, Putra, et al.
Publicado: (2024)
Learning to Answer from Correct Demonstrations
por: Joshi, Nirmit, et al.
Publicado: (2025)
por: Joshi, Nirmit, et al.
Publicado: (2025)
Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
por: Kim, Jung-hun, et al.
Publicado: (2017)
por: Kim, Jung-hun, et al.
Publicado: (2017)
From Oracle to Noisy Context: Mitigating Contextual Exposure Bias in Speech-LLMs
por: Guo, Xiaoyong, et al.
Publicado: (2026)
por: Guo, Xiaoyong, et al.
Publicado: (2026)
Adjudicator: Correcting Noisy Labels with a KG-Informed Council of LLM Agents
por: You, Doohee, et al.
Publicado: (2025)
por: You, Doohee, et al.
Publicado: (2025)
OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control
por: Jacob, Darryl, et al.
Publicado: (2026)
por: Jacob, Darryl, et al.
Publicado: (2026)
Score matching through the roof: linear, nonlinear, and latent variables causal discovery
por: Montagna, Francesco, et al.
Publicado: (2024)
por: Montagna, Francesco, et al.
Publicado: (2024)
Consistency Checks for Language Model Forecasters
por: Paleka, Daniel, et al.
Publicado: (2024)
por: Paleka, Daniel, et al.
Publicado: (2024)
CoLafier: Collaborative Noisy Label Purifier With Local Intrinsic Dimensionality Guidance
por: Zhang, Dongyu, et al.
Publicado: (2024)
por: Zhang, Dongyu, et al.
Publicado: (2024)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
por: Gupta, Raavi, et al.
Publicado: (2025)
por: Gupta, Raavi, et al.
Publicado: (2025)
Global Policy-Space Response Oracles for Two-Player Zero-Sum Games
por: Zhang, Junyu, et al.
Publicado: (2026)
por: Zhang, Junyu, et al.
Publicado: (2026)
Understanding LLM-Driven Test Oracle Generation
por: Bodicoat, Adam, et al.
Publicado: (2026)
por: Bodicoat, Adam, et al.
Publicado: (2026)
Evolution without an Oracle: Driving Effective Evolution with LLM Judges
por: Zhao, Zhe, et al.
Publicado: (2025)
por: Zhao, Zhe, et al.
Publicado: (2025)
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
por: Wang, Haoran, et al.
Publicado: (2026)
por: Wang, Haoran, et al.
Publicado: (2026)
Enhancing Health Fact-Checking with LLM-Generated Synthetic Data
por: Zhang, Jingze, et al.
Publicado: (2025)
por: Zhang, Jingze, et al.
Publicado: (2025)
CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy Labels
por: Hu, Ruofan, et al.
Publicado: (2025)
por: Hu, Ruofan, et al.
Publicado: (2025)
MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing
por: Bhargava, Nishant, et al.
Publicado: (2026)
por: Bhargava, Nishant, et al.
Publicado: (2026)
BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking
por: Park, Hyunkyung, et al.
Publicado: (2026)
por: Park, Hyunkyung, et al.
Publicado: (2026)
AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment
por: Aghaebrahimian, Ahmad
Publicado: (2025)
por: Aghaebrahimian, Ahmad
Publicado: (2025)
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
por: Tsimpourlas, Foivos, et al.
Publicado: (2024)
por: Tsimpourlas, Foivos, et al.
Publicado: (2024)
Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding
por: Kim, Suyoung, et al.
Publicado: (2024)
por: Kim, Suyoung, et al.
Publicado: (2024)
V15 - Recursive Symbolic Intelligence - φ-Ache Recursive Collapsors and the Collapse Oracle Engine
por: Foster, Camaron
Publicado: (2025)
por: Foster, Camaron
Publicado: (2025)
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks
por: Feng, Yu, et al.
Publicado: (2025)
por: Feng, Yu, et al.
Publicado: (2025)
GLaMoR: Consistency Checking of OWL Ontologies using Graph Language Models
por: Mücke, Justin, et al.
Publicado: (2025)
por: Mücke, Justin, et al.
Publicado: (2025)
Beyond Retrieval: Improving Evidence Quality for LLM-based Multimodal Fact-Checking
por: Ou, Haoran, et al.
Publicado: (2025)
por: Ou, Haoran, et al.
Publicado: (2025)
HYPERHEURIST: A Simulated Annealing-Based Control Framework for LLM-Driven Code Generation in Optimized Hardware Design
por: Ahir, Shiva, et al.
Publicado: (2026)
por: Ahir, Shiva, et al.
Publicado: (2026)
OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking
por: Ma, Yiding, et al.
Publicado: (2026)
por: Ma, Yiding, et al.
Publicado: (2026)
LoRA as Oracle
por: Arazzi, Marco, et al.
Publicado: (2026)
por: Arazzi, Marco, et al.
Publicado: (2026)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
por: Vazhentsev, Artem, et al.
Publicado: (2026)
por: Vazhentsev, Artem, et al.
Publicado: (2026)
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities
por: Oh, Changdae, et al.
Publicado: (2026)
por: Oh, Changdae, et al.
Publicado: (2026)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
por: Lin, Shuyi, et al.
Publicado: (2025)
por: Lin, Shuyi, et al.
Publicado: (2025)
Harnessing Consistency for Robust Test-Time LLM Ensemble
por: Zeng, Zhichen, et al.
Publicado: (2025)
por: Zeng, Zhichen, et al.
Publicado: (2025)
Ejemplares similares
-
From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples?
por: Hiremath, Sujai, et al.
Publicado: (2025) -
Training Large Language Models To Reason In Parallel With Global Forking Tokens
por: Jia, Sheng, et al.
Publicado: (2025) -
A Quantitative Characterization of Forgetting in Post-Training
por: Balasubramanian, Krishnakumar, et al.
Publicado: (2026) -
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
por: Park, Junhyung, et al.
Publicado: (2024) -
A Classical View on Benign Overfitting: The Role of Sample Size
por: Park, Junhyung, et al.
Publicado: (2025)