Probabilistic Soundness Guarantees in LLM Reasoning Chains
Fuente:
arXiv
Salvato in:
| Autori principali: | You, Weiqiu, Xue, Anton, Havaldar, Shreya, Rao, Delip, Jin, Helen, Callison-Burch, Chris, Wong, Eric |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
di: Rao, Delip, et al.
Pubblicazione: (2025)
di: Rao, Delip, et al.
Pubblicazione: (2025)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
di: Rao, Delip, et al.
Pubblicazione: (2026)
di: Rao, Delip, et al.
Pubblicazione: (2026)
Autorubric: Unifying Rubric-based LLM Evaluation
di: Rao, Delip, et al.
Pubblicazione: (2026)
di: Rao, Delip, et al.
Pubblicazione: (2026)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
di: Rao, Delip, et al.
Pubblicazione: (2026)
di: Rao, Delip, et al.
Pubblicazione: (2026)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
di: Rao, Delip, et al.
Pubblicazione: (2026)
di: Rao, Delip, et al.
Pubblicazione: (2026)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
di: Rao, Delip, et al.
Pubblicazione: (2024)
di: Rao, Delip, et al.
Pubblicazione: (2024)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
di: Rao, Delip, et al.
Pubblicazione: (2026)
di: Rao, Delip, et al.
Pubblicazione: (2026)
Probabilistic Stability Guarantees for Feature Attributions
di: Jin, Helen, et al.
Pubblicazione: (2025)
di: Jin, Helen, et al.
Pubblicazione: (2025)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
di: Rao, Delip, et al.
Pubblicazione: (2026)
di: Rao, Delip, et al.
Pubblicazione: (2026)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
di: Patel, Ajay, et al.
Pubblicazione: (2024)
di: Patel, Ajay, et al.
Pubblicazione: (2024)
Adaptively profiling models with task elicitation
di: Brown, Davis, et al.
Pubblicazione: (2025)
di: Brown, Davis, et al.
Pubblicazione: (2025)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
di: Liu, Muxin, et al.
Pubblicazione: (2026)
di: Liu, Muxin, et al.
Pubblicazione: (2026)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
di: Patel, Ajay, et al.
Pubblicazione: (2026)
di: Patel, Ajay, et al.
Pubblicazione: (2026)
Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models
di: You, Weiqiu, et al.
Pubblicazione: (2024)
di: You, Weiqiu, et al.
Pubblicazione: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
di: Zhu, Andrew, et al.
Pubblicazione: (2025)
di: Zhu, Andrew, et al.
Pubblicazione: (2025)
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
di: Patel, Ajay, et al.
Pubblicazione: (2024)
di: Patel, Ajay, et al.
Pubblicazione: (2024)
The FIX Benchmark: Extracting Features Interpretable to eXperts
di: Jin, Helen, et al.
Pubblicazione: (2024)
di: Jin, Helen, et al.
Pubblicazione: (2024)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
di: Dugan, Liam, et al.
Pubblicazione: (2025)
di: Dugan, Liam, et al.
Pubblicazione: (2025)
Towards Style Alignment in Cross-Cultural Translation
di: Havaldar, Shreya, et al.
Pubblicazione: (2025)
di: Havaldar, Shreya, et al.
Pubblicazione: (2025)
Comparing Styles across Languages: A Cross-Cultural Exploration of Politeness
di: Havaldar, Shreya, et al.
Pubblicazione: (2023)
di: Havaldar, Shreya, et al.
Pubblicazione: (2023)
Large Language Models Can Self-Improve At Web Agent Tasks
di: Patel, Ajay, et al.
Pubblicazione: (2024)
di: Patel, Ajay, et al.
Pubblicazione: (2024)
Domain Gating Ensemble Networks for AI-Generated Text Detection
di: Tripathi, Arihant, et al.
Pubblicazione: (2025)
di: Tripathi, Arihant, et al.
Pubblicazione: (2025)
T-FIX: Text-Based Explanations with Features Interpretable to eXperts
di: Havaldar, Shreya, et al.
Pubblicazione: (2025)
di: Havaldar, Shreya, et al.
Pubblicazione: (2025)
Lexical Hints of Accuracy in LLM Reasoning Chains
di: Vanhoyweghen, Arne, et al.
Pubblicazione: (2025)
di: Vanhoyweghen, Arne, et al.
Pubblicazione: (2025)
Detecting LLM-Generated Text with Performance Guarantees
di: Zhou, Hongyi, et al.
Pubblicazione: (2026)
di: Zhou, Hongyi, et al.
Pubblicazione: (2026)
Sum-of-Checks: Structured Reasoning for Surgical Safety with Large Vision-Language Models
di: You, Weiqiu, et al.
Pubblicazione: (2026)
di: You, Weiqiu, et al.
Pubblicazione: (2026)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
di: Zhu, Andrew, et al.
Pubblicazione: (2025)
di: Zhu, Andrew, et al.
Pubblicazione: (2025)
Soundness-Aware Level: A Microscopic Signature that Predicts LLM Reasoning Potential
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
di: Jin, Meiqing, et al.
Pubblicazione: (2025)
di: Jin, Meiqing, et al.
Pubblicazione: (2025)
Towards Probabilistically-Sound Beam Search with Masked Language Models
di: Brooks, Creston, et al.
Pubblicazione: (2024)
di: Brooks, Creston, et al.
Pubblicazione: (2024)
PaCE: Parsimonious Concept Engineering for Large Language Models
di: Luo, Jinqi, et al.
Pubblicazione: (2024)
di: Luo, Jinqi, et al.
Pubblicazione: (2024)
Thinking with Knowledge Graphs: Enhancing LLM Reasoning Through Structured Data
di: Wu, Xue, et al.
Pubblicazione: (2024)
di: Wu, Xue, et al.
Pubblicazione: (2024)
Towards Faithful Model Explanation in NLP: A Survey
di: Lyu, Qing, et al.
Pubblicazione: (2022)
di: Lyu, Qing, et al.
Pubblicazione: (2022)
Uncovering Differences in Persuasive Language in Russian versus English Wikipedia
di: Li, Bryan, et al.
Pubblicazione: (2024)
di: Li, Bryan, et al.
Pubblicazione: (2024)
This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models
di: Li, Bryan, et al.
Pubblicazione: (2023)
di: Li, Bryan, et al.
Pubblicazione: (2023)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
di: Patel, Ajay, et al.
Pubblicazione: (2022)
di: Patel, Ajay, et al.
Pubblicazione: (2022)
Counterfactual Generation with Identifiability Guarantees
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
Probabilistic Topic Modelling with Transformer Representations
di: Reuter, Arik, et al.
Pubblicazione: (2024)
di: Reuter, Arik, et al.
Pubblicazione: (2024)
Probabilistic Reasoning with LLMs for k-anonymity Estimation
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
di: Rao, Delip, et al.
Pubblicazione: (2025) -
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
di: Rao, Delip, et al.
Pubblicazione: (2026) -
Autorubric: Unifying Rubric-based LLM Evaluation
di: Rao, Delip, et al.
Pubblicazione: (2026) -
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
di: Rao, Delip, et al.
Pubblicazione: (2026) -
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
di: Rao, Delip, et al.
Pubblicazione: (2026)