Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Zhen, Trivedi, Shubhendu, Sun, Jimeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
von: Lin, Zhen, et al.
Veröffentlicht: (2023)
von: Lin, Zhen, et al.
Veröffentlicht: (2023)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation
von: Lin, Qinhong, et al.
Veröffentlicht: (2024)
von: Lin, Qinhong, et al.
Veröffentlicht: (2024)
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
DiffScore: Text Evaluation Beyond Autoregressive Likelihood
von: Lai, Wen, et al.
Veröffentlicht: (2026)
von: Lai, Wen, et al.
Veröffentlicht: (2026)
Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning
von: Wu, John, et al.
Veröffentlicht: (2024)
von: Wu, John, et al.
Veröffentlicht: (2024)
Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
von: Chen, Shiting, et al.
Veröffentlicht: (2025)
von: Chen, Shiting, et al.
Veröffentlicht: (2025)
AutoRD: An Automatic and End-to-End System for Rare Disease Knowledge Graph Construction Based on Ontologies-enhanced Large Language Models
von: Cao, Lang, et al.
Veröffentlicht: (2024)
von: Cao, Lang, et al.
Veröffentlicht: (2024)
SUGAR: Leveraging Contextual Confidence for Smarter Retrieval
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2025)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2025)
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
von: Das, Trisha, et al.
Veröffentlicht: (2024)
von: Das, Trisha, et al.
Veröffentlicht: (2024)
TTM-RE: Memory-Augmented Document-Level Relation Extraction
von: Gao, Chufan, et al.
Veröffentlicht: (2024)
von: Gao, Chufan, et al.
Veröffentlicht: (2024)
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research
von: Wang, Zifeng, et al.
Veröffentlicht: (2025)
von: Wang, Zifeng, et al.
Veröffentlicht: (2025)
Meta-DiffuB: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration
von: Chuang, Yun-Yen, et al.
Veröffentlicht: (2024)
von: Chuang, Yun-Yen, et al.
Veröffentlicht: (2024)
TaSR-RAG: Taxonomy-guided Structured Reasoning for Retrieval-Augmented Generation
von: Sun, Jiashuo, et al.
Veröffentlicht: (2026)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2026)
Panacea: A foundation model for clinical trial search, summarization, design, and recruitment
von: Lin, Jiacheng, et al.
Veröffentlicht: (2024)
von: Lin, Jiacheng, et al.
Veröffentlicht: (2024)
Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity
von: Jing, Bowen, et al.
Veröffentlicht: (2025)
von: Jing, Bowen, et al.
Veröffentlicht: (2025)
Enhancing RWKV-based Language Models for Long-Sequence Text Generation
von: Pan, Xinghan
Veröffentlicht: (2025)
von: Pan, Xinghan
Veröffentlicht: (2025)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
von: Cai, Yida, et al.
Veröffentlicht: (2025)
von: Cai, Yida, et al.
Veröffentlicht: (2025)
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
von: Chakravarty, Abhirup, et al.
Veröffentlicht: (2025)
von: Chakravarty, Abhirup, et al.
Veröffentlicht: (2025)
Contextualization Distillation from Large Language Model for Knowledge Graph Completion
von: Li, Dawei, et al.
Veröffentlicht: (2024)
von: Li, Dawei, et al.
Veröffentlicht: (2024)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
von: Balkır, Esma, et al.
Veröffentlicht: (2026)
von: Balkır, Esma, et al.
Veröffentlicht: (2026)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
Multi-Perspective Consistency Enhances Confidence Estimation in Large Language Models
von: Wang, Pei, et al.
Veröffentlicht: (2024)
von: Wang, Pei, et al.
Veröffentlicht: (2024)
Graph Alignment Topology as an Inductive Bias for Grounding Detection
von: Landes, Paul, et al.
Veröffentlicht: (2026)
von: Landes, Paul, et al.
Veröffentlicht: (2026)
CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis
von: Zhen, Hao, et al.
Veröffentlicht: (2025)
von: Zhen, Hao, et al.
Veröffentlicht: (2025)
Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
von: Zhao, Zheng, et al.
Veröffentlicht: (2024)
von: Zhao, Zheng, et al.
Veröffentlicht: (2024)
Neural Contextual Reinforcement Framework for Logical Structure Language Generation
von: Irvin, Marcus, et al.
Veröffentlicht: (2025)
von: Irvin, Marcus, et al.
Veröffentlicht: (2025)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
von: Pathak, Manas, et al.
Veröffentlicht: (2026)
von: Pathak, Manas, et al.
Veröffentlicht: (2026)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
von: Qi, Siya, et al.
Veröffentlicht: (2024)
von: Qi, Siya, et al.
Veröffentlicht: (2024)
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
von: Matys, Piotr, et al.
Veröffentlicht: (2025)
von: Matys, Piotr, et al.
Veröffentlicht: (2025)
Sarang at DEFACTIFY 4.0: Detecting AI-Generated Text Using Noised Data and an Ensemble of DeBERTa Models
von: Trivedi, Avinash, et al.
Veröffentlicht: (2025)
von: Trivedi, Avinash, et al.
Veröffentlicht: (2025)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
von: Chu, SeongYeub, et al.
Veröffentlicht: (2024)
von: Chu, SeongYeub, et al.
Veröffentlicht: (2024)
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
von: Song, Yiliang, et al.
Veröffentlicht: (2026)
von: Song, Yiliang, et al.
Veröffentlicht: (2026)
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
Fine-Tuning Medical Language Models for Enhanced Long-Contextual Understanding and Domain Expertise
von: Yang, Qimin, et al.
Veröffentlicht: (2024)
von: Yang, Qimin, et al.
Veröffentlicht: (2024)
Position Paper: Generalized grammar rules and structure-based generalization beyond classical equivariance for lexical tasks and transduction
von: Petrache, Mircea, et al.
Veröffentlicht: (2024)
von: Petrache, Mircea, et al.
Veröffentlicht: (2024)
s3: You Don't Need That Much Data to Train a Search Agent via RL
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
von: Lin, Zhen, et al.
Veröffentlicht: (2023) -
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025) -
Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation
von: Lin, Qinhong, et al.
Veröffentlicht: (2024) -
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024) -
DiffScore: Text Evaluation Beyond Autoregressive Likelihood
von: Lai, Wen, et al.
Veröffentlicht: (2026)