STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Guanghui, Yu, Jinze, Zhang, Xing, Jiang, Dayuan, Song, Yin, Deb, Tomal, Liu, Xuefeng, He, Peiyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
by: Huang, Minghui
Published: (2025)
by: Huang, Minghui
Published: (2025)
GEM: Graph-Enhanced Mixture-of-Experts with ReAct Agents for Dialogue State Tracking
by: Zhu, Ziqi, et al.
Published: (2026)
by: Zhu, Ziqi, et al.
Published: (2026)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
by: Cavalin, Paulo, et al.
Published: (2025)
by: Cavalin, Paulo, et al.
Published: (2025)
GenAI-DrawIO-Creator: A Framework for Automated Diagram Generation
by: Yu, Jinze, et al.
Published: (2026)
by: Yu, Jinze, et al.
Published: (2026)
Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
ADEQA: A Question Answer based approach for joint ADE-Suspect Extraction using Sequence-To-Sequence Transformers
by: Arannil, Vinayak, et al.
Published: (2024)
by: Arannil, Vinayak, et al.
Published: (2024)
Meaning Typed Prompting: A Technique for Efficient, Reliable Structured Output Generation
by: Irugalbandara, Chandra
Published: (2024)
by: Irugalbandara, Chandra
Published: (2024)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
by: Hong, Zhaochen, et al.
Published: (2025)
by: Hong, Zhaochen, et al.
Published: (2025)
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models
by: Singh, Abhinav Kumar, et al.
Published: (2026)
by: Singh, Abhinav Kumar, et al.
Published: (2026)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
by: Yao, Jing, et al.
Published: (2024)
by: Yao, Jing, et al.
Published: (2024)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
by: Tao, Zhen, et al.
Published: (2024)
by: Tao, Zhen, et al.
Published: (2024)
RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
by: Wang, Haofeng, et al.
Published: (2025)
by: Wang, Haofeng, et al.
Published: (2025)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
by: Yu, Fangyi, et al.
Published: (2025)
by: Yu, Fangyi, et al.
Published: (2025)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
by: Pandey, Atharva, et al.
Published: (2025)
by: Pandey, Atharva, et al.
Published: (2025)
Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams
by: Tang, Xiuxiu, et al.
Published: (2026)
by: Tang, Xiuxiu, et al.
Published: (2026)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026)
by: Galeone, Cosimo, et al.
Published: (2026)
Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
by: Bologna, Federica, et al.
Published: (2026)
by: Bologna, Federica, et al.
Published: (2026)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
by: Xu, Yunqi, et al.
Published: (2024)
by: Xu, Yunqi, et al.
Published: (2024)
Large Language Model Reasoning Failures
by: Song, Peiyang, et al.
Published: (2026)
by: Song, Peiyang, et al.
Published: (2026)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models
by: Han, Pengrui, et al.
Published: (2024)
by: Han, Pengrui, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
Harnessing Consistency for Robust Test-Time LLM Ensemble
by: Zeng, Zhichen, et al.
Published: (2025)
by: Zeng, Zhichen, et al.
Published: (2025)
An LLM Maturity Model for Reliable and Transparent Text-to-Query
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
by: Yoo, Yongmin, et al.
Published: (2025)
by: Yoo, Yongmin, et al.
Published: (2025)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
by: Kostić, Bogdan, et al.
Published: (2026)
by: Kostić, Bogdan, et al.
Published: (2026)
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
by: Li, Keyu, et al.
Published: (2025)
by: Li, Keyu, et al.
Published: (2025)
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
by: Chen, Sijia, et al.
Published: (2026)
by: Chen, Sijia, et al.
Published: (2026)
MSI-Agent: Incorporating Multi-Scale Insight into Embodied Agents for Superior Planning and Decision-Making
by: Fu, Dayuan, et al.
Published: (2024)
by: Fu, Dayuan, et al.
Published: (2024)
Output-Space Search: Targeting LLM Generations in a Frozen Encoder-Defined Output Space
by: Materzok, Tobias
Published: (2026)
by: Materzok, Tobias
Published: (2026)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024)
by: Liang, Jiashuo, et al.
Published: (2024)
Similar Items
-
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
by: Huang, Minghui
Published: (2025) -
GEM: Graph-Enhanced Mixture-of-Experts with ReAct Agents for Dialogue State Tracking
by: Zhu, Ziqi, et al.
Published: (2026) -
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026) -
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
by: Zhang, Xing, et al.
Published: (2026) -
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
by: Zhang, Xing, et al.
Published: (2026)