Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
Fuente:
arXiv
Saved in:
| Main Authors: | Gero, Zelalem, Singh, Chandan, Xie, Yiqing, Zhang, Sheng, Subramanian, Praveen, Vozila, Paul, Naumann, Tristan, Gao, Jianfeng, Poon, Hoifung |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
by: Xie, Yiqing, et al.
Published: (2023)
by: Xie, Yiqing, et al.
Published: (2023)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
Exploring Scaling Laws for EHR Foundation Models
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
Test-Time Learning with an Evolving Library
by: Xu, Weijia, et al.
Published: (2026)
by: Xu, Weijia, et al.
Published: (2026)
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
by: Liu, Qianchu, et al.
Published: (2025)
by: Liu, Qianchu, et al.
Published: (2025)
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
by: Singh, Chandan, et al.
Published: (2026)
by: Singh, Chandan, et al.
Published: (2026)
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
by: Ossowski, Timothy, et al.
Published: (2025)
by: Ossowski, Timothy, et al.
Published: (2025)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
by: Huang, James Y., et al.
Published: (2025)
by: Huang, James Y., et al.
Published: (2025)
TRIALSCOPE: A Unifying Causal Framework for Scaling Real-World Evidence Generation with Biomedical Language Models
by: González, Javier, et al.
Published: (2023)
by: González, Javier, et al.
Published: (2023)
AURAD: Anatomy-Pathology Unified Radiology Synthesis with Progressive Representations
by: Ding, Shuhan, et al.
Published: (2025)
by: Ding, Shuhan, et al.
Published: (2025)
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
by: Ossowski, Timothy, et al.
Published: (2026)
by: Ossowski, Timothy, et al.
Published: (2026)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
by: Huang, James Y., et al.
Published: (2025)
by: Huang, James Y., et al.
Published: (2025)
T-Rex: Text-assisted Retrosynthesis Prediction
by: Liu, Yifeng, et al.
Published: (2024)
by: Liu, Yifeng, et al.
Published: (2024)
Text Generation Beyond Discrete Token Sampling
by: Zhuang, Yufan, et al.
Published: (2025)
by: Zhuang, Yufan, et al.
Published: (2025)
Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers
by: Gong, Linyuan, et al.
Published: (2023)
by: Gong, Linyuan, et al.
Published: (2023)
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
by: Zhou, Wenxuan, et al.
Published: (2023)
by: Zhou, Wenxuan, et al.
Published: (2023)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
by: Xu, Nan, et al.
Published: (2024)
by: Xu, Nan, et al.
Published: (2024)
Learning Sparse Visual Representations via Spatial-Semantic Factorization
by: Zhao, Theodore Zhengde, et al.
Published: (2026)
by: Zhao, Theodore Zhengde, et al.
Published: (2026)
Universal Abstraction: Harnessing Frontier Models to Structure Real-World Data at Scale
by: Wong, Cliff, et al.
Published: (2025)
by: Wong, Cliff, et al.
Published: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
by: Zhao, Theodore, et al.
Published: (2024)
by: Zhao, Theodore, et al.
Published: (2024)
Generative Medical Event Models Improve with Scale
by: Waxler, Shane, et al.
Published: (2025)
by: Waxler, Shane, et al.
Published: (2025)
Pareto Optimal Learning for Estimating Large Language Model Errors
by: Zhao, Theodore, et al.
Published: (2023)
by: Zhao, Theodore, et al.
Published: (2023)
Cautionary Tales on Synthetic Controls in Survival Analyses
by: Curth, Alicia, et al.
Published: (2023)
by: Curth, Alicia, et al.
Published: (2023)
Scaling medical imaging report generation with multimodal reinforcement learning
by: Liu, Qianchu, et al.
Published: (2026)
by: Liu, Qianchu, et al.
Published: (2026)
Improving LLM Reasoning with Homophily-aware Structural and Semantic Text-Attributed Graph Compression
by: Di, Zijun, et al.
Published: (2026)
by: Di, Zijun, et al.
Published: (2026)
On the Role of Summary Content Units in Text Summarization Evaluation
by: Nawrath, Marcel, et al.
Published: (2024)
by: Nawrath, Marcel, et al.
Published: (2024)
Offset Unlearning for Large Language Models
by: Huang, James Y., et al.
Published: (2024)
by: Huang, James Y., et al.
Published: (2024)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Vector-ICL: In-context Learning with Continuous Vector Representations
by: Zhuang, Yufan, et al.
Published: (2024)
by: Zhuang, Yufan, et al.
Published: (2024)
Learning a Decision Tree Algorithm with Transformers
by: Zhuang, Yufan, et al.
Published: (2024)
by: Zhuang, Yufan, et al.
Published: (2024)
TextNow Transition Programme: Evaluation Report and Executive Summary
by: Maxwell, Bronwen, et al.
Published: (2014)
by: Maxwell, Bronwen, et al.
Published: (2014)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
Investigation of entanglement in $N = Z$ nuclei within no-core shell model
by: Sarma, Chandan, et al.
Published: (2024)
by: Sarma, Chandan, et al.
Published: (2024)
Ab initio no-core shell-model study of $^{20-23}$Na isotopes
by: Sarma, Chandan, et al.
Published: (2023)
by: Sarma, Chandan, et al.
Published: (2023)
IEEE 802.11be Wi-Fi 7: Feature Summary and Performance Evaluation
by: Liu, Xiaoqian, et al.
Published: (2023)
by: Liu, Xiaoqian, et al.
Published: (2023)
Rethinking Interpretability in the Era of Large Language Models
by: Singh, Chandan, et al.
Published: (2024)
by: Singh, Chandan, et al.
Published: (2024)
CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation
by: Unell, Alyssa, et al.
Published: (2025)
by: Unell, Alyssa, et al.
Published: (2025)
STRUM-LLM: Attributed and Structured Contrastive Summarization
by: Gunel, Beliz, et al.
Published: (2024)
by: Gunel, Beliz, et al.
Published: (2024)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
Similar Items
-
DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
by: Xie, Yiqing, et al.
Published: (2023) -
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
by: Zhang, Sheng, et al.
Published: (2025) -
Exploring Scaling Laws for EHR Foundation Models
by: Zhang, Sheng, et al.
Published: (2025) -
Test-Time Learning with an Evolving Library
by: Xu, Weijia, et al.
Published: (2026) -
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
by: Liu, Qianchu, et al.
Published: (2025)