DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Yiqing, Zhang, Sheng, Cheng, Hao, Liu, Pengfei, Gero, Zelalem, Wong, Cliff, Naumann, Tristan, Poon, Hoifung, Rose, Carolyn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
von: Gero, Zelalem, et al.
Veröffentlicht: (2024)
von: Gero, Zelalem, et al.
Veröffentlicht: (2024)
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
TRIALSCOPE: A Unifying Causal Framework for Scaling Real-World Evidence Generation with Biomedical Language Models
von: González, Javier, et al.
Veröffentlicht: (2023)
von: González, Javier, et al.
Veröffentlicht: (2023)
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2025)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2025)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
Generative Medical Event Models Improve with Scale
von: Waxler, Shane, et al.
Veröffentlicht: (2025)
von: Waxler, Shane, et al.
Veröffentlicht: (2025)
Improving Model Factuality with Fine-grained Critique-based Evaluator
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
von: Xie, Yiqing, et al.
Veröffentlicht: (2025)
von: Xie, Yiqing, et al.
Veröffentlicht: (2025)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
T-Rex: Text-assisted Retrosynthesis Prediction
von: Liu, Yifeng, et al.
Veröffentlicht: (2024)
von: Liu, Yifeng, et al.
Veröffentlicht: (2024)
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
Universal Abstraction: Harnessing Frontier Models to Structure Real-World Data at Scale
von: Wong, Cliff, et al.
Veröffentlicht: (2025)
von: Wong, Cliff, et al.
Veröffentlicht: (2025)
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
Learning Sparse Visual Representations via Spatial-Semantic Factorization
von: Zhao, Theodore Zhengde, et al.
Veröffentlicht: (2026)
von: Zhao, Theodore Zhengde, et al.
Veröffentlicht: (2026)
CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation
von: Unell, Alyssa, et al.
Veröffentlicht: (2025)
von: Unell, Alyssa, et al.
Veröffentlicht: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
Pareto Optimal Learning for Estimating Large Language Model Errors
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
Cautionary Tales on Synthetic Controls in Survival Analyses
von: Curth, Alicia, et al.
Veröffentlicht: (2023)
von: Curth, Alicia, et al.
Veröffentlicht: (2023)
Scaling medical imaging report generation with multimodal reinforcement learning
von: Liu, Qianchu, et al.
Veröffentlicht: (2026)
von: Liu, Qianchu, et al.
Veröffentlicht: (2026)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
von: Afzal, Anum, et al.
Veröffentlicht: (2025)
von: Afzal, Anum, et al.
Veröffentlicht: (2025)
ChartLens: Fine-grained Visual Attribution in Charts
von: Suri, Manan, et al.
Veröffentlicht: (2025)
von: Suri, Manan, et al.
Veröffentlicht: (2025)
Fine-grained Image Quality Assessment for Perceptual Image Restoration
von: Sheng, Xiangfei, et al.
Veröffentlicht: (2025)
von: Sheng, Xiangfei, et al.
Veröffentlicht: (2025)
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
von: Ossowski, Timothy, et al.
Veröffentlicht: (2026)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2026)
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
von: Shankar, Shreya, et al.
Veröffentlicht: (2024)
von: Shankar, Shreya, et al.
Veröffentlicht: (2024)
Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
Some aspects of the ecology of the limnoplankton, with special reference to the phytoplankton. [Translation from: Svensk Botanisk Tidskrift 13(2) 129-163, 1919.]
von: Naumann, E.
Veröffentlicht: (1919)
von: Naumann, E.
Veröffentlicht: (1919)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
von: Zhang, Shiyue, et al.
Veröffentlicht: (2024)
von: Zhang, Shiyue, et al.
Veröffentlicht: (2024)
Evaluating Representational Similarity Measures from the Lens of Functional Correspondence
von: Bo, Yiqing, et al.
Veröffentlicht: (2024)
von: Bo, Yiqing, et al.
Veröffentlicht: (2024)
A Perspective for Adapting Generalist AI to Specialized Medical AI Applications and Their Challenges
von: Wang, Zifeng, et al.
Veröffentlicht: (2024)
von: Wang, Zifeng, et al.
Veröffentlicht: (2024)
Offset Unlearning for Large Language Models
von: Huang, James Y., et al.
Veröffentlicht: (2024)
von: Huang, James Y., et al.
Veröffentlicht: (2024)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
von: Cui, Wanqing, et al.
Veröffentlicht: (2024)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
von: Zhang, Sheng, et al.
Veröffentlicht: (2023)
von: Zhang, Sheng, et al.
Veröffentlicht: (2023)
Fine-grained Text to Image Synthesis
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
von: Gero, Zelalem, et al.
Veröffentlicht: (2024) -
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025) -
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
von: Zhang, Sheng, et al.
Veröffentlicht: (2025) -
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025) -
TRIALSCOPE: A Unifying Causal Framework for Scaling Real-World Evidence Generation with Biomedical Language Models
von: González, Javier, et al.
Veröffentlicht: (2023)