$T^5Score$: A Methodology for Automatically Assessing the Quality of LLM Generated Multi-Document Topic Sets
Fuente:
arXiv
Saved in:
| Main Authors: | Trainin, Itamar, Abend, Omri |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
by: Trainin, Itamar, et al.
Published: (2026)
by: Trainin, Itamar, et al.
Published: (2026)
Assessing the Role of Lexical Semantics in Cross-lingual Transfer through Controlled Manipulations
by: Ilani, Roy, et al.
Published: (2024)
by: Ilani, Roy, et al.
Published: (2024)
Identifying Narrative Patterns and Outliers in Holocaust Testimonies Using Topic Modeling
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs
by: Wagner, Eitan, et al.
Published: (2025)
by: Wagner, Eitan, et al.
Published: (2025)
The Challenge and Reward of Fair Play in Narrative: A Computational Approach
by: Wagner, Eitan, et al.
Published: (2025)
by: Wagner, Eitan, et al.
Published: (2025)
Locally Measuring Cross-lingual Lexical Alignment: A Domain and Word Level Perspective
by: Karidi, Taelin, et al.
Published: (2024)
by: Karidi, Taelin, et al.
Published: (2024)
Mediocrity is the key for LLM as a Judge Anchor Selection
by: Don-Yehiya, Shachar, et al.
Published: (2026)
by: Don-Yehiya, Shachar, et al.
Published: (2026)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
by: Elmakies, Avishai, et al.
Published: (2025)
by: Elmakies, Avishai, et al.
Published: (2025)
CONTESTS: a Framework for Consistency Testing of Span Probabilities in Language Models
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
Unsupervised Location Mapping for Narrative Corpora
by: Wagner, Eitan, et al.
Published: (2025)
by: Wagner, Eitan, et al.
Published: (2025)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Computational Analysis of Character Development in Holocaust Testimonies
by: Shizgal, Esther, et al.
Published: (2024)
by: Shizgal, Esther, et al.
Published: (2024)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
by: Berger, Uri, et al.
Published: (2025)
by: Berger, Uri, et al.
Published: (2025)
Measuring Pragmatic Influence in Large Language Model Instructions
by: Geng, Yilin, et al.
Published: (2026)
by: Geng, Yilin, et al.
Published: (2026)
QA-Noun: Representing Nominal Semantics via Natural Language Question-Answer Pairs
by: Tseytlin, Maria, et al.
Published: (2025)
by: Tseytlin, Maria, et al.
Published: (2025)
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
Beyond Holistic Scores: Automatic Trait-Based Quality Scoring of Argumentative Essays
by: Favero, Lucile, et al.
Published: (2026)
by: Favero, Lucile, et al.
Published: (2026)
Simplify-This: A Comparative Analysis of Prompt-Based and Fine-Tuned LLMs
by: Cohen, Eilam, et al.
Published: (2026)
by: Cohen, Eilam, et al.
Published: (2026)
A Language-agnostic Model of Child Language Acquisition
by: Mahon, Louis, et al.
Published: (2024)
by: Mahon, Louis, et al.
Published: (2024)
Can Automatic Metrics Assess High-Quality Translations?
by: Agrawal, Sweta, et al.
Published: (2024)
by: Agrawal, Sweta, et al.
Published: (2024)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
by: Das, Sourya Dipta, et al.
Published: (2024)
by: Das, Sourya Dipta, et al.
Published: (2024)
Cross-linguistically Consistent Semantic and Syntactic Annotation of Child-directed Speech
by: Szubert, Ida, et al.
Published: (2021)
by: Szubert, Ida, et al.
Published: (2021)
LLM Reading Tea Leaves: Automatically Evaluating Topic Models with Large Language Models
by: Yang, Xiaohao, et al.
Published: (2024)
by: Yang, Xiaohao, et al.
Published: (2024)
AutoTM 2.0: Automatic Topic Modeling Framework for Documents Analysis
by: Khodorchenko, Maria, et al.
Published: (2024)
by: Khodorchenko, Maria, et al.
Published: (2024)
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
DiagGPT: An LLM-based and Multi-agent Dialogue System with Automatic Topic Management for Flexible Task-Oriented Dialogue
by: Cao, Lang
Published: (2023)
by: Cao, Lang
Published: (2023)
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
by: Latif, Ehsan, et al.
Published: (2023)
by: Latif, Ehsan, et al.
Published: (2023)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
by: Oleson, Jon
Published: (2024)
by: Oleson, Jon
Published: (2024)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
by: Li, Chuyuan, et al.
Published: (2025)
by: Li, Chuyuan, et al.
Published: (2025)
Methodological Rigour in Algorithm Application: An Illustration of Topic Modelling Algorithm
by: Amadoru, Malmi
Published: (2025)
by: Amadoru, Malmi
Published: (2025)
Generating Benchmarks for Factuality Evaluation of Language Models
by: Muhlgay, Dor, et al.
Published: (2023)
by: Muhlgay, Dor, et al.
Published: (2023)
NLP Verification: Towards a General Methodology for Certifying Robustness
by: Casadio, Marco, et al.
Published: (2024)
by: Casadio, Marco, et al.
Published: (2024)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
Similar Items
-
Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
by: Trainin, Itamar, et al.
Published: (2026) -
Assessing the Role of Lexical Semantics in Cross-lingual Transfer through Controlled Manipulations
by: Ilani, Roy, et al.
Published: (2024) -
Identifying Narrative Patterns and Outliers in Holocaust Testimonies Using Topic Modeling
by: Ifergan, Maxim, et al.
Published: (2024) -
Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs
by: Wagner, Eitan, et al.
Published: (2025) -
The Challenge and Reward of Fair Play in Narrative: A Computational Approach
by: Wagner, Eitan, et al.
Published: (2025)