$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Darrin, Maxime, Formont, Philippe, Cheung, Jackie Chi Kit, Piantanida, Pablo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
por: Darrin, Maxime, et al.
Publicado: (2024)
por: Darrin, Maxime, et al.
Publicado: (2024)
When is an Embedding Model More Promising than Another?
por: Darrin, Maxime, et al.
Publicado: (2024)
por: Darrin, Maxime, et al.
Publicado: (2024)
MolRGen: A Training and Evaluation Setting for De Novo Molecular Generation with Reasonning Models
por: Formont, Philippe, et al.
Publicado: (2026)
por: Formont, Philippe, et al.
Publicado: (2026)
PreSumm: Predicting Summarization Performance Without Summarizing
por: Koniaev, Steven, et al.
Publicado: (2025)
por: Koniaev, Steven, et al.
Publicado: (2025)
Unsupervised Layer-wise Score Aggregation for Textual OOD Detection
por: Darrin, Maxime, et al.
Publicado: (2023)
por: Darrin, Maxime, et al.
Publicado: (2023)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
por: Piano, Cesare Spinoso-Di, et al.
Publicado: (2025)
por: Piano, Cesare Spinoso-Di, et al.
Publicado: (2025)
Learning Task-Agnostic Representations through Multi-Teacher Distillation
por: Formont, Philippe, et al.
Publicado: (2025)
por: Formont, Philippe, et al.
Publicado: (2025)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
por: Yu, Lei, et al.
Publicado: (2024)
por: Yu, Lei, et al.
Publicado: (2024)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
por: Flores, Lorenzo Jaime Yu, et al.
Publicado: (2025)
por: Flores, Lorenzo Jaime Yu, et al.
Publicado: (2025)
Rainproof: An Umbrella To Shield Text Generators From Out-Of-Distribution Data
por: Darrin, Maxime, et al.
Publicado: (2022)
por: Darrin, Maxime, et al.
Publicado: (2022)
A Strong Baseline for Molecular Few-Shot Learning
por: Formont, Philippe, et al.
Publicado: (2024)
por: Formont, Philippe, et al.
Publicado: (2024)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
por: Chae, Kyubyung, et al.
Publicado: (2024)
por: Chae, Kyubyung, et al.
Publicado: (2024)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
por: Gao, Jie, et al.
Publicado: (2026)
por: Gao, Jie, et al.
Publicado: (2026)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
por: Porada, Ian, et al.
Publicado: (2024)
por: Porada, Ian, et al.
Publicado: (2024)
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
por: Shayegh, Behzad, et al.
Publicado: (2024)
por: Shayegh, Behzad, et al.
Publicado: (2024)
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
por: AlShikh, Waseem, et al.
Publicado: (2025)
por: AlShikh, Waseem, et al.
Publicado: (2025)
STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
por: Chhetri, Tek Raj, et al.
Publicado: (2025)
por: Chhetri, Tek Raj, et al.
Publicado: (2025)
FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation
por: Otmakhova, Yulia, et al.
Publicado: (2025)
por: Otmakhova, Yulia, et al.
Publicado: (2025)
$\texttt{LM}^\texttt{2}$: A Simple Society of Language Models Solves Complex Reasoning
por: Juneja, Gurusha, et al.
Publicado: (2024)
por: Juneja, Gurusha, et al.
Publicado: (2024)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
por: Lu, Yen-Ju, et al.
Publicado: (2025)
por: Lu, Yen-Ju, et al.
Publicado: (2025)
Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review
por: Croxford, Emma, et al.
Publicado: (2024)
por: Croxford, Emma, et al.
Publicado: (2024)
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
por: Aly, Walid Mohamed, et al.
Publicado: (2025)
por: Aly, Walid Mohamed, et al.
Publicado: (2025)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
por: Bück-Kaeffer, Aurélien, et al.
Publicado: (2025)
por: Bück-Kaeffer, Aurélien, et al.
Publicado: (2025)
How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
por: Elkins, Sabina, et al.
Publicado: (2024)
por: Elkins, Sabina, et al.
Publicado: (2024)
STORYSUMM: Evaluating Faithfulness in Story Summarization
por: Subbiah, Melanie, et al.
Publicado: (2024)
por: Subbiah, Melanie, et al.
Publicado: (2024)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
por: Yuan, Dong, et al.
Publicado: (2024)
por: Yuan, Dong, et al.
Publicado: (2024)
Extrinsically-Focused Evaluation of Omissions in Medical Summarization
por: Schumacher, Elliot, et al.
Publicado: (2023)
por: Schumacher, Elliot, et al.
Publicado: (2023)
Information-Theoretic Distillation for Reference-less Summarization
por: Jung, Jaehun, et al.
Publicado: (2024)
por: Jung, Jaehun, et al.
Publicado: (2024)
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
por: Pallagani, Vishal, et al.
Publicado: (2024)
por: Pallagani, Vishal, et al.
Publicado: (2024)
QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
por: Ghebriout, Mohamed Imed Eddine, et al.
Publicado: (2025)
por: Ghebriout, Mohamed Imed Eddine, et al.
Publicado: (2025)
Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialog
por: Estienne, Lautaro, et al.
Publicado: (2025)
por: Estienne, Lautaro, et al.
Publicado: (2025)
$\texttt{SEM-CTRL}$: Semantically Controlled Decoding
por: Albinhassan, Mohammad, et al.
Publicado: (2025)
por: Albinhassan, Mohammad, et al.
Publicado: (2025)
Does This Summary Answer My Question? Modeling Query-Focused Summary Readers with Rational Speech Acts
por: Piano, Cesare Spinoso-Di, et al.
Publicado: (2024)
por: Piano, Cesare Spinoso-Di, et al.
Publicado: (2024)
Learning to Maximize Mutual Information for Chain-of-Thought Distillation
por: Chen, Xin, et al.
Publicado: (2024)
por: Chen, Xin, et al.
Publicado: (2024)
Think Carefully and Check Again! Meta-Generation Unlocking LLMs for Low-Resource Cross-Lingual Summarization
por: Li, Zhecheng, et al.
Publicado: (2024)
por: Li, Zhecheng, et al.
Publicado: (2024)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
por: He, Muyu, et al.
Publicado: (2026)
por: He, Muyu, et al.
Publicado: (2026)
Statistical Deficiency for Task Inclusion Estimation
por: Fosse, Loïc, et al.
Publicado: (2025)
por: Fosse, Loïc, et al.
Publicado: (2025)
FamiCom: Further Demystifying Prompts for Language Models with Task-Agnostic Performance Estimation
por: Li, Bangzheng, et al.
Publicado: (2024)
por: Li, Bangzheng, et al.
Publicado: (2024)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
por: Zhang, Shiyue, et al.
Publicado: (2024)
por: Zhang, Shiyue, et al.
Publicado: (2024)
Ejemplares similares
-
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
por: Darrin, Maxime, et al.
Publicado: (2024) -
When is an Embedding Model More Promising than Another?
por: Darrin, Maxime, et al.
Publicado: (2024) -
MolRGen: A Training and Evaluation Setting for De Novo Molecular Generation with Reasonning Models
por: Formont, Philippe, et al.
Publicado: (2026) -
PreSumm: Predicting Summarization Performance Without Summarizing
por: Koniaev, Steven, et al.
Publicado: (2025) -
Unsupervised Layer-wise Score Aggregation for Textual OOD Detection
por: Darrin, Maxime, et al.
Publicado: (2023)