*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Lemesle, Quentin, Jourdan, Léane, Munson, Daisy, Alain, Pierre, Chevelu, Jonathan, Delhay, Arnaud, Lolive, Damien |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Identifying Reliable Evaluation Metrics for Scientific Text Revision
por: Jourdan, Léane, et al.
Publicado: (2025)
por: Jourdan, Léane, et al.
Publicado: (2025)
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
por: Jourdan, Leane, et al.
Publicado: (2024)
por: Jourdan, Leane, et al.
Publicado: (2024)
Text revision in Scientific Writing Assistance: An Overview
por: Jourdan, Léane, et al.
Publicado: (2023)
por: Jourdan, Léane, et al.
Publicado: (2023)
Multi-level SSL Feature Gating for Audio Deepfake Detection
por: Tran, Hoan My, et al.
Publicado: (2025)
por: Tran, Hoan My, et al.
Publicado: (2025)
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
por: Jourdan, Léane, et al.
Publicado: (2025)
por: Jourdan, Léane, et al.
Publicado: (2025)
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
por: Jourdan, Léane, et al.
Publicado: (2026)
por: Jourdan, Léane, et al.
Publicado: (2026)
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
por: Jourdan, Fanny
Publicado: (2024)
por: Jourdan, Fanny
Publicado: (2024)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
por: Jourdan, Fanny, et al.
Publicado: (2025)
por: Jourdan, Fanny, et al.
Publicado: (2025)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
por: Fu, Xiao, et al.
Publicado: (2025)
por: Fu, Xiao, et al.
Publicado: (2025)
What do the metrics mean? A critical analysis of the use of Automated Evaluation Metrics in Interpreting
por: Downie, Jonathan, et al.
Publicado: (2026)
por: Downie, Jonathan, et al.
Publicado: (2026)
TAPS: Tool-Augmented Personalisation via Structured Tagging
por: Taktasheva, Ekaterina, et al.
Publicado: (2025)
por: Taktasheva, Ekaterina, et al.
Publicado: (2025)
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
por: Dowerah, Sandipana, et al.
Publicado: (2025)
por: Dowerah, Sandipana, et al.
Publicado: (2025)
Tau-Eval: A Unified Evaluation Framework for Useful and Private Text Anonymization
por: Loiseau, Gabriel, et al.
Publicado: (2025)
por: Loiseau, Gabriel, et al.
Publicado: (2025)
Evaluating the Evaluators: Are readability metrics good measures of readability?
por: Cachola, Isabel, et al.
Publicado: (2025)
por: Cachola, Isabel, et al.
Publicado: (2025)
Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
por: Dhawan, Aashish, et al.
Publicado: (2026)
por: Dhawan, Aashish, et al.
Publicado: (2026)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
por: Kirk, Hannah Rose, et al.
Publicado: (2026)
por: Kirk, Hannah Rose, et al.
Publicado: (2026)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
por: Ross, Candace, et al.
Publicado: (2024)
por: Ross, Candace, et al.
Publicado: (2024)
A Multilingual, Large-Scale Study of the Interplay between LLM Safeguards, Personalisation, and Disinformation
por: Leite, João A., et al.
Publicado: (2025)
por: Leite, João A., et al.
Publicado: (2025)
Waste Not, Want Not; Recycled Gumbel Noise Improves Consistency in Natural Language Generation
por: de Mijolla, Damien, et al.
Publicado: (2025)
por: de Mijolla, Damien, et al.
Publicado: (2025)
Towards Stable and Personalised Profiles for Lexical Alignment in Spoken Human-Agent Dialogue
por: Schaaij, Keara, et al.
Publicado: (2025)
por: Schaaij, Keara, et al.
Publicado: (2025)
Passive Learning of Lattice Automata from Recurrent Neural Networks
por: Slimi, Jaouhar, et al.
Publicado: (2025)
por: Slimi, Jaouhar, et al.
Publicado: (2025)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
por: Vázquez, Alain, et al.
Publicado: (2026)
por: Vázquez, Alain, et al.
Publicado: (2026)
Personalised Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code Generation
por: Chen, Hailin, et al.
Publicado: (2023)
por: Chen, Hailin, et al.
Publicado: (2023)
Structured Context Recomposition for Large Language Models Using Probabilistic Layer Realignment
por: Teel, Jonathan, et al.
Publicado: (2025)
por: Teel, Jonathan, et al.
Publicado: (2025)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
por: Chen, Lu, et al.
Publicado: (2024)
por: Chen, Lu, et al.
Publicado: (2024)
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
por: Piot, Paloma, et al.
Publicado: (2025)
por: Piot, Paloma, et al.
Publicado: (2025)
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
por: Sileo, Damien
Publicado: (2025)
por: Sileo, Damien
Publicado: (2025)
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
por: Sileo, Damien
Publicado: (2024)
por: Sileo, Damien
Publicado: (2024)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
por: Gajewska, Ewelina, et al.
Publicado: (2026)
por: Gajewska, Ewelina, et al.
Publicado: (2026)
MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts
por: Lanzeray, Etienne, et al.
Publicado: (2026)
por: Lanzeray, Etienne, et al.
Publicado: (2026)
LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
por: Gupta, Sammriddh, et al.
Publicado: (2025)
por: Gupta, Sammriddh, et al.
Publicado: (2025)
Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
por: Bernas, Raphael, et al.
Publicado: (2026)
por: Bernas, Raphael, et al.
Publicado: (2026)
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop
por: Zhao, Runcong, et al.
Publicado: (2025)
por: Zhao, Runcong, et al.
Publicado: (2025)
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation
por: Goto, Takumi, et al.
Publicado: (2025)
por: Goto, Takumi, et al.
Publicado: (2025)
ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability
por: Poché, Antonin, et al.
Publicado: (2025)
por: Poché, Antonin, et al.
Publicado: (2025)
AI for Monitoring and Classifying Data Used in Research Literature
por: Macalaba, Rafael, et al.
Publicado: (2026)
por: Macalaba, Rafael, et al.
Publicado: (2026)
Detecting Turkish Synonyms Used in Different Time Periods
por: Yazar, Umur Togay, et al.
Publicado: (2024)
por: Yazar, Umur Togay, et al.
Publicado: (2024)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
por: Vincent, Sebastian, et al.
Publicado: (2023)
por: Vincent, Sebastian, et al.
Publicado: (2023)
Developing a Mixed-Methods Pipeline for Community-Oriented Digitization of Kwak'wala Legacy Texts
por: Agarwal, Milind, et al.
Publicado: (2025)
por: Agarwal, Milind, et al.
Publicado: (2025)
Ink and Individuality: Crafting a Personalised Narrative in the Age of LLMs
por: Wasi, Azmine Toushik, et al.
Publicado: (2024)
por: Wasi, Azmine Toushik, et al.
Publicado: (2024)
Ejemplares similares
-
Identifying Reliable Evaluation Metrics for Scientific Text Revision
por: Jourdan, Léane, et al.
Publicado: (2025) -
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
por: Jourdan, Leane, et al.
Publicado: (2024) -
Text revision in Scientific Writing Assistance: An Overview
por: Jourdan, Léane, et al.
Publicado: (2023) -
Multi-level SSL Feature Gating for Audio Deepfake Detection
por: Tran, Hoan My, et al.
Publicado: (2025) -
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
por: Jourdan, Léane, et al.
Publicado: (2025)