*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Lemesle, Quentin, Jourdan, Léane, Munson, Daisy, Alain, Pierre, Chevelu, Jonathan, Delhay, Arnaud, Lolive, Damien |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Identifying Reliable Evaluation Metrics for Scientific Text Revision
di: Jourdan, Léane, et al.
Pubblicazione: (2025)
di: Jourdan, Léane, et al.
Pubblicazione: (2025)
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
di: Jourdan, Leane, et al.
Pubblicazione: (2024)
di: Jourdan, Leane, et al.
Pubblicazione: (2024)
Text revision in Scientific Writing Assistance: An Overview
di: Jourdan, Léane, et al.
Pubblicazione: (2023)
di: Jourdan, Léane, et al.
Pubblicazione: (2023)
Multi-level SSL Feature Gating for Audio Deepfake Detection
di: Tran, Hoan My, et al.
Pubblicazione: (2025)
di: Tran, Hoan My, et al.
Pubblicazione: (2025)
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
di: Jourdan, Léane, et al.
Pubblicazione: (2025)
di: Jourdan, Léane, et al.
Pubblicazione: (2025)
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
di: Jourdan, Léane, et al.
Pubblicazione: (2026)
di: Jourdan, Léane, et al.
Pubblicazione: (2026)
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
di: Jourdan, Fanny
Pubblicazione: (2024)
di: Jourdan, Fanny
Pubblicazione: (2024)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
di: Jourdan, Fanny, et al.
Pubblicazione: (2025)
di: Jourdan, Fanny, et al.
Pubblicazione: (2025)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
di: Fu, Xiao, et al.
Pubblicazione: (2025)
di: Fu, Xiao, et al.
Pubblicazione: (2025)
What do the metrics mean? A critical analysis of the use of Automated Evaluation Metrics in Interpreting
di: Downie, Jonathan, et al.
Pubblicazione: (2026)
di: Downie, Jonathan, et al.
Pubblicazione: (2026)
TAPS: Tool-Augmented Personalisation via Structured Tagging
di: Taktasheva, Ekaterina, et al.
Pubblicazione: (2025)
di: Taktasheva, Ekaterina, et al.
Pubblicazione: (2025)
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
di: Dowerah, Sandipana, et al.
Pubblicazione: (2025)
di: Dowerah, Sandipana, et al.
Pubblicazione: (2025)
Tau-Eval: A Unified Evaluation Framework for Useful and Private Text Anonymization
di: Loiseau, Gabriel, et al.
Pubblicazione: (2025)
di: Loiseau, Gabriel, et al.
Pubblicazione: (2025)
Evaluating the Evaluators: Are readability metrics good measures of readability?
di: Cachola, Isabel, et al.
Pubblicazione: (2025)
di: Cachola, Isabel, et al.
Pubblicazione: (2025)
Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2026)
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2026)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
di: Ross, Candace, et al.
Pubblicazione: (2024)
di: Ross, Candace, et al.
Pubblicazione: (2024)
A Multilingual, Large-Scale Study of the Interplay between LLM Safeguards, Personalisation, and Disinformation
di: Leite, João A., et al.
Pubblicazione: (2025)
di: Leite, João A., et al.
Pubblicazione: (2025)
Waste Not, Want Not; Recycled Gumbel Noise Improves Consistency in Natural Language Generation
di: de Mijolla, Damien, et al.
Pubblicazione: (2025)
di: de Mijolla, Damien, et al.
Pubblicazione: (2025)
Towards Stable and Personalised Profiles for Lexical Alignment in Spoken Human-Agent Dialogue
di: Schaaij, Keara, et al.
Pubblicazione: (2025)
di: Schaaij, Keara, et al.
Pubblicazione: (2025)
Passive Learning of Lattice Automata from Recurrent Neural Networks
di: Slimi, Jaouhar, et al.
Pubblicazione: (2025)
di: Slimi, Jaouhar, et al.
Pubblicazione: (2025)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
di: Vázquez, Alain, et al.
Pubblicazione: (2026)
di: Vázquez, Alain, et al.
Pubblicazione: (2026)
Personalised Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code Generation
di: Chen, Hailin, et al.
Pubblicazione: (2023)
di: Chen, Hailin, et al.
Pubblicazione: (2023)
Structured Context Recomposition for Large Language Models Using Probabilistic Layer Realignment
di: Teel, Jonathan, et al.
Pubblicazione: (2025)
di: Teel, Jonathan, et al.
Pubblicazione: (2025)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
di: Chen, Lu, et al.
Pubblicazione: (2024)
di: Chen, Lu, et al.
Pubblicazione: (2024)
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
di: Piot, Paloma, et al.
Pubblicazione: (2025)
di: Piot, Paloma, et al.
Pubblicazione: (2025)
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
di: Sileo, Damien
Pubblicazione: (2025)
di: Sileo, Damien
Pubblicazione: (2025)
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
di: Sileo, Damien
Pubblicazione: (2024)
di: Sileo, Damien
Pubblicazione: (2024)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
di: Gajewska, Ewelina, et al.
Pubblicazione: (2026)
di: Gajewska, Ewelina, et al.
Pubblicazione: (2026)
MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts
di: Lanzeray, Etienne, et al.
Pubblicazione: (2026)
di: Lanzeray, Etienne, et al.
Pubblicazione: (2026)
LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
di: Gupta, Sammriddh, et al.
Pubblicazione: (2025)
di: Gupta, Sammriddh, et al.
Pubblicazione: (2025)
Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics
di: Bernas, Raphael, et al.
Pubblicazione: (2026)
di: Bernas, Raphael, et al.
Pubblicazione: (2026)
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop
di: Zhao, Runcong, et al.
Pubblicazione: (2025)
di: Zhao, Runcong, et al.
Pubblicazione: (2025)
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation
di: Goto, Takumi, et al.
Pubblicazione: (2025)
di: Goto, Takumi, et al.
Pubblicazione: (2025)
ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability
di: Poché, Antonin, et al.
Pubblicazione: (2025)
di: Poché, Antonin, et al.
Pubblicazione: (2025)
AI for Monitoring and Classifying Data Used in Research Literature
di: Macalaba, Rafael, et al.
Pubblicazione: (2026)
di: Macalaba, Rafael, et al.
Pubblicazione: (2026)
Detecting Turkish Synonyms Used in Different Time Periods
di: Yazar, Umur Togay, et al.
Pubblicazione: (2024)
di: Yazar, Umur Togay, et al.
Pubblicazione: (2024)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
di: Vincent, Sebastian, et al.
Pubblicazione: (2023)
di: Vincent, Sebastian, et al.
Pubblicazione: (2023)
Developing a Mixed-Methods Pipeline for Community-Oriented Digitization of Kwak'wala Legacy Texts
di: Agarwal, Milind, et al.
Pubblicazione: (2025)
di: Agarwal, Milind, et al.
Pubblicazione: (2025)
Ink and Individuality: Crafting a Personalised Narrative in the Age of LLMs
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2024)
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Identifying Reliable Evaluation Metrics for Scientific Text Revision
di: Jourdan, Léane, et al.
Pubblicazione: (2025) -
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
di: Jourdan, Leane, et al.
Pubblicazione: (2024) -
Text revision in Scientific Writing Assistance: An Overview
di: Jourdan, Léane, et al.
Pubblicazione: (2023) -
Multi-level SSL Feature Gating for Audio Deepfake Detection
di: Tran, Hoan My, et al.
Pubblicazione: (2025) -
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
di: Jourdan, Léane, et al.
Pubblicazione: (2025)