Exploring Precision and Recall to assess the quality and diversity of LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Bronnec, Florian Le, Verine, Alexandre, Negrevergne, Benjamin, Chevaleyre, Yann, Allauzen, Alexandre |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
par: Verine, Alexandre, et autres
Publié: (2025)
par: Verine, Alexandre, et autres
Publié: (2025)
Improving Discriminator Guidance in Diffusion Models
par: Verine, Alexandre, et autres
Publié: (2025)
par: Verine, Alexandre, et autres
Publié: (2025)
On the expressivity of bi-Lipschitz normalizing flows
par: Verine, Alexandre, et autres
Publié: (2021)
par: Verine, Alexandre, et autres
Publié: (2021)
Beyond pass@k: Redundancy-Aware RLVR for Multi-Sample Code Generation
par: Florian, Le Bronnec, et autres
Publié: (2026)
par: Florian, Le Bronnec, et autres
Publié: (2026)
Optimal Budgeted Rejection Sampling for Generative Models
par: Verine, Alexandre, et autres
Publié: (2023)
par: Verine, Alexandre, et autres
Publié: (2023)
Equalized Generative Treatment: Matching f-divergences for Fairness in Generative Models
par: Verine, Alexandre, et autres
Publié: (2026)
par: Verine, Alexandre, et autres
Publié: (2026)
Lattice Climber Attack: Adversarial attacks for randomized mixtures of classifiers
par: Gnecco-Heredia, Lucas, et autres
Publié: (2025)
par: Gnecco-Heredia, Lucas, et autres
Publié: (2025)
LOCOST: State-Space Models for Long Document Abstractive Summarization
par: Bronnec, Florian Le, et autres
Publié: (2024)
par: Bronnec, Florian Le, et autres
Publié: (2024)
Conditional Distribution Quantization in Machine Learning
par: Delattre, Blaise, et autres
Publié: (2025)
par: Delattre, Blaise, et autres
Publié: (2025)
Polynomial Mixing for Efficient Self-supervised Speech Encoders
par: Feillet, Eva, et autres
Publié: (2026)
par: Feillet, Eva, et autres
Publié: (2026)
SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text Generation
par: Duong, Song, et autres
Publié: (2025)
par: Duong, Song, et autres
Publié: (2025)
Chain and Causal Attention for Efficient Entity Tracking
par: Fagnou, Erwan, et autres
Publié: (2024)
par: Fagnou, Erwan, et autres
Publié: (2024)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
par: Zhao, Hangyue, et autres
Publié: (2026)
par: Zhao, Hangyue, et autres
Publié: (2026)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
par: Fagnou, Erwan, et autres
Publié: (2026)
par: Fagnou, Erwan, et autres
Publié: (2026)
Unveiling the Role of Randomization in Multiclass Adversarial Classification: Insights from Graph Theory
par: Gnecco-Heredia, Lucas, et autres
Publié: (2025)
par: Gnecco-Heredia, Lucas, et autres
Publié: (2025)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
par: Chughtai, Bilal, et autres
Publié: (2024)
par: Chughtai, Bilal, et autres
Publié: (2024)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
par: Lei, Ge, et autres
Publié: (2025)
par: Lei, Ge, et autres
Publié: (2025)
Reasoning Boosts Opinion Alignment in LLMs
par: Berdoz, Frédéric, et autres
Publié: (2026)
par: Berdoz, Frédéric, et autres
Publié: (2026)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
par: Pink, Mathis, et autres
Publié: (2024)
par: Pink, Mathis, et autres
Publié: (2024)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
par: Yuan, Jiaqing, et autres
Publié: (2024)
par: Yuan, Jiaqing, et autres
Publié: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
par: Skean, Oscar, et autres
Publié: (2024)
par: Skean, Oscar, et autres
Publié: (2024)
LLMs can learn self-restraint through iterative self-reflection
par: Piché, Alexandre, et autres
Publié: (2024)
par: Piché, Alexandre, et autres
Publié: (2024)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
par: Delattre, Blaise, et autres
Publié: (2024)
par: Delattre, Blaise, et autres
Publié: (2024)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
par: Zheng, Kunhao, et autres
Publié: (2026)
par: Zheng, Kunhao, et autres
Publié: (2026)
LLM In-Context Recall is Prompt Dependent
par: Machlab, Daniel, et autres
Publié: (2024)
par: Machlab, Daniel, et autres
Publié: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
par: Smith, Matthew L., et autres
Publié: (2026)
par: Smith, Matthew L., et autres
Publié: (2026)
Spectral Collapse in Diffusion Inversion
par: Bourriez, Nicolas, et autres
Publié: (2026)
par: Bourriez, Nicolas, et autres
Publié: (2026)
Learning To Sample From Diffusion Models Via Inverse Reinforcement Learning
par: Bourdrez, Constant, et autres
Publié: (2026)
par: Bourdrez, Constant, et autres
Publié: (2026)
Mimetic Initialization Helps State Space Models Learn to Recall
par: Trockman, Asher, et autres
Publié: (2024)
par: Trockman, Asher, et autres
Publié: (2024)
Exploring Contrastive Learning for Long-Tailed Multi-Label Text Classification
par: Audibert, Alexandre, et autres
Publié: (2024)
par: Audibert, Alexandre, et autres
Publié: (2024)
Learn or Recall? Revisiting Incremental Learning with Pre-trained Language Models
par: Zheng, Junhao, et autres
Publié: (2023)
par: Zheng, Junhao, et autres
Publié: (2023)
Hidden State Poisoning Attacks against Mamba-based Language Models
par: Mercier, Alexandre Le, et autres
Publié: (2026)
par: Mercier, Alexandre Le, et autres
Publié: (2026)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
par: Wang, Qianli, et autres
Publié: (2025)
par: Wang, Qianli, et autres
Publié: (2025)
The Lipschitz-Variance-Margin Tradeoff for Enhanced Randomized Smoothing
par: Delattre, Blaise, et autres
Publié: (2023)
par: Delattre, Blaise, et autres
Publié: (2023)
Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration
par: Guan, Zhong, et autres
Publié: (2025)
par: Guan, Zhong, et autres
Publié: (2025)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
par: Gao, Peiheng, et autres
Publié: (2025)
par: Gao, Peiheng, et autres
Publié: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
par: Deng, Jianing, et autres
Publié: (2026)
par: Deng, Jianing, et autres
Publié: (2026)
Scaling Laws for Precision
par: Kumar, Tanishq, et autres
Publié: (2024)
par: Kumar, Tanishq, et autres
Publié: (2024)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
par: Ghazanfari, Sara, et autres
Publié: (2024)
par: Ghazanfari, Sara, et autres
Publié: (2024)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
par: Vasudeva, Bhavya, et autres
Publié: (2026)
par: Vasudeva, Bhavya, et autres
Publié: (2026)
Documents similaires
-
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
par: Verine, Alexandre, et autres
Publié: (2025) -
Improving Discriminator Guidance in Diffusion Models
par: Verine, Alexandre, et autres
Publié: (2025) -
On the expressivity of bi-Lipschitz normalizing flows
par: Verine, Alexandre, et autres
Publié: (2021) -
Beyond pass@k: Redundancy-Aware RLVR for Multi-Sample Code Generation
par: Florian, Le Bronnec, et autres
Publié: (2026) -
Optimal Budgeted Rejection Sampling for Generative Models
par: Verine, Alexandre, et autres
Publié: (2023)