Exploring Precision and Recall to assess the quality and diversity of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Bronnec, Florian Le, Verine, Alexandre, Negrevergne, Benjamin, Chevaleyre, Yann, Allauzen, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
by: Verine, Alexandre, et al.
Published: (2025)
by: Verine, Alexandre, et al.
Published: (2025)
Improving Discriminator Guidance in Diffusion Models
by: Verine, Alexandre, et al.
Published: (2025)
by: Verine, Alexandre, et al.
Published: (2025)
On the expressivity of bi-Lipschitz normalizing flows
by: Verine, Alexandre, et al.
Published: (2021)
by: Verine, Alexandre, et al.
Published: (2021)
Beyond pass@k: Redundancy-Aware RLVR for Multi-Sample Code Generation
by: Florian, Le Bronnec, et al.
Published: (2026)
by: Florian, Le Bronnec, et al.
Published: (2026)
Optimal Budgeted Rejection Sampling for Generative Models
by: Verine, Alexandre, et al.
Published: (2023)
by: Verine, Alexandre, et al.
Published: (2023)
Equalized Generative Treatment: Matching f-divergences for Fairness in Generative Models
by: Verine, Alexandre, et al.
Published: (2026)
by: Verine, Alexandre, et al.
Published: (2026)
Lattice Climber Attack: Adversarial attacks for randomized mixtures of classifiers
by: Gnecco-Heredia, Lucas, et al.
Published: (2025)
by: Gnecco-Heredia, Lucas, et al.
Published: (2025)
LOCOST: State-Space Models for Long Document Abstractive Summarization
by: Bronnec, Florian Le, et al.
Published: (2024)
by: Bronnec, Florian Le, et al.
Published: (2024)
Conditional Distribution Quantization in Machine Learning
by: Delattre, Blaise, et al.
Published: (2025)
by: Delattre, Blaise, et al.
Published: (2025)
Polynomial Mixing for Efficient Self-supervised Speech Encoders
by: Feillet, Eva, et al.
Published: (2026)
by: Feillet, Eva, et al.
Published: (2026)
SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text Generation
by: Duong, Song, et al.
Published: (2025)
by: Duong, Song, et al.
Published: (2025)
Chain and Causal Attention for Efficient Entity Tracking
by: Fagnou, Erwan, et al.
Published: (2024)
by: Fagnou, Erwan, et al.
Published: (2024)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
by: Zhao, Hangyue, et al.
Published: (2026)
by: Zhao, Hangyue, et al.
Published: (2026)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
by: Fagnou, Erwan, et al.
Published: (2026)
by: Fagnou, Erwan, et al.
Published: (2026)
Unveiling the Role of Randomization in Multiclass Adversarial Classification: Insights from Graph Theory
by: Gnecco-Heredia, Lucas, et al.
Published: (2025)
by: Gnecco-Heredia, Lucas, et al.
Published: (2025)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024)
by: Chughtai, Bilal, et al.
Published: (2024)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
by: Pink, Mathis, et al.
Published: (2024)
by: Pink, Mathis, et al.
Published: (2024)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)
by: Skean, Oscar, et al.
Published: (2024)
LLMs can learn self-restraint through iterative self-reflection
by: Piché, Alexandre, et al.
Published: (2024)
by: Piché, Alexandre, et al.
Published: (2024)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
by: Delattre, Blaise, et al.
Published: (2024)
by: Delattre, Blaise, et al.
Published: (2024)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
by: Zheng, Kunhao, et al.
Published: (2026)
by: Zheng, Kunhao, et al.
Published: (2026)
LLM In-Context Recall is Prompt Dependent
by: Machlab, Daniel, et al.
Published: (2024)
by: Machlab, Daniel, et al.
Published: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
by: Smith, Matthew L., et al.
Published: (2026)
by: Smith, Matthew L., et al.
Published: (2026)
Spectral Collapse in Diffusion Inversion
by: Bourriez, Nicolas, et al.
Published: (2026)
by: Bourriez, Nicolas, et al.
Published: (2026)
Learning To Sample From Diffusion Models Via Inverse Reinforcement Learning
by: Bourdrez, Constant, et al.
Published: (2026)
by: Bourdrez, Constant, et al.
Published: (2026)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Exploring Contrastive Learning for Long-Tailed Multi-Label Text Classification
by: Audibert, Alexandre, et al.
Published: (2024)
by: Audibert, Alexandre, et al.
Published: (2024)
Learn or Recall? Revisiting Incremental Learning with Pre-trained Language Models
by: Zheng, Junhao, et al.
Published: (2023)
by: Zheng, Junhao, et al.
Published: (2023)
Hidden State Poisoning Attacks against Mamba-based Language Models
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
The Lipschitz-Variance-Margin Tradeoff for Enhanced Randomized Smoothing
by: Delattre, Blaise, et al.
Published: (2023)
by: Delattre, Blaise, et al.
Published: (2023)
Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration
by: Guan, Zhong, et al.
Published: (2025)
by: Guan, Zhong, et al.
Published: (2025)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
by: Gao, Peiheng, et al.
Published: (2025)
by: Gao, Peiheng, et al.
Published: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
by: Deng, Jianing, et al.
Published: (2026)
by: Deng, Jianing, et al.
Published: (2026)
Scaling Laws for Precision
by: Kumar, Tanishq, et al.
Published: (2024)
by: Kumar, Tanishq, et al.
Published: (2024)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Similar Items
-
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
by: Verine, Alexandre, et al.
Published: (2025) -
Improving Discriminator Guidance in Diffusion Models
by: Verine, Alexandre, et al.
Published: (2025) -
On the expressivity of bi-Lipschitz normalizing flows
by: Verine, Alexandre, et al.
Published: (2021) -
Beyond pass@k: Redundancy-Aware RLVR for Multi-Sample Code Generation
by: Florian, Le Bronnec, et al.
Published: (2026) -
Optimal Budgeted Rejection Sampling for Generative Models
by: Verine, Alexandre, et al.
Published: (2023)