PIE: Performance Interval Estimation for Free-Form Generation Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Hsu, Chi-Yang, Braylan, Alexander, Su, Yiheng, Lease, Matthew, Alonso, Omar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM REgression with a Latent Iterative State Head
por: Su, Yiheng, et al.
Publicado: (2026)
por: Su, Yiheng, et al.
Publicado: (2026)
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
por: Su, Yiheng, et al.
Publicado: (2023)
por: Su, Yiheng, et al.
Publicado: (2023)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
por: Wu, Mian, et al.
Publicado: (2025)
por: Wu, Mian, et al.
Publicado: (2025)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
por: Li, Zongxia, et al.
Publicado: (2025)
por: Li, Zongxia, et al.
Publicado: (2025)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
por: Si, Chongjie, et al.
Publicado: (2025)
por: Si, Chongjie, et al.
Publicado: (2025)
TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
por: Naim, Omar, et al.
Publicado: (2025)
por: Naim, Omar, et al.
Publicado: (2025)
VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs
por: Amballa, Avinash, et al.
Publicado: (2025)
por: Amballa, Avinash, et al.
Publicado: (2025)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
por: Wang, Zhiyuan, et al.
Publicado: (2024)
por: Wang, Zhiyuan, et al.
Publicado: (2024)
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
por: Wang, Shun, et al.
Publicado: (2025)
por: Wang, Shun, et al.
Publicado: (2025)
From Symbolic Tasks to Code Generation: Diversification Yields Better Task Performers
por: Zhang, Dylan, et al.
Publicado: (2024)
por: Zhang, Dylan, et al.
Publicado: (2024)
TabDLM: Free-Form Tabular Data Generation via Joint Numerical-Language Diffusion
por: Cai, Donghong, et al.
Publicado: (2026)
por: Cai, Donghong, et al.
Publicado: (2026)
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
por: Chen, Yan-Lun, et al.
Publicado: (2025)
por: Chen, Yan-Lun, et al.
Publicado: (2025)
Diversity Enhances an LLM's Performance in RAG and Long-context Task
por: Wang, Zhichao, et al.
Publicado: (2025)
por: Wang, Zhichao, et al.
Publicado: (2025)
Scaling Laws for Downstream Task Performance of Large Language Models
por: Isik, Berivan, et al.
Publicado: (2024)
por: Isik, Berivan, et al.
Publicado: (2024)
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
por: Tian, Yufei, et al.
Publicado: (2024)
por: Tian, Yufei, et al.
Publicado: (2024)
Reference-Free Rating of LLM Responses via Latent Information
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
por: He, Jianfeng, et al.
Publicado: (2024)
por: He, Jianfeng, et al.
Publicado: (2024)
Optimizing Multi-Task Learning for Enhanced Performance in Large Language Models
por: Qi, Zhen, et al.
Publicado: (2024)
por: Qi, Zhen, et al.
Publicado: (2024)
How Does Code Pretraining Affect Language Model Task Performance?
por: Petty, Jackson, et al.
Publicado: (2024)
por: Petty, Jackson, et al.
Publicado: (2024)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
por: Yan, Hao, et al.
Publicado: (2026)
por: Yan, Hao, et al.
Publicado: (2026)
Towards Automated Kernel Generation in the Era of LLMs
por: Yu, Yang, et al.
Publicado: (2026)
por: Yu, Yang, et al.
Publicado: (2026)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
por: Niu, Tianyi, et al.
Publicado: (2026)
por: Niu, Tianyi, et al.
Publicado: (2026)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
por: Bhowmik, Shimanto, et al.
Publicado: (2025)
por: Bhowmik, Shimanto, et al.
Publicado: (2025)
Exploring the Impact of a Transformer's Latent Space Geometry on Downstream Task Performance
por: Marbut, Anna C., et al.
Publicado: (2024)
por: Marbut, Anna C., et al.
Publicado: (2024)
Linguistic Calibration of Long-Form Generations
por: Band, Neil, et al.
Publicado: (2024)
por: Band, Neil, et al.
Publicado: (2024)
Towards Understanding Multi-Task Learning (Generalization) of LLMs via Detecting and Exploring Task-Specific Neurons
por: Leng, Yongqi, et al.
Publicado: (2024)
por: Leng, Yongqi, et al.
Publicado: (2024)
An Automatic Prompt Generation System for Tabular Data Tasks
por: Akella, Ashlesha, et al.
Publicado: (2024)
por: Akella, Ashlesha, et al.
Publicado: (2024)
ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events
por: Islakoglu, Duygu Sezen, et al.
Publicado: (2025)
por: Islakoglu, Duygu Sezen, et al.
Publicado: (2025)
Robustness as an Emergent Property of Task Performance
por: Ashury-Tahan, Shir, et al.
Publicado: (2026)
por: Ashury-Tahan, Shir, et al.
Publicado: (2026)
Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks
por: Li, Zongqian, et al.
Publicado: (2026)
por: Li, Zongqian, et al.
Publicado: (2026)
Converting Transformers into DGNNs Form
por: Zhang, Jie, et al.
Publicado: (2025)
por: Zhang, Jie, et al.
Publicado: (2025)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
por: Dialameh, Maryam, et al.
Publicado: (2025)
por: Dialameh, Maryam, et al.
Publicado: (2025)
Detecting LLM-Generated Text with Performance Guarantees
por: Zhou, Hongyi, et al.
Publicado: (2026)
por: Zhou, Hongyi, et al.
Publicado: (2026)
Automatic Generation of Python Programs Using Context-Free Grammars
por: Yamani, Kamel, et al.
Publicado: (2024)
por: Yamani, Kamel, et al.
Publicado: (2024)
Gumbel Distillation for Parallel Text Generation
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
Generative AI Meets Semantic Communication: Evolution and Revolution of Communication Tasks
por: Grassucci, Eleonora, et al.
Publicado: (2024)
por: Grassucci, Eleonora, et al.
Publicado: (2024)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
por: Nayak, Nihal V., et al.
Publicado: (2024)
por: Nayak, Nihal V., et al.
Publicado: (2024)
Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
por: Sabbaghi, Mahdi, et al.
Publicado: (2024)
por: Sabbaghi, Mahdi, et al.
Publicado: (2024)
Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
por: Cox, Kyle, et al.
Publicado: (2025)
por: Cox, Kyle, et al.
Publicado: (2025)
Memory Is All You Need: Testing How Model Memory Affects LLM Performance in Annotation Tasks
por: Timoneda, Joan C., et al.
Publicado: (2025)
por: Timoneda, Joan C., et al.
Publicado: (2025)
Ejemplares similares
-
LLM REgression with a Latent Iterative State Head
por: Su, Yiheng, et al.
Publicado: (2026) -
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
por: Su, Yiheng, et al.
Publicado: (2023) -
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
por: Wu, Mian, et al.
Publicado: (2025) -
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
por: Li, Zongxia, et al.
Publicado: (2025) -
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
por: Si, Chongjie, et al.
Publicado: (2025)