The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Carlsson, Fredrik, Liu, Fangyu, Ward, Daniel, Kurfali, Murathan, Nivre, Joakim |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
di: Kurfalı, Murathan, et al.
Pubblicazione: (2025)
di: Kurfalı, Murathan, et al.
Pubblicazione: (2025)
Language Bias under Conflicting Information in Multilingual LLMs
di: Östling, Robert, et al.
Pubblicazione: (2026)
di: Östling, Robert, et al.
Pubblicazione: (2026)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
di: Gogoulou, Evangelia, et al.
Pubblicazione: (2025)
di: Gogoulou, Evangelia, et al.
Pubblicazione: (2025)
Generating Planning Feedback for Open-Ended Programming Exercises with LLMs
di: Demirtaş, Mehmet Arif, et al.
Pubblicazione: (2025)
di: Demirtaş, Mehmet Arif, et al.
Pubblicazione: (2025)
Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
di: Li, Meimingwei, et al.
Pubblicazione: (2026)
di: Li, Meimingwei, et al.
Pubblicazione: (2026)
Reverse-Engineered Reasoning for Open-Ended Generation
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
di: Xie, Tianbao, et al.
Pubblicazione: (2024)
Open-Ended Wargames with Large Language Models
di: Hogan, Daniel P., et al.
Pubblicazione: (2024)
di: Hogan, Daniel P., et al.
Pubblicazione: (2024)
Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation
di: Zhou, Yuxuan, et al.
Pubblicazione: (2024)
di: Zhou, Yuxuan, et al.
Pubblicazione: (2024)
Learning Approximate and Exact Numeral Systems via Reinforcement Learning
di: Carlsson, Emil, et al.
Pubblicazione: (2021)
di: Carlsson, Emil, et al.
Pubblicazione: (2021)
Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations
di: Yu, Sunny, et al.
Pubblicazione: (2025)
di: Yu, Sunny, et al.
Pubblicazione: (2025)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
di: Jeune, Pierre Le, et al.
Pubblicazione: (2026)
di: Jeune, Pierre Le, et al.
Pubblicazione: (2026)
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
di: Wang, Pengkai, et al.
Pubblicazione: (2025)
di: Wang, Pengkai, et al.
Pubblicazione: (2025)
A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds
di: Cui, Christopher Z., et al.
Pubblicazione: (2024)
di: Cui, Christopher Z., et al.
Pubblicazione: (2024)
Lightweight Connective Detection Using Gradient Boosting
di: Er, Mustafa Erolcan, et al.
Pubblicazione: (2024)
di: Er, Mustafa Erolcan, et al.
Pubblicazione: (2024)
MIRROR: A Novel Approach for the Automated Evaluation of Open-Ended Question Generation
di: Deroy, Aniket, et al.
Pubblicazione: (2024)
di: Deroy, Aniket, et al.
Pubblicazione: (2024)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
di: Demchak, Nathaniel, et al.
Pubblicazione: (2024)
di: Demchak, Nathaniel, et al.
Pubblicazione: (2024)
ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generation
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2026)
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2026)
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
di: Guo, Zhihan, et al.
Pubblicazione: (2025)
di: Guo, Zhihan, et al.
Pubblicazione: (2025)
GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
di: Hutson, Dylan, et al.
Pubblicazione: (2025)
di: Hutson, Dylan, et al.
Pubblicazione: (2025)
Towards Open-Ended Discovery for Low-Resource NLP
di: Dossou, Bonaventure F. P., et al.
Pubblicazione: (2025)
di: Dossou, Bonaventure F. P., et al.
Pubblicazione: (2025)
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
di: Wang, Jize, et al.
Pubblicazione: (2026)
di: Wang, Jize, et al.
Pubblicazione: (2026)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
di: Samvelyan, Mikayel, et al.
Pubblicazione: (2024)
di: Samvelyan, Mikayel, et al.
Pubblicazione: (2024)
AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
di: Li, Yishan, et al.
Pubblicazione: (2026)
di: Li, Yishan, et al.
Pubblicazione: (2026)
An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems
di: Shaik, Hashmath, et al.
Pubblicazione: (2024)
di: Shaik, Hashmath, et al.
Pubblicazione: (2024)
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
di: Liu, Wanlong, et al.
Pubblicazione: (2026)
di: Liu, Wanlong, et al.
Pubblicazione: (2026)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
di: Yang, Zixuan, et al.
Pubblicazione: (2026)
di: Yang, Zixuan, et al.
Pubblicazione: (2026)
VUB-HYDR/Wikimpacts: Wikimpacts 1.0 database
di: Shorouq, et al.
Pubblicazione: (2025)
di: Shorouq, et al.
Pubblicazione: (2025)
O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
di: Mei, Jianbiao, et al.
Pubblicazione: (2025)
di: Mei, Jianbiao, et al.
Pubblicazione: (2025)
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
di: Lu, Xiaotian, et al.
Pubblicazione: (2024)
di: Lu, Xiaotian, et al.
Pubblicazione: (2024)
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
di: An, Subin, et al.
Pubblicazione: (2025)
di: An, Subin, et al.
Pubblicazione: (2025)
Decoding Open-Ended Information Seeking Goals from Eye Movements in Reading
di: Hadar, Cfir Avraham, et al.
Pubblicazione: (2025)
di: Hadar, Cfir Avraham, et al.
Pubblicazione: (2025)
O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL
di: Yao, Yi, et al.
Pubblicazione: (2026)
di: Yao, Yi, et al.
Pubblicazione: (2026)
Dreaming in Code for Curriculum Learning in Open-Ended Worlds
di: Mitsides, Konstantinos, et al.
Pubblicazione: (2026)
di: Mitsides, Konstantinos, et al.
Pubblicazione: (2026)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
di: Iyer, Vivek, et al.
Pubblicazione: (2025)
di: Iyer, Vivek, et al.
Pubblicazione: (2025)
Basis Vector Metric: A Method for Robust Open-Ended State Change Detection
di: Oprea, David, et al.
Pubblicazione: (2025)
di: Oprea, David, et al.
Pubblicazione: (2025)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
di: Keluskar, Aryan, et al.
Pubblicazione: (2024)
di: Keluskar, Aryan, et al.
Pubblicazione: (2024)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
di: Shin, Jisu, et al.
Pubblicazione: (2025)
di: Shin, Jisu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
di: Kurfalı, Murathan, et al.
Pubblicazione: (2025) -
Language Bias under Conflicting Information in Multilingual LLMs
di: Östling, Robert, et al.
Pubblicazione: (2026) -
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
di: Gogoulou, Evangelia, et al.
Pubblicazione: (2025) -
Generating Planning Feedback for Open-Ended Programming Exercises with LLMs
di: Demirtaş, Mehmet Arif, et al.
Pubblicazione: (2025) -
Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
di: Li, Meimingwei, et al.
Pubblicazione: (2026)