CurLL: A Developmental Framework to Evaluate Continual Learning in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kalyan, Pavan, Mishra, Shubhra, Lokam, Satya, Goyal, Navin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
by: Mishra, Shubhra, et al.
Published: (2024)
by: Mishra, Shubhra, et al.
Published: (2024)
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024)
by: Gulati, Aryan, et al.
Published: (2024)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
Developmental Predictive Coding Model for Early Infancy Mono and Bilingual Vocal Continual Learning
by: Chen, Xiaodan, et al.
Published: (2024)
by: Chen, Xiaodan, et al.
Published: (2024)
SCOPE: Language Models as One-Time Teacher for Hierarchical Planning in Text Environments
by: Lu, Haoye, et al.
Published: (2025)
by: Lu, Haoye, et al.
Published: (2025)
DeepLL: Considering Linear Logic for the Analysis of Deep Learning Experiments
by: Papoulias, Nick
Published: (2024)
by: Papoulias, Nick
Published: (2024)
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
by: Shu, Yiheng, et al.
Published: (2026)
by: Shu, Yiheng, et al.
Published: (2026)
Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations
by: Cherukuri, Kalyan, et al.
Published: (2026)
by: Cherukuri, Kalyan, et al.
Published: (2026)
Learning Tractable Distributions Of Language Model Continuations
by: Yidou-Weng, Gwen, et al.
Published: (2025)
by: Yidou-Weng, Gwen, et al.
Published: (2025)
IDALC: A Semi-Supervised Framework for Intent Detection and Active Learning based Correction
by: Mullick, Ankan, et al.
Published: (2025)
by: Mullick, Ankan, et al.
Published: (2025)
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
Cost-of-Pass: An Economic Framework for Evaluating Language Models
by: Erol, Mehmet Hamza, et al.
Published: (2025)
by: Erol, Mehmet Hamza, et al.
Published: (2025)
The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models
by: Jamshidi, Saeid, et al.
Published: (2025)
by: Jamshidi, Saeid, et al.
Published: (2025)
Elsevier Arena: Human Evaluation of Chemistry/Biology/Health Foundational Large Language Models
by: Thorne, Camilo, et al.
Published: (2024)
by: Thorne, Camilo, et al.
Published: (2024)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Where Should I Study? Biased Language Models Decide! Evaluating Fairness in LMs for Academic Recommendations
by: Shailya, Krithi, et al.
Published: (2025)
by: Shailya, Krithi, et al.
Published: (2025)
A Scalable Tool for Measuring Manner and Result Verbs in Developmental Language Research
by: Singh, Divyesh Pratap, et al.
Published: (2026)
by: Singh, Divyesh Pratap, et al.
Published: (2026)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025)
by: Ghosh, Shubhra, et al.
Published: (2025)
ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models
by: Narsupalli, Yaswanth, et al.
Published: (2024)
by: Narsupalli, Yaswanth, et al.
Published: (2024)
LLMs as On-demand Customizable Service
by: Sarkar, Souvika, et al.
Published: (2024)
by: Sarkar, Souvika, et al.
Published: (2024)
Convergence of Outputs When Two Large Language Models Interact in a Multi-Agentic Setup
by: Maiti, Aniruddha, et al.
Published: (2025)
by: Maiti, Aniruddha, et al.
Published: (2025)
CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
by: Mishra, Manit, et al.
Published: (2024)
by: Mishra, Manit, et al.
Published: (2024)
Continual Learning Using Only Large Language Model Prompting
by: Qiu, Jiabao, et al.
Published: (2024)
by: Qiu, Jiabao, et al.
Published: (2024)
Continual Learning in Large Language Models: Methods, Challenges, and Opportunities
by: Chen, Hongyang, et al.
Published: (2026)
by: Chen, Hongyang, et al.
Published: (2026)
KIF: Knowledge Identification and Fusion for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2024)
by: Feng, Yujie, et al.
Published: (2024)
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
Towards Automatic Continual Learning: A Self-Adaptive Framework for Continual Instruction Tuning
by: Lin, Peiyi, et al.
Published: (2025)
by: Lin, Peiyi, et al.
Published: (2025)
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
by: Jones, Jaylen, et al.
Published: (2024)
by: Jones, Jaylen, et al.
Published: (2024)
IndicEval: A Bilingual Indian Educational Evaluation Framework for Large Language Models
by: Bharti, Saurabh, et al.
Published: (2026)
by: Bharti, Saurabh, et al.
Published: (2026)
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
by: Pimentel, Marco AF, et al.
Published: (2024)
by: Pimentel, Marco AF, et al.
Published: (2024)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
by: Olson, Matthew Lyle, et al.
Published: (2026)
by: Olson, Matthew Lyle, et al.
Published: (2026)
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)
by: Parmar, Mihir, et al.
Published: (2025)
A Quantum Inspired Variational Kernel and Explainable AI Framework for Cross Region Solar and Wind Energy Forecasting
by: Manjunath, Pavan, et al.
Published: (2026)
by: Manjunath, Pavan, et al.
Published: (2026)
CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models
by: Li, Dongfang, et al.
Published: (2024)
by: Li, Dongfang, et al.
Published: (2024)
A Scalable Framework for Evaluating Health Language Models
by: Mallinar, Neil, et al.
Published: (2025)
by: Mallinar, Neil, et al.
Published: (2025)
Similar Items
-
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
by: Aggarwal, Divyanshu, et al.
Published: (2024) -
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
by: Mishra, Shubhra, et al.
Published: (2024) -
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024) -
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024) -
Developmental Predictive Coding Model for Early Infancy Mono and Bilingual Vocal Continual Learning
by: Chen, Xiaodan, et al.
Published: (2024)