Pretrained Generative Language Models as General Learning Frameworks for Sequence-Based Tasks
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Fauber, Ben |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accurate Prediction of Ligand-Protein Interaction Affinities with Fine-Tuned Small Language Models
von: Fauber, Ben
Veröffentlicht: (2024)
von: Fauber, Ben
Veröffentlicht: (2024)
Learning the Latent Rules of a Game from Data: A Chess Story
von: Fauber, Ben
Veröffentlicht: (2024)
von: Fauber, Ben
Veröffentlicht: (2024)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
Language and Experience: A Computational Model of Social Learning in Complex Tasks
von: Colas, Cédric, et al.
Veröffentlicht: (2025)
von: Colas, Cédric, et al.
Veröffentlicht: (2025)
Episodic Memories Generation and Evaluation Benchmark for Large Language Models
von: Huet, Alexis, et al.
Veröffentlicht: (2025)
von: Huet, Alexis, et al.
Veröffentlicht: (2025)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Pretraining Large Language Models with NVFP4
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
FastDiSS: Few-step Match Many-step Diffusion Language Model on Sequence-to-Sequence Generation--Full Version
von: Nguyen-Cong, Dat, et al.
Veröffentlicht: (2026)
von: Nguyen-Cong, Dat, et al.
Veröffentlicht: (2026)
SPADE: Faster Drug Discovery by Learning from Sparse Data
von: Nandakumar, Rahul, et al.
Veröffentlicht: (2026)
von: Nandakumar, Rahul, et al.
Veröffentlicht: (2026)
Patent Language Model Pretraining with ModernBERT
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks
von: Zhu, Yinghao, et al.
Veröffentlicht: (2024)
von: Zhu, Yinghao, et al.
Veröffentlicht: (2024)
FutureFill: Fast Generation from Convolutional Sequence Models
von: Agarwal, Naman, et al.
Veröffentlicht: (2024)
von: Agarwal, Naman, et al.
Veröffentlicht: (2024)
Sequence-to-Sequence Spanish Pre-trained Language Models
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
In-context Pretraining: Language Modeling Beyond Document Boundaries
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
von: Bayazit, Deniz, et al.
Veröffentlicht: (2023)
GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
von: Stepanov, Ihor, et al.
Veröffentlicht: (2025)
von: Stepanov, Ihor, et al.
Veröffentlicht: (2025)
Filter-then-Generate: Large Language Models with Structure-Text Adapter for Knowledge Graph Completion
von: Liu, Ben, et al.
Veröffentlicht: (2024)
von: Liu, Ben, et al.
Veröffentlicht: (2024)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
von: Yu, Zichun, et al.
Veröffentlicht: (2026)
von: Yu, Zichun, et al.
Veröffentlicht: (2026)
A Dual-Space Framework for General Knowledge Distillation of Large Language Models
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
The Mouth is Not the Brain: Bridging Energy-Based World Models and Language Generation
von: Niimi, Junichiro
Veröffentlicht: (2026)
von: Niimi, Junichiro
Veröffentlicht: (2026)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
von: Wu, Mian, et al.
Veröffentlicht: (2025)
von: Wu, Mian, et al.
Veröffentlicht: (2025)
RuAG: Learned-rule-augmented Generation for Large Language Models
von: Zhang, Yudi, et al.
Veröffentlicht: (2024)
von: Zhang, Yudi, et al.
Veröffentlicht: (2024)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks
von: Van Huynh, Tin, et al.
Veröffentlicht: (2026)
von: Van Huynh, Tin, et al.
Veröffentlicht: (2026)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
von: Neill, James O', et al.
Veröffentlicht: (2025)
von: Neill, James O', et al.
Veröffentlicht: (2025)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
von: Zhu, Alan, et al.
Veröffentlicht: (2025)
von: Zhu, Alan, et al.
Veröffentlicht: (2025)
Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy
von: Cullhed, Eric
Veröffentlicht: (2024)
von: Cullhed, Eric
Veröffentlicht: (2024)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
von: Ge, Ce, et al.
Veröffentlicht: (2024)
von: Ge, Ce, et al.
Veröffentlicht: (2024)
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
von: Xiaomi, LLM-Core, et al.
Veröffentlicht: (2025)
von: Xiaomi, LLM-Core, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Accurate Prediction of Ligand-Protein Interaction Affinities with Fine-Tuned Small Language Models
von: Fauber, Ben
Veröffentlicht: (2024) -
Learning the Latent Rules of a Game from Data: A Chess Story
von: Fauber, Ben
Veröffentlicht: (2024) -
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
von: Wang, Xinyi, et al.
Veröffentlicht: (2024) -
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
von: Li, Yixiao, et al.
Veröffentlicht: (2025) -
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)