Efficient Continual Learning for Small Language Models with a Discrete Key-Value Bottleneck
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Diera, Andor, Galke, Lukas, Karl, Fabian, Scherp, Ansgar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Isotropy Matters: Soft-ZCA Whitening of Embeddings for Semantic Code Search
von: Diera, Andor, et al.
Veröffentlicht: (2024)
von: Diera, Andor, et al.
Veröffentlicht: (2024)
Do Language Models Encode Semantic Relations? Probing and Sparse Feature Analysis
von: Diera, Andor, et al.
Veröffentlicht: (2026)
von: Diera, Andor, et al.
Veröffentlicht: (2026)
Memorization of Named Entities in Fine-tuned BERT Models
von: Diera, Andor, et al.
Veröffentlicht: (2022)
von: Diera, Andor, et al.
Veröffentlicht: (2022)
Are We Really Making Much Progress in Text Classification? A Comparative Review
von: Galke, Lukas, et al.
Veröffentlicht: (2022)
von: Galke, Lukas, et al.
Veröffentlicht: (2022)
CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents
von: Karl, Fabian, et al.
Veröffentlicht: (2025)
von: Karl, Fabian, et al.
Veröffentlicht: (2025)
Multi-View Structural Graph Summaries
von: Frank, Jonatan, et al.
Veröffentlicht: (2024)
von: Frank, Jonatan, et al.
Veröffentlicht: (2024)
Semantic Source Code Segmentation using Small and Large Language Models
von: Dahou, Abdelhalim, et al.
Veröffentlicht: (2025)
von: Dahou, Abdelhalim, et al.
Veröffentlicht: (2025)
POWN: Prototypical Open-World Node Classification
von: Hoffmann, Marcel, et al.
Veröffentlicht: (2024)
von: Hoffmann, Marcel, et al.
Veröffentlicht: (2024)
Gumbel-MPNN: Graph Rewiring with Gumbel-Softmax
von: Hoffmann, Marcel, et al.
Veröffentlicht: (2025)
von: Hoffmann, Marcel, et al.
Veröffentlicht: (2025)
A Transformer-based Autoregressive Decoder Architecture for Hierarchical Text Classification
von: Yousef, Younes, et al.
Veröffentlicht: (2025)
von: Yousef, Younes, et al.
Veröffentlicht: (2025)
Your Extreme Multi-label Classifier is Secretly a Hierarchical Text Classifier for Free
von: Bertalis, Nerijus, et al.
Veröffentlicht: (2024)
von: Bertalis, Nerijus, et al.
Veröffentlicht: (2024)
Isolating Culture Neurons in Multilingual Large Language Models
von: Namazifard, Danial, et al.
Veröffentlicht: (2025)
von: Namazifard, Danial, et al.
Veröffentlicht: (2025)
Text Role Classification in Scientific Charts Using Multimodal Transformers
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
Learning and communication pressures in neural networks: Lessons from emergent communication
von: Galke, Lukas, et al.
Veröffentlicht: (2024)
von: Galke, Lukas, et al.
Veröffentlicht: (2024)
On the Anatomy of Real-World R Code for Static Analysis
von: Sihler, Florian, et al.
Veröffentlicht: (2024)
von: Sihler, Florian, et al.
Veröffentlicht: (2024)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
von: Torrielli, Federico, et al.
Veröffentlicht: (2026)
von: Torrielli, Federico, et al.
Veröffentlicht: (2026)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
von: Dang, Thao Anh, et al.
Veröffentlicht: (2024)
von: Dang, Thao Anh, et al.
Veröffentlicht: (2024)
Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences
von: Seidlmayer, Eva, et al.
Veröffentlicht: (2025)
von: Seidlmayer, Eva, et al.
Veröffentlicht: (2025)
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models
von: Lin, Zhenghao, et al.
Veröffentlicht: (2025)
von: Lin, Zhenghao, et al.
Veröffentlicht: (2025)
What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure
von: Galke, Lukas, et al.
Veröffentlicht: (2023)
von: Galke, Lukas, et al.
Veröffentlicht: (2023)
Training Language Models to Use Prolog as a Tool
von: Mellgren, Niklas, et al.
Veröffentlicht: (2025)
von: Mellgren, Niklas, et al.
Veröffentlicht: (2025)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
Chain of Summaries: Summarization Through Iterative Questioning
von: Brach, William, et al.
Veröffentlicht: (2025)
von: Brach, William, et al.
Veröffentlicht: (2025)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2022)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2022)
GLaMoR: Consistency Checking of OWL Ontologies using Graph Language Models
von: Mücke, Justin, et al.
Veröffentlicht: (2025)
von: Mücke, Justin, et al.
Veröffentlicht: (2025)
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects
von: Zhang, Jun, et al.
Veröffentlicht: (2026)
von: Zhang, Jun, et al.
Veröffentlicht: (2026)
Steering Information Utility in Key-Value Memory for Language Model Post-Training
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025)
Policy Learning with a Language Bottleneck
von: Srivastava, Megha, et al.
Veröffentlicht: (2024)
von: Srivastava, Megha, et al.
Veröffentlicht: (2024)
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
von: Beltoft, Stine Lyngsø, et al.
Veröffentlicht: (2026)
von: Beltoft, Stine Lyngsø, et al.
Veröffentlicht: (2026)
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
von: Barmina, Gianluca, et al.
Veröffentlicht: (2025)
von: Barmina, Gianluca, et al.
Veröffentlicht: (2025)
Concept Bottleneck Large Language Models
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
von: Wang, Minghan, et al.
Veröffentlicht: (2025)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
Small Languages, Big Models: A Study of Continual Training on Languages of Norway
von: Samuel, David, et al.
Veröffentlicht: (2024)
von: Samuel, David, et al.
Veröffentlicht: (2024)
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
Rehearsal-Free Modular and Compositional Continual Learning for Language Models
von: Wang, Mingyang, et al.
Veröffentlicht: (2024)
von: Wang, Mingyang, et al.
Veröffentlicht: (2024)
SDUs DAISY: A Benchmark for Danish Culture
von: Nielsen, Jacob, et al.
Veröffentlicht: (2026)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2026)
LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Isotropy Matters: Soft-ZCA Whitening of Embeddings for Semantic Code Search
von: Diera, Andor, et al.
Veröffentlicht: (2024) -
Do Language Models Encode Semantic Relations? Probing and Sparse Feature Analysis
von: Diera, Andor, et al.
Veröffentlicht: (2026) -
Memorization of Named Entities in Fine-tuned BERT Models
von: Diera, Andor, et al.
Veröffentlicht: (2022) -
Are We Really Making Much Progress in Text Classification? A Comparative Review
von: Galke, Lukas, et al.
Veröffentlicht: (2022) -
CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents
von: Karl, Fabian, et al.
Veröffentlicht: (2025)