Pretraining with hierarchical memories: separating long-tail and common knowledge
Fuente:
arXiv
Guardado en:
| Autores principales: | Pouransari, Hadi, Grangier, David, Thomas, C, Kirchhof, Michael, Tuzel, Oncel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MobileCLIP2: Improving Multi-Modal Reinforced Training
por: Faghri, Fartash, et al.
Publicado: (2025)
por: Faghri, Fartash, et al.
Publicado: (2025)
Learning to Reason for Hallucination Span Detection
por: Su, Hsuan, et al.
Publicado: (2025)
por: Su, Hsuan, et al.
Publicado: (2025)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
por: Lu, Yen-Ju, et al.
Publicado: (2025)
por: Lu, Yen-Ju, et al.
Publicado: (2025)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2023)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2023)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
por: Li, Jeffrey, et al.
Publicado: (2025)
por: Li, Jeffrey, et al.
Publicado: (2025)
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
por: Kirchhof, Michael, et al.
Publicado: (2025)
por: Kirchhof, Michael, et al.
Publicado: (2025)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
por: Öncel, Fırat, et al.
Publicado: (2024)
por: Öncel, Fırat, et al.
Publicado: (2024)
TiC-CLIP: Continual Training of CLIP Models
por: Garg, Saurabh, et al.
Publicado: (2023)
por: Garg, Saurabh, et al.
Publicado: (2023)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
por: Pouransari, Hadi, et al.
Publicado: (2024)
por: Pouransari, Hadi, et al.
Publicado: (2024)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
por: Mehta, Sachin, et al.
Publicado: (2024)
por: Mehta, Sachin, et al.
Publicado: (2024)
FastVLM: Efficient Vision Encoding for Vision Language Models
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
The Geometries of Truth Are Orthogonal Across Tasks
por: Azizian, Waiss, et al.
Publicado: (2025)
por: Azizian, Waiss, et al.
Publicado: (2025)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
por: Hsieh, Cheng-Yu, et al.
Publicado: (2025)
por: Hsieh, Cheng-Yu, et al.
Publicado: (2025)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
por: Santilli, Andrea, et al.
Publicado: (2025)
por: Santilli, Andrea, et al.
Publicado: (2025)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
por: Kirchhof, Michael, et al.
Publicado: (2025)
por: Kirchhof, Michael, et al.
Publicado: (2025)
Uncertainties of Latent Representations in Computer Vision
por: Kirchhof, Michael
Publicado: (2024)
por: Kirchhof, Michael
Publicado: (2024)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
por: Choudhury, Deepro, et al.
Publicado: (2025)
por: Choudhury, Deepro, et al.
Publicado: (2025)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
por: Ablin, Pierre, et al.
Publicado: (2025)
por: Ablin, Pierre, et al.
Publicado: (2025)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
por: Li, Jeffrey, et al.
Publicado: (2026)
por: Li, Jeffrey, et al.
Publicado: (2026)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
por: Grangier, David, et al.
Publicado: (2024)
por: Grangier, David, et al.
Publicado: (2024)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
por: Singh, Karan, et al.
Publicado: (2026)
por: Singh, Karan, et al.
Publicado: (2026)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
por: Vemulapalli, Raviteja, et al.
Publicado: (2023)
por: Vemulapalli, Raviteja, et al.
Publicado: (2023)
Pretrained Hybrids with MAD Skills
por: Roberts, Nicholas, et al.
Publicado: (2024)
por: Roberts, Nicholas, et al.
Publicado: (2024)
Pretraining Large Language Models with NVFP4
por: NVIDIA, et al.
Publicado: (2025)
por: NVIDIA, et al.
Publicado: (2025)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
por: Tice, Cameron, et al.
Publicado: (2026)
por: Tice, Cameron, et al.
Publicado: (2026)
TPTT: Transforming Pretrained Transformers into Titans
por: Furfaro, Fabien
Publicado: (2025)
por: Furfaro, Fabien
Publicado: (2025)
RLP: Reinforcement as a Pretraining Objective
por: Hatamizadeh, Ali, et al.
Publicado: (2025)
por: Hatamizadeh, Ali, et al.
Publicado: (2025)
Memorization Dynamics of Fill-in-the-Middle Pretraining
por: von Arx, Tobias, et al.
Publicado: (2026)
por: von Arx, Tobias, et al.
Publicado: (2026)
Linguistic Blind Spots of Large Language Models
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
Tool Unlearning for Tool-Augmented LLMs
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
FairFlow: Mitigating Dataset Biases through Undecided Learning
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
Fresh in memory: Training-order recency is linearly encoded in language model activations
por: Krasheninnikov, Dmitrii, et al.
Publicado: (2025)
por: Krasheninnikov, Dmitrii, et al.
Publicado: (2025)
Patent Language Model Pretraining with ModernBERT
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
Output Embedding Centering for Stable LLM Pretraining
por: Stollenwerk, Felix, et al.
Publicado: (2026)
por: Stollenwerk, Felix, et al.
Publicado: (2026)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
por: Ali, Mehdi, et al.
Publicado: (2025)
por: Ali, Mehdi, et al.
Publicado: (2025)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
por: Mirzadeh, Iman, et al.
Publicado: (2024)
por: Mirzadeh, Iman, et al.
Publicado: (2024)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
por: Ni, Kangqi, et al.
Publicado: (2025)
por: Ni, Kangqi, et al.
Publicado: (2025)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
por: Foroutan, Negar, et al.
Publicado: (2025)
por: Foroutan, Negar, et al.
Publicado: (2025)
In-context Pretraining: Language Modeling Beyond Document Boundaries
por: Shi, Weijia, et al.
Publicado: (2023)
por: Shi, Weijia, et al.
Publicado: (2023)
Ejemplares similares
-
MobileCLIP2: Improving Multi-Modal Reinforced Training
por: Faghri, Fartash, et al.
Publicado: (2025) -
Learning to Reason for Hallucination Span Detection
por: Su, Hsuan, et al.
Publicado: (2025) -
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024) -
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
por: Lu, Yen-Ju, et al.
Publicado: (2025) -
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2023)