$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khandelwal, Apoorv, Yun, Tian, Nayak, Nihal V., Merullo, Jack, Bach, Stephen H., Sun, Chen, Pavlick, Ellie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Do Language Models Compose Functions?
von: Khandelwal, Apoorv, et al.
Veröffentlicht: (2025)
von: Khandelwal, Apoorv, et al.
Veröffentlicht: (2025)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2024)
von: Merullo, Jack, et al.
Veröffentlicht: (2024)
Circuit Component Reuse Across Tasks in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
Transformer Mechanisms Mimic Frontostriatal Gating Operations When Trained on Human Working Memory Tasks
von: Traylor, Aaron, et al.
Veröffentlicht: (2024)
von: Traylor, Aaron, et al.
Veröffentlicht: (2024)
Transferring Linear Features Across Language Models With Model Stitching
von: Chen, Alan, et al.
Veröffentlicht: (2025)
von: Chen, Alan, et al.
Veröffentlicht: (2025)
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting
von: Anand, Suraj, et al.
Veröffentlicht: (2024)
von: Anand, Suraj, et al.
Veröffentlicht: (2024)
What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models
von: Yun, Tian, et al.
Veröffentlicht: (2025)
von: Yun, Tian, et al.
Veröffentlicht: (2025)
mOthello: When Do Cross-Lingual Representation Alignment and Cross-Lingual Transfer Emerge in Multilingual Models?
von: Hua, Tianze, et al.
Veröffentlicht: (2024)
von: Hua, Tianze, et al.
Veröffentlicht: (2024)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
Does Training on Synthetic Data Make Models Less Robust?
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
Source-Modality Monitoring in Vision-Language Models
von: Hua, Etha Tianze, et al.
Veröffentlicht: (2026)
von: Hua, Etha Tianze, et al.
Veröffentlicht: (2026)
Video Finetuning Improves Reasoning Between Frames
von: Yang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Yang, Ruiqi, et al.
Veröffentlicht: (2025)
From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
von: Serre, Thomas, et al.
Veröffentlicht: (2025)
von: Serre, Thomas, et al.
Veröffentlicht: (2025)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
von: Nayak, Nihal V., et al.
Veröffentlicht: (2024)
von: Nayak, Nihal V., et al.
Veröffentlicht: (2024)
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
von: Abdullahi, Tassallah, et al.
Veröffentlicht: (2025)
von: Abdullahi, Tassallah, et al.
Veröffentlicht: (2025)
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
von: Kordi, Yeganeh, et al.
Veröffentlicht: (2025)
von: Kordi, Yeganeh, et al.
Veröffentlicht: (2025)
LLMs model how humans induce logically structured rules
von: Loo, Alyssa, et al.
Veröffentlicht: (2025)
von: Loo, Alyssa, et al.
Veröffentlicht: (2025)
Handling and Interpreting Missing Modalities in Patient Clinical Trajectories via Autoregressive Sequence Modeling
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
A Knapsack by Any Other Name: Presentation impacts LLM performance on NP-hard problems
von: Duchnowski, Alex, et al.
Veröffentlicht: (2025)
von: Duchnowski, Alex, et al.
Veröffentlicht: (2025)
Instilling Inductive Biases with Subnetworks
von: Zhang, Enyan, et al.
Veröffentlicht: (2023)
von: Zhang, Enyan, et al.
Veröffentlicht: (2023)
The dynamic interplay between in-context and in-weight learning in humans and neural networks
von: Russin, Jacob, et al.
Veröffentlicht: (2024)
von: Russin, Jacob, et al.
Veröffentlicht: (2024)
Uncovering Intermediate Variables in Transformers using Circuit Probing
von: Lepori, Michael A., et al.
Veröffentlicht: (2023)
von: Lepori, Michael A., et al.
Veröffentlicht: (2023)
Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval Models
von: Chen, Catherine, et al.
Veröffentlicht: (2024)
von: Chen, Catherine, et al.
Veröffentlicht: (2024)
LLMs as Models for Analogical Reasoning
von: Musker, Sam, et al.
Veröffentlicht: (2024)
von: Musker, Sam, et al.
Veröffentlicht: (2024)
Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline
von: Lu, Meng, et al.
Veröffentlicht: (2025)
von: Lu, Meng, et al.
Veröffentlicht: (2025)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
I Have No Mouth, and I Must Rhyme: Uncovering Internal Phonetic Representations in LLaMA 3.2
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2025)
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2025)
FLM-101B: An Open LLM and How to Train It with $100K Budget
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
Revisiting Privacy-Utility Trade-off for DP Training with Pre-existing Knowledge
von: Zheng, Yu, et al.
Veröffentlicht: (2024)
von: Zheng, Yu, et al.
Veröffentlicht: (2024)
Spin-force from a Nitrogen-Vacancy ensemble drives a 100 mg levitated resonator
von: Nayak, Anshuman, et al.
Veröffentlicht: (2026)
von: Nayak, Anshuman, et al.
Veröffentlicht: (2026)
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
von: Zhang, Ruochen, et al.
Veröffentlicht: (2024)
von: Zhang, Ruochen, et al.
Veröffentlicht: (2024)
A Convex Loss Function for Set Prediction with Optimal Trade-offs Between Size and Conditional Coverage
von: Bach, Francis
Veröffentlicht: (2025)
von: Bach, Francis
Veröffentlicht: (2025)
From Memorization to Reasoning in the Spectrum of Loss Curvature
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
TV100: A TV Series Dataset that Pre-Trained CLIP Has Not Seen
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
von: Wang, Weizhi, et al.
Veröffentlicht: (2025)
von: Wang, Weizhi, et al.
Veröffentlicht: (2025)
Positive Pre‐Transplant Respiratory Viral PCR is Associated With Increased Day 100 Transplant‐Related Mortality in Pediatric HSCT Recipients
von: Jane Trainor, et al.
Veröffentlicht: (2025)
von: Jane Trainor, et al.
Veröffentlicht: (2025)
smithkunieda/Trade-TGN: v1.0.0
von: smithkunieda
Veröffentlicht: (2026)
von: smithkunieda
Veröffentlicht: (2026)
Are LLMs Models of Distributional Semantics? A Case Study on Quantifiers
von: Enyan, Zhang, et al.
Veröffentlicht: (2024)
von: Enyan, Zhang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Do Language Models Compose Functions?
von: Khandelwal, Apoorv, et al.
Veröffentlicht: (2025) -
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
von: Lewis, Martha, et al.
Veröffentlicht: (2022) -
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2024) -
Circuit Component Reuse Across Tasks in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2023) -
Language Models Implement Simple Word2Vec-style Vector Arithmetic
von: Merullo, Jack, et al.
Veröffentlicht: (2023)