Salvato in:
| Autori principali: | Kristiansen, Gus, Sandler, Mark, Zhmoginov, Andrey, Miller, Nolan, Goyal, Anirudh, Lee, Jihwan, Vladymyrov, Max |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2408.09310 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Contextually Guided Transformers via Low-Rank Adaptation
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
Continual HyperTransformer: A Meta-Learner for Continual Few-Shot Learning
di: Vladymyrov, Max, et al.
Pubblicazione: (2023)
di: Vladymyrov, Max, et al.
Pubblicazione: (2023)
Long Context In-Context Compression by Getting to the Gist of Gisting
di: Petrov, Aleksandar, et al.
Pubblicazione: (2025)
di: Petrov, Aleksandar, et al.
Pubblicazione: (2025)
Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
Learning and Unlearning of Fabricated Knowledge in Language Models
di: Sun, Chen, et al.
Pubblicazione: (2024)
di: Sun, Chen, et al.
Pubblicazione: (2024)
Linear Transformers are Versatile In-Context Learners
di: Vladymyrov, Max, et al.
Pubblicazione: (2024)
di: Vladymyrov, Max, et al.
Pubblicazione: (2024)
How new data permeates LLM knowledge and how to dilute it
di: Sun, Chen, et al.
Pubblicazione: (2025)
di: Sun, Chen, et al.
Pubblicazione: (2025)
Uncovering mesa-optimization algorithms in Transformers
di: von Oswald, Johannes, et al.
Pubblicazione: (2023)
di: von Oswald, Johannes, et al.
Pubblicazione: (2023)
Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
di: Subramanyam, Anirudh, et al.
Pubblicazione: (2025)
di: Subramanyam, Anirudh, et al.
Pubblicazione: (2025)
MELODI: Exploring Memory Compression for Long Contexts
di: Chen, Yinpeng, et al.
Pubblicazione: (2024)
di: Chen, Yinpeng, et al.
Pubblicazione: (2024)
Non-Convex Optimization with Spectral Radius Regularization
di: Sandler, Adam, et al.
Pubblicazione: (2021)
di: Sandler, Adam, et al.
Pubblicazione: (2021)
DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
Can Models Learn Skill Composition from Examples?
di: Zhao, Haoyu, et al.
Pubblicazione: (2024)
di: Zhao, Haoyu, et al.
Pubblicazione: (2024)
Fast Differentiable Modal Simulation of Non-linear Strings, Membranes, and Plates
di: Diaz, Rodrigo, et al.
Pubblicazione: (2025)
di: Diaz, Rodrigo, et al.
Pubblicazione: (2025)
Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
di: Goyal, Sachin, et al.
Pubblicazione: (2025)
di: Goyal, Sachin, et al.
Pubblicazione: (2025)
On the Impossibility of Retrain Equivalence in Machine Unlearning
di: Yu, Jiatong, et al.
Pubblicazione: (2025)
di: Yu, Jiatong, et al.
Pubblicazione: (2025)
LLM Agents for Bargaining with Utility-based Feedback
di: Oh, Jihwan
Pubblicazione: (2025)
di: Oh, Jihwan
Pubblicazione: (2025)
Heterogeneous Federated Learning with Prototype Alignment and Upscaling
di: Lee, Gyuejeong, et al.
Pubblicazione: (2025)
di: Lee, Gyuejeong, et al.
Pubblicazione: (2025)
Towards Efficient Modelling of String Dynamics: A Comparison of State Space and Koopman based Deep Learning Methods
di: Diaz, Rodrigo, et al.
Pubblicazione: (2024)
di: Diaz, Rodrigo, et al.
Pubblicazione: (2024)
Faraday: Synthetic Smart Meter Generator for the smart grid
di: Chai, Sheng, et al.
Pubblicazione: (2024)
di: Chai, Sheng, et al.
Pubblicazione: (2024)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
di: Kaur, Simran, et al.
Pubblicazione: (2024)
di: Kaur, Simran, et al.
Pubblicazione: (2024)
$α$-TCVAE: On the relationship between Disentanglement and Diversity
di: Meo, Cristian, et al.
Pubblicazione: (2024)
di: Meo, Cristian, et al.
Pubblicazione: (2024)
Improving Sparse Memory Finetuning
di: Goyal, Satyam, et al.
Pubblicazione: (2026)
di: Goyal, Satyam, et al.
Pubblicazione: (2026)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
di: Didolkar, Aniket, et al.
Pubblicazione: (2025)
di: Didolkar, Aniket, et al.
Pubblicazione: (2025)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
di: Gupta, Prakhar, et al.
Pubblicazione: (2026)
di: Gupta, Prakhar, et al.
Pubblicazione: (2026)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2023)
di: Prabhudesai, Mihir, et al.
Pubblicazione: (2023)
Latent Diffusion Pretraining for Crystal Property Prediction
di: Mukherjee, Shrimon, et al.
Pubblicazione: (2026)
di: Mukherjee, Shrimon, et al.
Pubblicazione: (2026)
Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models
di: Dang, Xingyu, et al.
Pubblicazione: (2026)
di: Dang, Xingyu, et al.
Pubblicazione: (2026)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
Partial Inverse Design of High-Performance Concrete Using Cooperative Neural Networks for Constraint-Aware Mix Generation
di: Nugraha, Agung, et al.
Pubblicazione: (2025)
di: Nugraha, Agung, et al.
Pubblicazione: (2025)
Towards Diverse Evaluation of Class Incremental Learning: A Representation Learning Perspective
di: Cha, Sungmin, et al.
Pubblicazione: (2022)
di: Cha, Sungmin, et al.
Pubblicazione: (2022)
Evaluation of Neural Surrogates for Physical Modelling Synthesis of Nonlinear Elastic Plates
di: Martin, Carlos De La Vega, et al.
Pubblicazione: (2025)
di: Martin, Carlos De La Vega, et al.
Pubblicazione: (2025)
Detecting Pretraining Data from Large Language Models
di: Shi, Weijia, et al.
Pubblicazione: (2023)
di: Shi, Weijia, et al.
Pubblicazione: (2023)
When Should We Introduce Safety Interventions During Pretraining?
di: Sam, Dylan, et al.
Pubblicazione: (2026)
di: Sam, Dylan, et al.
Pubblicazione: (2026)
BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames
di: Mark, Max Sobol, et al.
Pubblicazione: (2026)
di: Mark, Max Sobol, et al.
Pubblicazione: (2026)
ARCLE: The Abstraction and Reasoning Corpus Learning Environment for Reinforcement Learning
di: Lee, Hosung, et al.
Pubblicazione: (2024)
di: Lee, Hosung, et al.
Pubblicazione: (2024)
A Granular Study of Safety Pretraining under Model Abliteration
di: Agnihotri, Shashank, et al.
Pubblicazione: (2025)
di: Agnihotri, Shashank, et al.
Pubblicazione: (2025)
Robust and Consistent Ski Rental with Distributional Advice
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
Benchmarking Optimizers for Large Language Model Pretraining
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Contextually Guided Transformers via Low-Rank Adaptation
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025) -
Continual HyperTransformer: A Meta-Learner for Continual Few-Shot Learning
di: Vladymyrov, Max, et al.
Pubblicazione: (2023) -
Long Context In-Context Compression by Getting to the Gist of Gisting
di: Petrov, Aleksandar, et al.
Pubblicazione: (2025) -
Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025) -
Learning and Unlearning of Fabricated Knowledge in Language Models
di: Sun, Chen, et al.
Pubblicazione: (2024)