Modular Deep Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pfeiffer, Jonas, Ruder, Sebastian, Vulić, Ivan, Ponti, Edoardo Maria |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
von: Caccia, Lucas, et al.
Veröffentlicht: (2025)
von: Caccia, Lucas, et al.
Veröffentlicht: (2025)
Emergent Communication Pretraining for Few-Shot Machine Translation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020)
Scaling Sparse Fine-Tuning to Large Language Models
von: Ansell, Alan, et al.
Veröffentlicht: (2024)
von: Ansell, Alan, et al.
Veröffentlicht: (2024)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
von: Nawrot, Piotr, et al.
Veröffentlicht: (2025)
von: Nawrot, Piotr, et al.
Veröffentlicht: (2025)
Modular Multi-Task Learning for Chemical Reaction Prediction
von: Pang, Jiayun, et al.
Veröffentlicht: (2026)
von: Pang, Jiayun, et al.
Veröffentlicht: (2026)
Zero-Shot Tokenizer Transfer
von: Minixhofer, Benjamin, et al.
Veröffentlicht: (2024)
von: Minixhofer, Benjamin, et al.
Veröffentlicht: (2024)
Universal Cross-Tokenizer Distillation via Approximate Likelihood Matching
von: Minixhofer, Benjamin, et al.
Veröffentlicht: (2025)
von: Minixhofer, Benjamin, et al.
Veröffentlicht: (2025)
Specialising and Analysing Instruction-Tuned and Byte-Level Language Models for Organic Reaction Prediction
von: Pang, Jiayun, et al.
Veröffentlicht: (2024)
von: Pang, Jiayun, et al.
Veröffentlicht: (2024)
Towards Modular LLMs by Building and Reusing a Library of LoRAs
von: Ostapenko, Oleksiy, et al.
Veröffentlicht: (2024)
von: Ostapenko, Oleksiy, et al.
Veröffentlicht: (2024)
Adapting Time Series Foundation Models through Data Mixtures
von: Lee, Thomas L., et al.
Veröffentlicht: (2026)
von: Lee, Thomas L., et al.
Veröffentlicht: (2026)
Self-Augmented In-Context Learning for Unsupervised Word Translation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2024)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2024)
On the Independence Assumption in Neurosymbolic Learning
von: van Krieken, Emile, et al.
Veröffentlicht: (2024)
von: van Krieken, Emile, et al.
Veröffentlicht: (2024)
Neurosymbolic Diffusion Models
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
Neurosymbolic Reasoning Shortcuts under the Independence Assumption
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
von: van Krieken, Emile, et al.
Veröffentlicht: (2025)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
von: Ruder, Sebastian, et al.
Veröffentlicht: (2018)
Probing the Emergence of Cross-lingual Alignment during LLM Training
von: Wang, Hetong, et al.
Veröffentlicht: (2024)
von: Wang, Hetong, et al.
Veröffentlicht: (2024)
Inference-Time Hyper-Scaling with KV Cache Compression
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2025)
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2025)
DARE: Diverse Visual Question Answering with Robustness Evaluation
von: Sterz, Hannah, et al.
Veröffentlicht: (2024)
von: Sterz, Hannah, et al.
Veröffentlicht: (2024)
On Bilingual Lexicon Induction with Large Language Models
von: Li, Yaoyiran, et al.
Veröffentlicht: (2023)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2023)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
von: Cheng, Tianhao, et al.
Veröffentlicht: (2026)
von: Cheng, Tianhao, et al.
Veröffentlicht: (2026)
Modular Duality in Deep Learning
von: Bernstein, Jeremy, et al.
Veröffentlicht: (2024)
von: Bernstein, Jeremy, et al.
Veröffentlicht: (2024)
Graph data augmentation with Gromow-Wasserstein Barycenters
von: Ponti, Andrea
Veröffentlicht: (2024)
von: Ponti, Andrea
Veröffentlicht: (2024)
AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-Tuning
von: Zhou, Han, et al.
Veröffentlicht: (2023)
von: Zhou, Han, et al.
Veröffentlicht: (2023)
Agentic Policy Optimization via Instruction-Policy Co-Evolution
von: Zhou, Han, et al.
Veröffentlicht: (2025)
von: Zhou, Han, et al.
Veröffentlicht: (2025)
Improving Word Translation via Two-Stage Contrastive Learning
von: Li, Yaoyiran, et al.
Veröffentlicht: (2022)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2022)
Modular Deep-Learning-Based Early Warning System for Deadly Heatwave Prediction
von: Xu, Shangqing, et al.
Veröffentlicht: (2025)
von: Xu, Shangqing, et al.
Veröffentlicht: (2025)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
Model Merging by Uncertainty-Based Gradient Matching
von: Daheim, Nico, et al.
Veröffentlicht: (2023)
von: Daheim, Nico, et al.
Veröffentlicht: (2023)
PeakWeather: MeteoSwiss Weather Station Measurements for Spatiotemporal Deep Learning
von: Zambon, Daniele, et al.
Veröffentlicht: (2025)
von: Zambon, Daniele, et al.
Veröffentlicht: (2025)
Improving Bilingual Lexicon Induction with Cross-Encoder Reranking
von: Li, Yaoyiran, et al.
Veröffentlicht: (2022)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2022)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
Animal Behavior Analysis Methods Using Deep Learning: A Survey
von: Fazzari, Edoardo, et al.
Veröffentlicht: (2024)
von: Fazzari, Edoardo, et al.
Veröffentlicht: (2024)
Mixtures of In-Context Learners
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
Large Language Models are Miscalibrated In-Context Learners
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence Segmentation
von: Frohmann, Markus, et al.
Veröffentlicht: (2024)
von: Frohmann, Markus, et al.
Veröffentlicht: (2024)
Towards Optimal Adapter Placement for Efficient Transfer Learning
von: Nowak, Aleksandra I., et al.
Veröffentlicht: (2024)
von: Nowak, Aleksandra I., et al.
Veröffentlicht: (2024)
Deep Modularity Networks with Diversity-Preserving Regularization
von: Salehi, Yasmin, et al.
Veröffentlicht: (2025)
von: Salehi, Yasmin, et al.
Veröffentlicht: (2025)
How Does Quantization Affect Multilingual LLMs?
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
AdaSplash-2: Faster Differentiable Sparse Attention
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2026)
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2026)
GAS-Norm: Score-Driven Adaptive Normalization for Non-Stationary Time Series Forecasting in Deep Learning
von: Urettini, Edoardo, et al.
Veröffentlicht: (2024)
von: Urettini, Edoardo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
von: Caccia, Lucas, et al.
Veröffentlicht: (2025) -
Emergent Communication Pretraining for Few-Shot Machine Translation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020) -
Scaling Sparse Fine-Tuning to Large Language Models
von: Ansell, Alan, et al.
Veröffentlicht: (2024) -
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
von: Nawrot, Piotr, et al.
Veröffentlicht: (2025) -
Modular Multi-Task Learning for Chemical Reaction Prediction
von: Pang, Jiayun, et al.
Veröffentlicht: (2026)