Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
Fuente:
arXiv
Guardado en:
| Autores principales: | Hościłowicz, Jakub, Sowański, Marcin, Czubowski, Piotr, Janicki, Artur |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Large Language Models for Expansion of Spoken Language Understanding Systems to New Languages
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
por: Hoscilowicz, Jakub, et al.
Publicado: (2025)
por: Hoscilowicz, Jakub, et al.
Publicado: (2025)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
Relational Knowledge Distillation Using Fine-tuned Function Vectors
por: Kang, Andrea, et al.
Publicado: (2026)
por: Kang, Andrea, et al.
Publicado: (2026)
Large Language Models as Carriers of Hidden Messages
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
Improving Sampling Methods for Fine-tuning SentenceBERT in Text Streams
por: Garcia, Cristiano Mesquita, et al.
Publicado: (2024)
por: Garcia, Cristiano Mesquita, et al.
Publicado: (2024)
OpenAutoNLU: Open Source AutoML Library for NLU
por: Arshinov, Grigory, et al.
Publicado: (2026)
por: Arshinov, Grigory, et al.
Publicado: (2026)
Steerability of Instrumental-Convergence Tendencies in LLMs
por: Hoscilowicz, Jakub
Publicado: (2026)
por: Hoscilowicz, Jakub
Publicado: (2026)
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
por: Zhang, Ying, et al.
Publicado: (2024)
por: Zhang, Ying, et al.
Publicado: (2024)
An investigation of structures responsible for gender bias in BERT and DistilBERT
por: Leteno, Thibaud, et al.
Publicado: (2024)
por: Leteno, Thibaud, et al.
Publicado: (2024)
Instruction-tuned Language Models are Better Knowledge Learners
por: Jiang, Zhengbao, et al.
Publicado: (2024)
por: Jiang, Zhengbao, et al.
Publicado: (2024)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
por: Wu, Taiqiang, et al.
Publicado: (2023)
por: Wu, Taiqiang, et al.
Publicado: (2023)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
por: Zhou, Huichi, et al.
Publicado: (2025)
por: Zhou, Huichi, et al.
Publicado: (2025)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
por: Krajewski, Jakub, et al.
Publicado: (2025)
por: Krajewski, Jakub, et al.
Publicado: (2025)
Harnessing Large Language Models: Fine-tuned BERT for Detecting Charismatic Leadership Tactics in Natural Language
por: Saeid, Yasser, et al.
Publicado: (2024)
por: Saeid, Yasser, et al.
Publicado: (2024)
Energy and Carbon Considerations of Fine-Tuning BERT
por: Wang, Xiaorong, et al.
Publicado: (2023)
por: Wang, Xiaorong, et al.
Publicado: (2023)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
por: Yao, Kai, et al.
Publicado: (2024)
por: Yao, Kai, et al.
Publicado: (2024)
Harnessing Test-time Adaptation for NLU tasks Involving Dialects of English
por: Nguyen, Duke, et al.
Publicado: (2025)
por: Nguyen, Duke, et al.
Publicado: (2025)
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
por: Rabby, Gollam, et al.
Publicado: (2024)
por: Rabby, Gollam, et al.
Publicado: (2024)
Can Perplexity Predict Fine-tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali
por: Luitel, Nishant, et al.
Publicado: (2024)
por: Luitel, Nishant, et al.
Publicado: (2024)
Empowering Small-Scale Knowledge Graphs: A Strategy of Leveraging General-Purpose Knowledge Graphs for Enriched Embeddings
por: Sawczyn, Albert, et al.
Publicado: (2024)
por: Sawczyn, Albert, et al.
Publicado: (2024)
EpilepsyLLM: Domain-Specific Large Language Model Fine-tuned with Epilepsy Medical Knowledge
por: Zhao, Xuyang, et al.
Publicado: (2024)
por: Zhao, Xuyang, et al.
Publicado: (2024)
A Contextualized BERT model for Knowledge Graph Completion
por: Gul, Haji, et al.
Publicado: (2024)
por: Gul, Haji, et al.
Publicado: (2024)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
por: Bihani, Geetanjali, et al.
Publicado: (2024)
por: Bihani, Geetanjali, et al.
Publicado: (2024)
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
por: Song, Yuda, et al.
Publicado: (2024)
por: Song, Yuda, et al.
Publicado: (2024)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
por: Bai, Xueying, et al.
Publicado: (2024)
por: Bai, Xueying, et al.
Publicado: (2024)
Enhancing BERT Fine-Tuning for Sentiment Analysis in Lower-Resourced Languages
por: Kubík, Jozef, et al.
Publicado: (2025)
por: Kubík, Jozef, et al.
Publicado: (2025)
Building a Few-Shot Cross-Domain Multilingual NLU Model for Customer Care
por: Kumar, Saurabh, et al.
Publicado: (2025)
por: Kumar, Saurabh, et al.
Publicado: (2025)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
por: Jung, Hee-Jun, et al.
Publicado: (2022)
por: Jung, Hee-Jun, et al.
Publicado: (2022)
Enhancing TinyBERT for Financial Sentiment Analysis Using GPT-Augmented FinBERT Distillation
por: Thomas, Graison Jos
Publicado: (2024)
por: Thomas, Graison Jos
Publicado: (2024)
Deep Learning for Medical Text Processing: BERT Model Fine-Tuning and Comparative Study
por: Hu, Jiacheng, et al.
Publicado: (2024)
por: Hu, Jiacheng, et al.
Publicado: (2024)
Topic Modeling with Fine-tuning LLMs and Bag of Sentences
por: Schneider, Johannes
Publicado: (2024)
por: Schneider, Johannes
Publicado: (2024)
SEE: Continual Fine-tuning with Sequential Ensemble of Experts
por: Wang, Zhilin, et al.
Publicado: (2025)
por: Wang, Zhilin, et al.
Publicado: (2025)
On the Loss of Context-awareness in General Instruction Fine-tuning
por: Wang, Yihan, et al.
Publicado: (2024)
por: Wang, Yihan, et al.
Publicado: (2024)
$\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge
por: Park, Core Francisco, et al.
Publicado: (2025)
por: Park, Core Francisco, et al.
Publicado: (2025)
Compact Language Models via Pruning and Knowledge Distillation
por: Muralidharan, Saurav, et al.
Publicado: (2024)
por: Muralidharan, Saurav, et al.
Publicado: (2024)
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language
por: Gete, Dawit Ketema, et al.
Publicado: (2025)
por: Gete, Dawit Ketema, et al.
Publicado: (2025)
Efficient Ensemble for Fine-tuning Language Models on Multiple Datasets
por: Li, Dongyue, et al.
Publicado: (2025)
por: Li, Dongyue, et al.
Publicado: (2025)
LoRA vs Full Fine-tuning: An Illusion of Equivalence
por: Shuttleworth, Reece, et al.
Publicado: (2024)
por: Shuttleworth, Reece, et al.
Publicado: (2024)
Ejemplares similares
-
Large Language Models for Expansion of Spoken Language Understanding Systems to New Languages
por: Hoscilowicz, Jakub, et al.
Publicado: (2024) -
Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
por: Hoscilowicz, Jakub, et al.
Publicado: (2025) -
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
por: Hoscilowicz, Jakub, et al.
Publicado: (2024) -
ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents
por: Hoscilowicz, Jakub, et al.
Publicado: (2024) -
Relational Knowledge Distillation Using Fine-tuned Function Vectors
por: Kang, Andrea, et al.
Publicado: (2026)