Saved in:
| Main Authors: | Zheng, Aaron, Rana, Mansi, Stolcke, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.14398 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
by: Grosjean, Juri, et al.
Published: (2024)
by: Grosjean, Juri, et al.
Published: (2024)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
REFINE on Scarce Data: Retrieval Enhancement through Fine-Tuning via Model Fusion of Embedding Models
by: Gupta, Ambuje, et al.
Published: (2024)
by: Gupta, Ambuje, et al.
Published: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Memorization of Named Entities in Fine-tuned BERT Models
by: Diera, Andor, et al.
Published: (2022)
by: Diera, Andor, et al.
Published: (2022)
ACL: Aligned Contrastive Learning Improves BERT and Multi-exit BERT Fine-tuning
by: Li, Liz, et al.
Published: (2026)
by: Li, Liz, et al.
Published: (2026)
Bypassing Safety Guardrails in LLMs Using Humor
by: Cisneros-Velarde, Pedro
Published: (2025)
by: Cisneros-Velarde, Pedro
Published: (2025)
Detecting Bias in Large Language Models: Fine-tuned KcBERT
by: Lee, J. K., et al.
Published: (2024)
by: Lee, J. K., et al.
Published: (2024)
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
by: Ilin, Aleksei, et al.
Published: (2025)
by: Ilin, Aleksei, et al.
Published: (2025)
Fine-tuning BERT with Bidirectional LSTM for Fine-grained Movie Reviews Sentiment Analysis
by: Nkhata, Gibson, et al.
Published: (2025)
by: Nkhata, Gibson, et al.
Published: (2025)
Improving Sampling Methods for Fine-tuning SentenceBERT in Text Streams
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
Question-Answering System for Bangla: Fine-tuning BERT-Bangla for a Closed Domain
by: Roy, Subal Chandra, et al.
Published: (2024)
by: Roy, Subal Chandra, et al.
Published: (2024)
Improving Text Embeddings for Smaller Language Models Using Contrastive Fine-tuning
by: Ukarapol, Trapoom, et al.
Published: (2024)
by: Ukarapol, Trapoom, et al.
Published: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Few Dimensions are Enough: Fine-tuning BERT with Selected Dimensions Revealed Its Redundant Nature
by: Fukuhata, Shion, et al.
Published: (2025)
by: Fukuhata, Shion, et al.
Published: (2025)
Slot Filling as a Reasoning Task for SpeechLLMs
by: Hacioglu, Kadri, et al.
Published: (2025)
by: Hacioglu, Kadri, et al.
Published: (2025)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
by: Hościłowicz, Jakub, et al.
Published: (2023)
by: Hościłowicz, Jakub, et al.
Published: (2023)
CUE Vectors: Modular Training of Language Models Conditioned on Diverse Contextual Signals
by: Novotney, Scott, et al.
Published: (2022)
by: Novotney, Scott, et al.
Published: (2022)
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
by: Hacioglu, Kadri, et al.
Published: (2025)
by: Hacioglu, Kadri, et al.
Published: (2025)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Biomedical Entity Linking for Dutch: Fine-tuning a Self-alignment BERT Model on an Automatically Generated Wikipedia Corpus
by: Hartendorp, Fons, et al.
Published: (2024)
by: Hartendorp, Fons, et al.
Published: (2024)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
by: Yang, Yahan, et al.
Published: (2025)
by: Yang, Yahan, et al.
Published: (2025)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Zero-shot Slot Filling in the Age of LLMs for Dialogue Systems
by: Rana, Mansi, et al.
Published: (2024)
by: Rana, Mansi, et al.
Published: (2024)
ColBERT: Using BERT Sentence Embedding in Parallel Neural Networks for Computational Humor
by: Annamoradnejad, Issa, et al.
Published: (2020)
by: Annamoradnejad, Issa, et al.
Published: (2020)
InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline
by: Kugler, Kai, et al.
Published: (2021)
by: Kugler, Kai, et al.
Published: (2021)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
Harnessing Large Language Models: Fine-tuned BERT for Detecting Charismatic Leadership Tactics in Natural Language
by: Saeid, Yasser, et al.
Published: (2024)
by: Saeid, Yasser, et al.
Published: (2024)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
by: Jain, Samyak, et al.
Published: (2024)
by: Jain, Samyak, et al.
Published: (2024)
Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering
by: Chiniya, Purva, et al.
Published: (2026)
by: Chiniya, Purva, et al.
Published: (2026)
PARA: Parameter-Efficient Fine-tuning with Prompt Aware Representation Adjustment
by: Liu, Zequan, et al.
Published: (2025)
by: Liu, Zequan, et al.
Published: (2025)
Topic mining based on fine-tuning Sentence-BERT and LDA
by: Li, Jianheng, et al.
Published: (2025)
by: Li, Jianheng, et al.
Published: (2025)
CWTM: Leveraging Contextualized Word Embeddings from BERT for Neural Topic Modeling
by: Fang, Zheng, et al.
Published: (2023)
by: Fang, Zheng, et al.
Published: (2023)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
by: Huang, Wei-Chieh, et al.
Published: (2025)
by: Huang, Wei-Chieh, et al.
Published: (2025)
iBERT: Interpretable Embeddings via Sense Decomposition
by: Anand, Vishal, et al.
Published: (2025)
by: Anand, Vishal, et al.
Published: (2025)
GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning
by: Solatorio, Aivin V.
Published: (2024)
by: Solatorio, Aivin V.
Published: (2024)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
by: Wang, Guoli, et al.
Published: (2026)
by: Wang, Guoli, et al.
Published: (2026)
Similar Items
-
Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
by: Grosjean, Juri, et al.
Published: (2024) -
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026) -
REFINE on Scarce Data: Retrieval Enhancement through Fine-Tuning via Model Fusion of Embedding Models
by: Gupta, Ambuje, et al.
Published: (2024) -
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025) -
Memorization of Named Entities in Fine-tuned BERT Models
by: Diera, Andor, et al.
Published: (2022)