From Robustness to Improved Generalization and Calibration in Pre-trained Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jukić, Josip, Šnajder, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Disentangling Latent Shifts of In-Context Learning with Weak Supervision
by: Jukić, Josip, et al.
Published: (2024)
by: Jukić, Josip, et al.
Published: (2024)
Context Parametrization with Compositional Adapters
by: Jukić, Josip, et al.
Published: (2025)
by: Jukić, Josip, et al.
Published: (2025)
Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis
by: Jukić, Josip
Published: (2025)
by: Jukić, Josip
Published: (2025)
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
by: Jelenić, Fran, et al.
Published: (2023)
by: Jelenić, Fran, et al.
Published: (2023)
Claim Check-Worthiness Detection: How Well do LLMs Grasp Annotation Guidelines?
by: Majer, Laura, et al.
Published: (2024)
by: Majer, Laura, et al.
Published: (2024)
Looking Right is Sometimes Right: Investigating the Capabilities of Decoder-only LLMs for Sequence Labeling
by: Dukić, David, et al.
Published: (2024)
by: Dukić, David, et al.
Published: (2024)
Sequence Repetition Enhances Token Embeddings and Improves Sequence Labeling with Decoder-only Language Models
by: Kukić, Matija Luka, et al.
Published: (2026)
by: Kukić, Matija Luka, et al.
Published: (2026)
Supervised In-Context Fine-Tuning for Generative Sequence Labeling
by: Dukić, David, et al.
Published: (2025)
by: Dukić, David, et al.
Published: (2025)
Characterizing Linguistic Shifts in Croatian News via Diachronic Word Embeddings
by: Dukić, David, et al.
Published: (2025)
by: Dukić, David, et al.
Published: (2025)
Leveraging Open Information Extraction for More Robust Domain Transfer of Event Trigger Detection
by: Dukić, David, et al.
Published: (2023)
by: Dukić, David, et al.
Published: (2023)
Are ELECTRA's Sentence Embeddings Beyond Repair? The Case of Semantic Textual Similarity
by: Rep, Ivan, et al.
Published: (2024)
by: Rep, Ivan, et al.
Published: (2024)
LLMs for Targeted Sentiment in News Headlines: Exploring the Descriptive-Prescriptive Dilemma
by: Juroš, Jana, et al.
Published: (2024)
by: Juroš, Jana, et al.
Published: (2024)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
What Makes You CLIC: Detection of Croatian Clickbait Headlines
by: Anđelić, Marija, et al.
Published: (2025)
by: Anđelić, Marija, et al.
Published: (2025)
TRELM: Towards Robust and Efficient Pre-training for Knowledge-Enhanced Language Models
by: Yan, Junbing, et al.
Published: (2024)
by: Yan, Junbing, et al.
Published: (2024)
Aligning Pre-trained Models for Spoken Language Translation
by: Sedláček, Šimon, et al.
Published: (2024)
by: Sedláček, Šimon, et al.
Published: (2024)
TakeLab Retriever: AI-Driven Search Engine for Articles from Croatian News Outlets
by: Dukić, David, et al.
Published: (2024)
by: Dukić, David, et al.
Published: (2024)
Improving Transfer Learning for Sequence Labeling Tasks by Adapting Pre-trained Neural Language Models
by: Dukić, David
Published: (2025)
by: Dukić, David
Published: (2025)
From N-grams to Pre-trained Multilingual Models For Language Identification
by: Sindane, Thapelo, et al.
Published: (2024)
by: Sindane, Thapelo, et al.
Published: (2024)
On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase Generation
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
by: Naous, Tarek, et al.
Published: (2025)
by: Naous, Tarek, et al.
Published: (2025)
Pre-trained Language Models for Keyphrase Generation: A Thorough Empirical Study
by: Wu, Di, et al.
Published: (2022)
by: Wu, Di, et al.
Published: (2022)
Development of Cognitive Intelligence in Pre-trained Language Models
by: Shah, Raj Sanjay, et al.
Published: (2024)
by: Shah, Raj Sanjay, et al.
Published: (2024)
Metadata Conditioning Accelerates Language Model Pre-training
by: Gao, Tianyu, et al.
Published: (2025)
by: Gao, Tianyu, et al.
Published: (2025)
Evaluating Discourse Cohesion in Pre-trained Language Models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Performance Trade-offs of Optimizing Small Language Models for E-Commerce
by: Licardo, Josip Tomo, et al.
Published: (2025)
by: Licardo, Josip Tomo, et al.
Published: (2025)
Fine-tuning Pre-trained Language Models for Few-shot Intent Detection: Supervised Pre-training and Isotropization
by: Zhang, Haode, et al.
Published: (2022)
by: Zhang, Haode, et al.
Published: (2022)
B-cos LM: Efficiently Transforming Pre-trained Language Models for Improved Explainability
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
SongSage: A Large Musical Language Model with Lyric Generative Pre-training
by: Guo, Jiani, et al.
Published: (2026)
by: Guo, Jiani, et al.
Published: (2026)
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
by: Wan, Zhongwei, et al.
Published: (2022)
by: Wan, Zhongwei, et al.
Published: (2022)
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
by: Ye, Liqin, et al.
Published: (2025)
by: Ye, Liqin, et al.
Published: (2025)
RepCali: High Efficient Fine-tuning Via Representation Calibration in Latent Space for Pre-trained Language Models
by: Zhang, Fujun, et al.
Published: (2025)
by: Zhang, Fujun, et al.
Published: (2025)
Bag of Lies: Robustness in Continuous Pre-training BERT
by: Gevers, Ine, et al.
Published: (2024)
by: Gevers, Ine, et al.
Published: (2024)
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models
by: Zhang, Ruiqi, et al.
Published: (2026)
by: Zhang, Ruiqi, et al.
Published: (2026)
Stable Language Model Pre-training by Reducing Embedding Variability
by: Chung, Woojin, et al.
Published: (2024)
by: Chung, Woojin, et al.
Published: (2024)
Tracr-Injection: Distilling Algorithms into Pre-trained Language Models
by: Vergara-Browne, Tomás, et al.
Published: (2025)
by: Vergara-Browne, Tomás, et al.
Published: (2025)
How Ready Are Generative Pre-trained Large Language Models for Explaining Bengali Grammatical Errors?
by: Maity, Subhankar, et al.
Published: (2024)
by: Maity, Subhankar, et al.
Published: (2024)
Model Merging in Pre-training of Large Language Models
by: Li, Yunshui, et al.
Published: (2025)
by: Li, Yunshui, et al.
Published: (2025)
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
by: Kaing, Hour, et al.
Published: (2025)
by: Kaing, Hour, et al.
Published: (2025)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
Similar Items
-
Disentangling Latent Shifts of In-Context Learning with Weak Supervision
by: Jukić, Josip, et al.
Published: (2024) -
Context Parametrization with Compositional Adapters
by: Jukić, Josip, et al.
Published: (2025) -
Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis
by: Jukić, Josip
Published: (2025) -
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
by: Jelenić, Fran, et al.
Published: (2023) -
Claim Check-Worthiness Detection: How Well do LLMs Grasp Annotation Guidelines?
by: Majer, Laura, et al.
Published: (2024)