On Initializing Transformers with Pre-trained Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Ha Young, Balasubramanian, Niranjan, Kang, Byungkon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pre-training data selection for biomedical domain adaptation using journal impact metrics
by: Laï-king, Mathieu, et al.
Published: (2024)
by: Laï-king, Mathieu, et al.
Published: (2024)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
by: Fernández-González, Daniel, et al.
Published: (2026)
by: Fernández-González, Daniel, et al.
Published: (2026)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025)
by: Doghmash, Salam Thabet, et al.
Published: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
by: Kim, Songsoo, et al.
Published: (2025)
by: Kim, Songsoo, et al.
Published: (2025)
Influence-driven Curriculum Learning for Pre-training on Limited Data
by: Schoenegger, Loris, et al.
Published: (2025)
by: Schoenegger, Loris, et al.
Published: (2025)
Revisiting Word Embeddings in the LLM Era
by: Mahajan, Yash, et al.
Published: (2024)
by: Mahajan, Yash, et al.
Published: (2024)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs
by: Tytarenko, Stepan, et al.
Published: (2024)
by: Tytarenko, Stepan, et al.
Published: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
by: Xu, Wenjie, et al.
Published: (2023)
by: Xu, Wenjie, et al.
Published: (2023)
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
by: Kang, Migyeong, et al.
Published: (2026)
by: Kang, Migyeong, et al.
Published: (2026)
Zero-Shot Spam Email Classification Using Pre-trained Large Language Models
by: Rojas-Galeano, Sergio
Published: (2024)
by: Rojas-Galeano, Sergio
Published: (2024)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
by: Karinshak, Elise, et al.
Published: (2024)
by: Karinshak, Elise, et al.
Published: (2024)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)
by: He, Yanjin, et al.
Published: (2025)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
by: Gwak, Jiho, et al.
Published: (2025)
by: Gwak, Jiho, et al.
Published: (2025)
idT5: Indonesian Version of Multilingual T5 Transformer
by: Fuadi, Mukhlish, et al.
Published: (2023)
by: Fuadi, Mukhlish, et al.
Published: (2023)
Ensembling Multilingual Transformers for Robust Sentiment Analysis of Tweets
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2025)
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
myNER: Contextualized Burmese Named Entity Recognition with Bidirectional LSTM and fastText Embeddings via Joint Training with POS Tagging
by: Thant, Kaung Lwin, et al.
Published: (2025)
by: Thant, Kaung Lwin, et al.
Published: (2025)
Combining Denoising Autoencoders with Contrastive Learning to fine-tune Transformer Models
by: Lopez-Avila, Alejo, et al.
Published: (2024)
by: Lopez-Avila, Alejo, et al.
Published: (2024)
Linguistic Interpretability of Transformer-based Language Models: a systematic review
by: López-Otal, Miguel, et al.
Published: (2025)
by: López-Otal, Miguel, et al.
Published: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
Comparing Complex Concepts with Transformers: Matching Patent Claims Against Natural Language Text
by: Blume, Matthias, et al.
Published: (2024)
by: Blume, Matthias, et al.
Published: (2024)
Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
by: Riemenschneider, Frederick, et al.
Published: (2024)
by: Riemenschneider, Frederick, et al.
Published: (2024)
Please Make it Sound like Human: Encoder-Decoder vs. Decoder-Only Transformers for AI-to-Human Text Style Transfer
by: Paneru, Utsav
Published: (2026)
by: Paneru, Utsav
Published: (2026)
COA-GPT: Generative Pre-trained Transformers for Accelerated Course of Action Development in Military Operations
by: Goecks, Vinicius G., et al.
Published: (2024)
by: Goecks, Vinicius G., et al.
Published: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
by: Tereshchenko, Yehor, et al.
Published: (2025)
by: Tereshchenko, Yehor, et al.
Published: (2025)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models
by: Lee, Hwiyeong, et al.
Published: (2025)
by: Lee, Hwiyeong, et al.
Published: (2025)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
by: Seki, Yohei, et al.
Published: (2024)
by: Seki, Yohei, et al.
Published: (2024)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
by: Wang, Xintao, et al.
Published: (2026)
by: Wang, Xintao, et al.
Published: (2026)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
by: Pan, Xinghan
Published: (2025)
by: Pan, Xinghan
Published: (2025)
Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text Generation
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
Similar Items
-
Pre-training data selection for biomedical domain adaptation using journal impact metrics
by: Laï-king, Mathieu, et al.
Published: (2024) -
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
by: Fernández-González, Daniel, et al.
Published: (2026) -
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025) -
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
by: Kim, Songsoo, et al.
Published: (2025)