REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Seungmin, Lee, Jeonghwan, Lim, Hyunkuk, Kim, Sejoon, Sung, Mingi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Terminology Integration for LLM-based Translation in Specialized Domains
by: Kim, Sejoon, et al.
Published: (2024)
by: Kim, Sejoon, et al.
Published: (2024)
Context-Aware LLM Translation System Using Conversation Summarization and Dialogue History
by: Sung, Mingi, et al.
Published: (2024)
by: Sung, Mingi, et al.
Published: (2024)
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
by: Yun, Taeyang, et al.
Published: (2024)
by: Yun, Taeyang, et al.
Published: (2024)
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
by: Kim, SungHo, et al.
Published: (2026)
by: Kim, SungHo, et al.
Published: (2026)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
by: Kim, Yumin, et al.
Published: (2024)
by: Kim, Yumin, et al.
Published: (2024)
PreFT: Prefill-only finetuning for efficient inference
by: Lanpouthakoun, Andrew, et al.
Published: (2026)
by: Lanpouthakoun, Andrew, et al.
Published: (2026)
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
by: Kim, SungHo, et al.
Published: (2026)
by: Kim, SungHo, et al.
Published: (2026)
EmbeddingGemma: Powerful and Lightweight Text Representations
by: Vera, Henrique Schechter, et al.
Published: (2025)
by: Vera, Henrique Schechter, et al.
Published: (2025)
Denoising Table-Text Retrieval for Open-Domain Question Answering
by: Kang, Deokhyung, et al.
Published: (2024)
by: Kang, Deokhyung, et al.
Published: (2024)
Incorporating Domain Knowledge into Materials Tokenization
by: Oh, Yerim, et al.
Published: (2025)
by: Oh, Yerim, et al.
Published: (2025)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
by: Seo, Jean, et al.
Published: (2026)
by: Seo, Jean, et al.
Published: (2026)
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
by: Song, Hwanjun, et al.
Published: (2025)
by: Song, Hwanjun, et al.
Published: (2025)
Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise
by: Hong, Giwon, et al.
Published: (2023)
by: Hong, Giwon, et al.
Published: (2023)
Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features
by: Lee, Bruce W., et al.
Published: (2021)
by: Lee, Bruce W., et al.
Published: (2021)
Low-rank finetuning for LLMs: A fairness perspective
by: Das, Saswat, et al.
Published: (2024)
by: Das, Saswat, et al.
Published: (2024)
Ontology-Free General-Domain Knowledge Graph-to-Text Generation Dataset Synthesis using Large Language Model
by: Kim, Daehee, et al.
Published: (2024)
by: Kim, Daehee, et al.
Published: (2024)
Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
by: Lim, Seungseop, et al.
Published: (2025)
by: Lim, Seungseop, et al.
Published: (2025)
Confidence Regularized Masked Language Modeling using Text Length
by: Ji, Seunghyun, et al.
Published: (2025)
by: Ji, Seunghyun, et al.
Published: (2025)
Semantic Aware Linear Transfer by Recycling Pre-trained Language Models for Cross-lingual Transfer
by: Lee, Seungyoon, et al.
Published: (2025)
by: Lee, Seungyoon, et al.
Published: (2025)
Analysis of Utterance Embeddings and Clustering Methods Related to Intent Induction for Task-Oriented Dialogue
by: Park, Jeiyoon, et al.
Published: (2022)
by: Park, Jeiyoon, et al.
Published: (2022)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
STAGE: Simplified Text-Attributed Graph Embeddings Using Pre-trained LLMs
by: Zolnai-Lucas, Aaron, et al.
Published: (2024)
by: Zolnai-Lucas, Aaron, et al.
Published: (2024)
Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games
by: Lim, Seungwon, et al.
Published: (2025)
by: Lim, Seungwon, et al.
Published: (2025)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
by: Ban, Minjeong, et al.
Published: (2026)
by: Ban, Minjeong, et al.
Published: (2026)
Quantifying Positional Biases in Text Embedding Models
by: Lee, Reagan J., et al.
Published: (2024)
by: Lee, Reagan J., et al.
Published: (2024)
BioBridge: Unified Bio-Embedding with Bridging Modality in Code-Switched EMR
by: Jeon, Jangyeong, et al.
Published: (2024)
by: Jeon, Jangyeong, et al.
Published: (2024)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
by: Moon, Sehwan, et al.
Published: (2025)
by: Moon, Sehwan, et al.
Published: (2025)
Pre-Storage Reasoning for Episodic Memory: Shifting Inference Burden to Memory for Personalized Dialogue
by: Kim, Sangyeop, et al.
Published: (2025)
by: Kim, Sangyeop, et al.
Published: (2025)
DSG-KD: Knowledge Distillation from Domain-Specific to General Language Models
by: Cho, Sangyeon, et al.
Published: (2024)
by: Cho, Sangyeon, et al.
Published: (2024)
Length-Aware Rotary Position Embedding for Text-Speech Alignment
by: Kim, Hyeongju, et al.
Published: (2025)
by: Kim, Hyeongju, et al.
Published: (2025)
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)
by: Lampinen, Andrew K., et al.
Published: (2025)
RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation
by: Kim, Kiseung, et al.
Published: (2024)
by: Kim, Kiseung, et al.
Published: (2024)
THEME: Enhancing Thematic Investing with Semantic Stock Representations and Temporal Dynamics
by: Lee, Hoyoung, et al.
Published: (2025)
by: Lee, Hoyoung, et al.
Published: (2025)
QIME: Constructing Interpretable Medical Text Embeddings via Ontology-Grounded Questions
by: Tang, Yixuan, et al.
Published: (2026)
by: Tang, Yixuan, et al.
Published: (2026)
Memory-efficient Energy-adaptive Inference of Pre-Trained Models on Batteryless Embedded Systems
by: Farina, Pietro, et al.
Published: (2024)
by: Farina, Pietro, et al.
Published: (2024)
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
by: Kim, SungHo, et al.
Published: (2025)
by: Kim, SungHo, et al.
Published: (2025)
Algorithm Selection with Zero Domain Knowledge via Text Embeddings
by: Szeider, Stefan
Published: (2026)
by: Szeider, Stefan
Published: (2026)
Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling
by: Byeon, Junhyeong, et al.
Published: (2026)
by: Byeon, Junhyeong, et al.
Published: (2026)
MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
Prompt-based Learning for Text Readability Assessment
by: Lee, Bruce W., et al.
Published: (2023)
by: Lee, Bruce W., et al.
Published: (2023)
Similar Items
-
Efficient Terminology Integration for LLM-based Translation in Specialized Domains
by: Kim, Sejoon, et al.
Published: (2024) -
Context-Aware LLM Translation System Using Conversation Summarization and Dialogue History
by: Sung, Mingi, et al.
Published: (2024) -
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
by: Yun, Taeyang, et al.
Published: (2024) -
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
by: Kim, SungHo, et al.
Published: (2026) -
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
by: Kim, Yumin, et al.
Published: (2024)