Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Them
Fuente:
arXiv
Guardado en:
| Autores principales: | Brinner, Marc, Mustafa, Tarek Al, Zarrieß, Sina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
por: Brinner, Marc, et al.
Publicado: (2025)
por: Brinner, Marc, et al.
Publicado: (2025)
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
por: Brinner, Marc, et al.
Publicado: (2025)
por: Brinner, Marc, et al.
Publicado: (2025)
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
por: Brinner, Marc, et al.
Publicado: (2025)
por: Brinner, Marc, et al.
Publicado: (2025)
Model Interpretability and Rationale Extraction by Input Mask Optimization
por: Brinner, Marc, et al.
Publicado: (2025)
por: Brinner, Marc, et al.
Publicado: (2025)
Efficient Scientific Full Text Classification: The Case of EICAT Impact Assessments
por: Brinner, Marc Felix, et al.
Publicado: (2025)
por: Brinner, Marc Felix, et al.
Publicado: (2025)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
por: Sieker, Judith, et al.
Publicado: (2026)
por: Sieker, Judith, et al.
Publicado: (2026)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
por: Kraus, Oliver, et al.
Publicado: (2026)
por: Kraus, Oliver, et al.
Publicado: (2026)
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
por: Pezeshkpour, Pouya, et al.
Publicado: (2025)
por: Pezeshkpour, Pouya, et al.
Publicado: (2025)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
por: Jadon, Aryan, et al.
Publicado: (2025)
por: Jadon, Aryan, et al.
Publicado: (2025)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
por: Wuhrmann, Arthur, et al.
Publicado: (2025)
por: Wuhrmann, Arthur, et al.
Publicado: (2025)
How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
por: Xiao, Teng, et al.
Publicado: (2024)
por: Xiao, Teng, et al.
Publicado: (2024)
Efficient Sample-Specific Encoder Perturbations
por: Fathullah, Yassir, et al.
Publicado: (2024)
por: Fathullah, Yassir, et al.
Publicado: (2024)
A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection
por: Reis, Philipp, et al.
Publicado: (2026)
por: Reis, Philipp, et al.
Publicado: (2026)
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
por: Arakelyan, Erik
Publicado: (2025)
por: Arakelyan, Erik
Publicado: (2025)
Anka: A Domain-Specific Language for Reliable LLM Code Generation
por: Mazrouei, Saif Khalfan Saif Al
Publicado: (2025)
por: Mazrouei, Saif Khalfan Saif Al
Publicado: (2025)
Discursive Circuits: How Do Language Models Understand Discourse Relations?
por: Miao, Yisong, et al.
Publicado: (2025)
por: Miao, Yisong, et al.
Publicado: (2025)
(How) Do Language Models Track State?
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
Clinical Reading Comprehension with Encoder-Decoder Models Enhanced by Direct Preference Optimization
por: Nahian, Md Sultan Al, et al.
Publicado: (2024)
por: Nahian, Md Sultan Al, et al.
Publicado: (2024)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
por: Li, Siting, et al.
Publicado: (2024)
por: Li, Siting, et al.
Publicado: (2024)
EpilepsyLLM: Domain-Specific Large Language Model Fine-tuned with Epilepsy Medical Knowledge
por: Zhao, Xuyang, et al.
Publicado: (2024)
por: Zhao, Xuyang, et al.
Publicado: (2024)
LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?
por: Tamo, J. Ben, et al.
Publicado: (2026)
por: Tamo, J. Ben, et al.
Publicado: (2026)
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
por: Buzeta, Nicolas, et al.
Publicado: (2026)
por: Buzeta, Nicolas, et al.
Publicado: (2026)
Aggregated Knowledge Model: Enhancing Domain-Specific QA with Fine-Tuned and Retrieval-Augmented Generation Models
por: Liu, Fengchen, et al.
Publicado: (2024)
por: Liu, Fengchen, et al.
Publicado: (2024)
TeleDoCTR: Domain-Specific and Contextual Troubleshooting for Telecommunications
por: Trabelsi, Mohamed, et al.
Publicado: (2026)
por: Trabelsi, Mohamed, et al.
Publicado: (2026)
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
por: Liang, Yiqing, et al.
Publicado: (2025)
por: Liang, Yiqing, et al.
Publicado: (2025)
Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases
por: Li, Jiarui, et al.
Publicado: (2024)
por: Li, Jiarui, et al.
Publicado: (2024)
Do Large Language Models Know How Much They Know?
por: Prato, Gabriele, et al.
Publicado: (2025)
por: Prato, Gabriele, et al.
Publicado: (2025)
Enhancing Clinical Documentation with Synthetic Data: Leveraging Generative Models for Improved Accuracy
por: Biswas, Anjanava, et al.
Publicado: (2024)
por: Biswas, Anjanava, et al.
Publicado: (2024)
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
por: Gill, Waris, et al.
Publicado: (2025)
por: Gill, Waris, et al.
Publicado: (2025)
HUKUKBERT: Domain-Specific Language Model for Turkish Law
por: Öztürk, Mehmet Utku, et al.
Publicado: (2026)
por: Öztürk, Mehmet Utku, et al.
Publicado: (2026)
Enhancing Robustness in Biomedical NLI Models: A Probing Approach for Clinical Trials
por: Mustafa, Ata
Publicado: (2024)
por: Mustafa, Ata
Publicado: (2024)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
por: Yang, Chenghao, et al.
Publicado: (2025)
por: Yang, Chenghao, et al.
Publicado: (2025)
How Can Large Language Models Understand Spatial-Temporal Data?
por: Liu, Lei, et al.
Publicado: (2024)
por: Liu, Lei, et al.
Publicado: (2024)
Resilience through Scene Context in Visual Referring Expression Generation
por: Junker, Simeon, et al.
Publicado: (2024)
por: Junker, Simeon, et al.
Publicado: (2024)
How Far Can Unsupervised RLVR Scale LLM Training?
por: He, Bingxiang, et al.
Publicado: (2026)
por: He, Bingxiang, et al.
Publicado: (2026)
How to Make Large Language Models Generate 100% Valid Molecules?
por: Tao, Wen, et al.
Publicado: (2025)
por: Tao, Wen, et al.
Publicado: (2025)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
por: Hu, Lanxiang, et al.
Publicado: (2024)
por: Hu, Lanxiang, et al.
Publicado: (2024)
TURNA: A Turkish Encoder-Decoder Language Model for Enhanced Understanding and Generation
por: Uludoğan, Gökçe, et al.
Publicado: (2024)
por: Uludoğan, Gökçe, et al.
Publicado: (2024)
Enhancing Large Language Models with Domain-Specific Knowledge: The Case in Topological Materials
por: Xu, HuangChao, et al.
Publicado: (2024)
por: Xu, HuangChao, et al.
Publicado: (2024)
The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions
por: Wagner, Stefan Sylvius, et al.
Publicado: (2024)
por: Wagner, Stefan Sylvius, et al.
Publicado: (2024)
Ejemplares similares
-
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
por: Brinner, Marc, et al.
Publicado: (2025) -
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
por: Brinner, Marc, et al.
Publicado: (2025) -
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
por: Brinner, Marc, et al.
Publicado: (2025) -
Model Interpretability and Rationale Extraction by Input Mask Optimization
por: Brinner, Marc, et al.
Publicado: (2025) -
Efficient Scientific Full Text Classification: The Case of EICAT Impact Assessments
por: Brinner, Marc Felix, et al.
Publicado: (2025)