From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ascione, Grazia Sveva, Tamagnone, Nicolò |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A comparative analysis of embedding models for patent similarity
von: Ascione, Grazia Sveva, et al.
Veröffentlicht: (2024)
von: Ascione, Grazia Sveva, et al.
Veröffentlicht: (2024)
Presenting Terrorizer: an algorithm for consolidating company names in patent assignees
von: Ascione, Grazia Sveva, et al.
Veröffentlicht: (2024)
von: Ascione, Grazia Sveva, et al.
Veröffentlicht: (2024)
Tell me the truth: A system to measure the trustworthiness of Large Language Models
von: Lipizzi, Carlo
Veröffentlicht: (2024)
von: Lipizzi, Carlo
Veröffentlicht: (2024)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
von: Dong, Peijie, et al.
Veröffentlicht: (2024)
von: Dong, Peijie, et al.
Veröffentlicht: (2024)
Matching domain experts by training from scratch on domain knowledge
von: Luo, Xiaoliang, et al.
Veröffentlicht: (2024)
von: Luo, Xiaoliang, et al.
Veröffentlicht: (2024)
Evaluating the Performance of Large Language Models for SDG Mapping (Technical Report)
von: Yin, Hui, et al.
Veröffentlicht: (2024)
von: Yin, Hui, et al.
Veröffentlicht: (2024)
On the performativity of SDG classifications in large bibliometric databases
von: Ottaviani, Matteo, et al.
Veröffentlicht: (2024)
von: Ottaviani, Matteo, et al.
Veröffentlicht: (2024)
Creating Suspenseful Stories: Iterative Planning with Large Language Models
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
Investigating the translation capabilities of Large Language Models trained on parallel data only
von: Gilabert, Javier García, et al.
Veröffentlicht: (2024)
von: Gilabert, Javier García, et al.
Veröffentlicht: (2024)
Evaluating open-source Large Language Models for automated fact-checking
von: Fontana, Nicolo', et al.
Veröffentlicht: (2025)
von: Fontana, Nicolo', et al.
Veröffentlicht: (2025)
Scrambled text: training Language Models to correct OCR errors using synthetic data
von: Bourne, Jonathan
Veröffentlicht: (2024)
von: Bourne, Jonathan
Veröffentlicht: (2024)
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2024)
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2024)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
von: Yang, Lei, et al.
Veröffentlicht: (2025)
von: Yang, Lei, et al.
Veröffentlicht: (2025)
Past Meets Present: Creating Historical Analogy with Large Language Models
von: Li, Nianqi, et al.
Veröffentlicht: (2024)
von: Li, Nianqi, et al.
Veröffentlicht: (2024)
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
von: Pieler, Michael, et al.
Veröffentlicht: (2024)
von: Pieler, Michael, et al.
Veröffentlicht: (2024)
Self-training Large Language Models through Knowledge Detection
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2024)
Model Merging in Pre-training of Large Language Models
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
From Robustness to Improved Generalization and Calibration in Pre-trained Language Models
von: Jukić, Josip, et al.
Veröffentlicht: (2024)
von: Jukić, Josip, et al.
Veröffentlicht: (2024)
Can Large Language Models Create New Knowledge for Spatial Reasoning Tasks?
von: Greatrix, Thomas, et al.
Veröffentlicht: (2024)
von: Greatrix, Thomas, et al.
Veröffentlicht: (2024)
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
von: Yano, Kazuki, et al.
Veröffentlicht: (2025)
von: Yano, Kazuki, et al.
Veröffentlicht: (2025)
Cross-layer Attention Sharing for Pre-trained Large Language Models
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
Examining Forgetting in Continual Pre-training of Aligned Large Language Models
von: Li, Chen-An, et al.
Veröffentlicht: (2024)
von: Li, Chen-An, et al.
Veröffentlicht: (2024)
Comparing Code Explanations Created by Students and Large Language Models
von: Leinonen, Juho, et al.
Veröffentlicht: (2023)
von: Leinonen, Juho, et al.
Veröffentlicht: (2023)
A Survey on Post-training of Large Language Models
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
von: Malin, Ben, et al.
Veröffentlicht: (2025)
von: Malin, Ben, et al.
Veröffentlicht: (2025)
Large Language Models for Biomedical Text Simplification: Promising But Not There Yet
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
von: Tu, Zhijun, et al.
Veröffentlicht: (2025)
von: Tu, Zhijun, et al.
Veröffentlicht: (2025)
OntoTune: Ontology-Driven Self-training for Aligning Large Language Models
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
Efficient Continual Pre-training for Building Domain Specific Large Language Models
von: Xie, Yong, et al.
Veröffentlicht: (2023)
von: Xie, Yong, et al.
Veröffentlicht: (2023)
Language evolution 'in silico': From large-scale data to artificial agents creating languages from scratch
von: Thomas Brochhagen
Veröffentlicht: (2025)
von: Thomas Brochhagen
Veröffentlicht: (2025)
Pre-trained Large Language Models for Financial Sentiment Analysis
von: Luo, Wei, et al.
Veröffentlicht: (2024)
von: Luo, Wei, et al.
Veröffentlicht: (2024)
Spike No More: Stabilizing the Pre-training of Large Language Models
von: Takase, Sho, et al.
Veröffentlicht: (2023)
von: Takase, Sho, et al.
Veröffentlicht: (2023)
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
SongSage: A Large Musical Language Model with Lyric Generative Pre-training
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
Rational Metareasoning for Large Language Models
von: De Sabbata, C. Nicolò, et al.
Veröffentlicht: (2024)
von: De Sabbata, C. Nicolò, et al.
Veröffentlicht: (2024)
How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
von: Elkins, Sabina, et al.
Veröffentlicht: (2024)
von: Elkins, Sabina, et al.
Veröffentlicht: (2024)
Investigating Large Language Models and Control Mechanisms to Improve Text Readability of Biomedical Abstracts
von: Li, Zihao, et al.
Veröffentlicht: (2023)
von: Li, Zihao, et al.
Veröffentlicht: (2023)
From N-grams to Pre-trained Multilingual Models For Language Identification
von: Sindane, Thapelo, et al.
Veröffentlicht: (2024)
von: Sindane, Thapelo, et al.
Veröffentlicht: (2024)
Modelling Language using Large Language Models
von: Grindrod, Jumbly
Veröffentlicht: (2024)
von: Grindrod, Jumbly
Veröffentlicht: (2024)
Ähnliche Einträge
-
A comparative analysis of embedding models for patent similarity
von: Ascione, Grazia Sveva, et al.
Veröffentlicht: (2024) -
Presenting Terrorizer: an algorithm for consolidating company names in patent assignees
von: Ascione, Grazia Sveva, et al.
Veröffentlicht: (2024) -
Tell me the truth: A system to measure the trustworthiness of Large Language Models
von: Lipizzi, Carlo
Veröffentlicht: (2024) -
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
von: Dong, Peijie, et al.
Veröffentlicht: (2024) -
Matching domain experts by training from scratch on domain knowledge
von: Luo, Xiaoliang, et al.
Veröffentlicht: (2024)