Curriculum Recommendations Using Transformer Base Model with InfoNCE Loss And Language Switching Method
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Xiaonan, Yuan, Bin, Mo, Yongyao, Song, Tianbo, Li, Shulin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
di: Li, Yongchang, et al.
Pubblicazione: (2024)
di: Li, Yongchang, et al.
Pubblicazione: (2024)
Optimization Strategies for Enhancing Resource Efficiency in Transformers & Large Language Models
di: Wallace, Tom, et al.
Pubblicazione: (2025)
di: Wallace, Tom, et al.
Pubblicazione: (2025)
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024)
di: Goldin, Gili, et al.
Pubblicazione: (2024)
InfoNCE Induces Gaussian Distribution
di: Betser, Roy, et al.
Pubblicazione: (2026)
di: Betser, Roy, et al.
Pubblicazione: (2026)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
di: Land, Sander, et al.
Pubblicazione: (2024)
di: Land, Sander, et al.
Pubblicazione: (2024)
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2025)
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2025)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
di: Yao, Huaiyuan, et al.
Pubblicazione: (2026)
di: Yao, Huaiyuan, et al.
Pubblicazione: (2026)
Contrastive Learning of Preferences with a Contextual InfoNCE Loss
di: Bertram, Timo, et al.
Pubblicazione: (2024)
di: Bertram, Timo, et al.
Pubblicazione: (2024)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
di: Andrylie, Lyzander Marciano, et al.
Pubblicazione: (2025)
di: Andrylie, Lyzander Marciano, et al.
Pubblicazione: (2025)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
di: Chen, Jiaju, et al.
Pubblicazione: (2025)
di: Chen, Jiaju, et al.
Pubblicazione: (2025)
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2023)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2023)
HInter: Exposing Hidden Intersectional Bias in Large Language Models
di: Souani, Badr, et al.
Pubblicazione: (2025)
di: Souani, Badr, et al.
Pubblicazione: (2025)
Transforming and Combining Rewards for Aligning Large Language Models
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
T-VEC: A Telecom-Specific Vectorization Model with Enhanced Semantic Understanding via Deep Triplet Loss Fine-Tuning
di: Ethiraj, Vignesh, et al.
Pubblicazione: (2025)
di: Ethiraj, Vignesh, et al.
Pubblicazione: (2025)
IteRABRe: Iterative Recovery-Aided Block Reduction
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2025)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2025)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
di: Li, Mingda, et al.
Pubblicazione: (2024)
di: Li, Mingda, et al.
Pubblicazione: (2024)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
di: Ikoma, Hayato, et al.
Pubblicazione: (2025)
di: Ikoma, Hayato, et al.
Pubblicazione: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models
di: Gouliev, Zaur, et al.
Pubblicazione: (2025)
di: Gouliev, Zaur, et al.
Pubblicazione: (2025)
A Legal Framework for Natural Language Processing Model Training in Portugal
di: Almeida, Rúben, et al.
Pubblicazione: (2024)
di: Almeida, Rúben, et al.
Pubblicazione: (2024)
A Primer on Large Language Models and their Limitations
di: Johnson, Sandra, et al.
Pubblicazione: (2024)
di: Johnson, Sandra, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Public Health Classification and Extraction Tasks
di: Harris, Joshua, et al.
Pubblicazione: (2024)
di: Harris, Joshua, et al.
Pubblicazione: (2024)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
di: Truong, Sang T., et al.
Pubblicazione: (2024)
di: Truong, Sang T., et al.
Pubblicazione: (2024)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
Evaluating Pixel Language Models on Non-Standardized Languages
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2024)
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2024)
Review GIDE -- Restaurant Review Gastrointestinal Illness Detection and Extraction with Large Language Models
di: Laurence, Timothy, et al.
Pubblicazione: (2025)
di: Laurence, Timothy, et al.
Pubblicazione: (2025)
EvidenceMap: Learning Evidence Analysis to Unleash the Power of Small Language Models for Biomedical Question Answering
di: Zong, Chang, et al.
Pubblicazione: (2025)
di: Zong, Chang, et al.
Pubblicazione: (2025)
Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
di: Fang, Xi, et al.
Pubblicazione: (2024)
di: Fang, Xi, et al.
Pubblicazione: (2024)
ProSwitch: Knowledge-Guided Instruction Tuning to Switch Between Professional and Non-Professional Responses
di: Zong, Chang, et al.
Pubblicazione: (2024)
di: Zong, Chang, et al.
Pubblicazione: (2024)
The Superalignment of Superhuman Intelligence with Large Language Models
di: Huang, Minlie, et al.
Pubblicazione: (2024)
di: Huang, Minlie, et al.
Pubblicazione: (2024)
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models
di: Safavi-Naini, Seyed Amir Ahmad, et al.
Pubblicazione: (2024)
di: Safavi-Naini, Seyed Amir Ahmad, et al.
Pubblicazione: (2024)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
di: Platt, Nolan, et al.
Pubblicazione: (2025)
di: Platt, Nolan, et al.
Pubblicazione: (2025)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
di: Yu, Haeun, et al.
Pubblicazione: (2024)
di: Yu, Haeun, et al.
Pubblicazione: (2024)
Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations
di: Bhattacharya, Shamik, et al.
Pubblicazione: (2026)
di: Bhattacharya, Shamik, et al.
Pubblicazione: (2026)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation
di: Li, Xuan, et al.
Pubblicazione: (2026)
di: Li, Xuan, et al.
Pubblicazione: (2026)
Which Pieces Does Unigram Tokenization Really Need?
di: Land, Sander, et al.
Pubblicazione: (2025)
di: Land, Sander, et al.
Pubblicazione: (2025)
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2026)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2026)
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
Documenti analoghi
-
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
di: Li, Yongchang, et al.
Pubblicazione: (2024) -
Optimization Strategies for Enhancing Resource Efficiency in Transformers & Large Language Models
di: Wallace, Tom, et al.
Pubblicazione: (2025) -
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024) -
InfoNCE Induces Gaussian Distribution
di: Betser, Roy, et al.
Pubblicazione: (2026) -
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
di: Land, Sander, et al.
Pubblicazione: (2024)