Bilingual Adaptation of Monolingual Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gosal, Gurpreet, Xu, Yishi, Ramakrishnan, Gokul, Joshi, Rituraj, Sheinin, Avraham, Zhiming, Chen, Mishra, Biswajit, Vassilieva, Natalia, Hestness, Joel, Sengupta, Neha, Sahu, Sunil Kumar, Jia, Bokang, Pandit, Onkar, Katipomu, Satheesh, Kamboj, Samta, Ghosh, Samujjwal, Pal, Rahul, Mullah, Parvez, Doraiswamy, Soundar, Chami, Mohamed El Karim, Nakov, Preslav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
von: Goffinet, Etienne, et al.
Veröffentlicht: (2025)
von: Goffinet, Etienne, et al.
Veröffentlicht: (2025)
Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi
von: Choudhury, Monojit, et al.
Veröffentlicht: (2025)
von: Choudhury, Monojit, et al.
Veröffentlicht: (2025)
Nomad: Autonomous Exploration and Discovery
von: Jia, Bokang, et al.
Veröffentlicht: (2026)
von: Jia, Bokang, et al.
Veröffentlicht: (2026)
Sherkala-Chat: Building a State-of-the-Art LLM for Kazakh in a Moderately Resourced Setting
von: Koto, Fajri, et al.
Veröffentlicht: (2025)
von: Koto, Fajri, et al.
Veröffentlicht: (2025)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2025)
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2025)
The Information Search
von: Doraiswamy, Uma
Veröffentlicht: (2011)
von: Doraiswamy, Uma
Veröffentlicht: (2011)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
von: Sundriyal, Megha, et al.
Veröffentlicht: (2023)
von: Sundriyal, Megha, et al.
Veröffentlicht: (2023)
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition
von: Alhindi, Tariq, et al.
Veröffentlicht: (2023)
von: Alhindi, Tariq, et al.
Veröffentlicht: (2023)
Rethinking STS and NLI in Large Language Models
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
DenoiseRank: Learning to Rank by Diffusion Models
von: Wang, Ying, et al.
Veröffentlicht: (2026)
von: Wang, Ying, et al.
Veröffentlicht: (2026)
Adapting Fake News Detection to the Era of Large Language Models
von: Su, Jinyan, et al.
Veröffentlicht: (2023)
von: Su, Jinyan, et al.
Veröffentlicht: (2023)
Corpus Poisoning via Approximate Greedy Gradient Descent
von: Su, Jinyan, et al.
Veröffentlicht: (2024)
von: Su, Jinyan, et al.
Veröffentlicht: (2024)
Work‐Related Migration and Couples in Long‐Distance Marriages: Mindfulness, Marital Quality, Satisfaction, and Happiness
von: Samta P. Pandya
Veröffentlicht: (2025)
von: Samta P. Pandya
Veröffentlicht: (2025)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2025)
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2025)
How Does Prefix Matter in Reasoning Model Tuning?
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2026)
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2026)
Оригинал и карикатура: как британская «профилактическая юстиция» (preventive justice) не приживается на континенте Оторванность от родной почвы выявляет полный абсурд
von: Vassilieva, Elena
Veröffentlicht: (2025)
von: Vassilieva, Elena
Veröffentlicht: (2025)
Создание «уязвимых групп» как механизм обхода правовых гарантий и Прав Человека - Кому выгодно «НЕравенство всех перед судом»?
von: Vassilieva, Elena
Veröffentlicht: (2025)
von: Vassilieva, Elena
Veröffentlicht: (2025)
Creating "Vulnerable Groups" as a Mechanism to Circumvent Legal Guarantees and Human Rights - Who Benefits from "INequality Before the Law"?
von: Vassilieva, Elena
Veröffentlicht: (2025)
von: Vassilieva, Elena
Veröffentlicht: (2025)
Британская модель профилактической юстиции (preventive justice) как новая форма правового подчинения (Claim of Originality)
von: Vassilieva, Elena
Veröffentlicht: (2025)
von: Vassilieva, Elena
Veröffentlicht: (2025)
Claim 5 - The Blind Spot of British Law: Judicial Criminalisation Without Realising It
von: Vassilieva, Elena
Veröffentlicht: (2025)
von: Vassilieva, Elena
Veröffentlicht: (2025)
От «защиты жертв» к системному вредительству: Превентивный надзор как закрепленный в законе обход процессуальных норм
von: Vassilieva, Elena
Veröffentlicht: (2025)
von: Vassilieva, Elena
Veröffentlicht: (2025)
The Impact of Economic Policy Uncertainty and ESG Reporting on Financial Performance of Hospitality Companies
von: Samta Jain, et al.
Veröffentlicht: (2025)
von: Samta Jain, et al.
Veröffentlicht: (2025)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
von: Bates, Luke, et al.
Veröffentlicht: (2025)
von: Bates, Luke, et al.
Veröffentlicht: (2025)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
von: Elozeiri, Kareem, et al.
Veröffentlicht: (2025)
von: Elozeiri, Kareem, et al.
Veröffentlicht: (2025)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
von: Thareja, Rushil, et al.
Veröffentlicht: (2025)
von: Thareja, Rushil, et al.
Veröffentlicht: (2025)
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
von: Vashurin, Roman, et al.
Veröffentlicht: (2025)
von: Vashurin, Roman, et al.
Veröffentlicht: (2025)
Missci: Reconstructing Fallacies in Misrepresented Science
von: Glockner, Max, et al.
Veröffentlicht: (2024)
von: Glockner, Max, et al.
Veröffentlicht: (2024)
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
von: Ivanov, Petar, et al.
Veröffentlicht: (2023)
von: Ivanov, Petar, et al.
Veröffentlicht: (2023)
Multimodal Large Language Models to Support Real-World Fact-Checking
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
von: Glockner, Max, et al.
Veröffentlicht: (2024)
von: Glockner, Max, et al.
Veröffentlicht: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
von: Agarwal, Siddhant, et al.
Veröffentlicht: (2024)
von: Agarwal, Siddhant, et al.
Veröffentlicht: (2024)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
Fast or Better? Balancing Accuracy and Cost in Retrieval-Augmented Generation with Flexible User Control
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
von: Goffinet, Etienne, et al.
Veröffentlicht: (2025) -
Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi
von: Choudhury, Monojit, et al.
Veröffentlicht: (2025) -
Nomad: Autonomous Exploration and Discovery
von: Jia, Bokang, et al.
Veröffentlicht: (2026) -
Sherkala-Chat: Building a State-of-the-Art LLM for Kazakh in a Moderately Resourced Setting
von: Koto, Fajri, et al.
Veröffentlicht: (2025) -
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)