BgGPT 1.0: Extending English-centric LLMs to other languages
Fuente:
arXiv
Saved in:
| Main Authors: | Alexandrov, Anton, Raychev, Veselin, Dimitrov, Dimitar I., Zhang, Ce, Vechev, Martin, Toutanova, Kristina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Catastrophic Forgetting in Language Transfer via Model Merging
by: Alexandrov, Anton, et al.
Published: (2024)
by: Alexandrov, Anton, et al.
Published: (2024)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
by: Thillen, Alex, et al.
Published: (2026)
by: Thillen, Alex, et al.
Published: (2026)
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
by: Petrov, Ivo, et al.
Published: (2026)
by: Petrov, Ivo, et al.
Published: (2026)
DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models
by: Berabi, Berkay, et al.
Published: (2024)
by: Berabi, Berkay, et al.
Published: (2024)
Modular Synthesis of Efficient Quantum Uncomputation
by: Venev, Hristo, et al.
Published: (2024)
by: Venev, Hristo, et al.
Published: (2024)
BaxBench: Can LLMs Generate Correct and Secure Backends?
by: Vero, Mark, et al.
Published: (2025)
by: Vero, Mark, et al.
Published: (2025)
Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
by: Yukhymenko, Hanna, et al.
Published: (2026)
by: Yukhymenko, Hanna, et al.
Published: (2026)
Customizing Static Analysis using Codesearch
by: Hayoun, Avi, et al.
Published: (2024)
by: Hayoun, Avi, et al.
Published: (2024)
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
Coding Agents Don't Know When to Act
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
by: Dékány, Csaba, et al.
Published: (2025)
by: Dékány, Csaba, et al.
Published: (2025)
Hiding in Plain Sight: Disguising Data Stealing Attacks in Federated Learning
by: Garov, Kostadin, et al.
Published: (2023)
by: Garov, Kostadin, et al.
Published: (2023)
SPEAR++: Scaling Gradient Inversion via Sparsely-Used Dictionary Learning
by: Bakarsky, Alexander, et al.
Published: (2025)
by: Bakarsky, Alexander, et al.
Published: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
SPEAR:Exact Gradient Inversion of Batches in Federated Learning
by: Dimitrov, Dimitar I., et al.
Published: (2024)
by: Dimitrov, Dimitar I., et al.
Published: (2024)
Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
by: Shaw, Peter, et al.
Published: (2025)
by: Shaw, Peter, et al.
Published: (2025)
A Unified Approach to Routing and Cascading for LLMs
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
by: Mündler, Niels, et al.
Published: (2025)
by: Mündler, Niels, et al.
Published: (2025)
GRAIN: Exact Graph Reconstruction from Gradients
by: Drencheva, Maria, et al.
Published: (2025)
by: Drencheva, Maria, et al.
Published: (2025)
Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
by: Kew, Tannon, et al.
Published: (2023)
by: Kew, Tannon, et al.
Published: (2023)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
by: Petrov, Ivo, et al.
Published: (2024)
by: Petrov, Ivo, et al.
Published: (2024)
Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
by: Beurer-Kellner, Luca, et al.
Published: (2024)
by: Beurer-Kellner, Luca, et al.
Published: (2024)
Negation in English and other languages
by: Jespersen, Otto, et al.
Published: (2026)
by: Jespersen, Otto, et al.
Published: (2026)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
by: Balauca, Ada-Astrid, et al.
Published: (2024)
by: Balauca, Ada-Astrid, et al.
Published: (2024)
FMI_SU_Yotkova_Kastreva at SemEval-2026 Task 13: Lightweight Detection of LLM-Generated Code via Stylometric Signals
by: Yotkova, Elitsa, et al.
Published: (2026)
by: Yotkova, Elitsa, et al.
Published: (2026)
A record of Achaearanea tabulata from the Balkan Peninsula (Araneae: Theridiidae)
by: Dimitrov, Dimitar
Published: (1994)
by: Dimitrov, Dimitar
Published: (1994)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
by: Zhu, Wang, et al.
Published: (2023)
by: Zhu, Wang, et al.
Published: (2023)
ALTA: Compiler-Based Analysis of Transformers
by: Shaw, Peter, et al.
Published: (2024)
by: Shaw, Peter, et al.
Published: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
Dissecting Paraphrases: The Impact of Prompt Syntax and supplementary Information on Knowledge Retrieval from Pretrained Language Models
by: Linzbach, Stephan, et al.
Published: (2024)
by: Linzbach, Stephan, et al.
Published: (2024)
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
by: Hiss, Hanno, et al.
Published: (2026)
by: Hiss, Hanno, et al.
Published: (2026)
Large Language Models for Code: Security Hardening and Adversarial Testing
by: He, Jingxuan, et al.
Published: (2023)
by: He, Jingxuan, et al.
Published: (2023)
Can Grammarly and ChatGPT accelerate language change? AI-powered technologies and their impact on the English language: wordiness vs. conciseness
by: Rudnicka, Karolina
Published: (2025)
by: Rudnicka, Karolina
Published: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
CueBuddy: helping non-native English speakers navigate English-centric STEM education
by: Gupta, Pranav
Published: (2025)
by: Gupta, Pranav
Published: (2025)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024)
by: Beshirov, Angel, et al.
Published: (2024)
ConStat: Performance-Based Contamination Detection in Large Language Models
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
Similar Items
-
Mitigating Catastrophic Forgetting in Language Transfer via Model Merging
by: Alexandrov, Anton, et al.
Published: (2024) -
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
by: Thillen, Alex, et al.
Published: (2026) -
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
by: Petrov, Ivo, et al.
Published: (2026) -
DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models
by: Berabi, Berkay, et al.
Published: (2024) -
Modular Synthesis of Efficient Quantum Uncomputation
by: Venev, Hristo, et al.
Published: (2024)