Smart Bilingual Focused Crawling of Parallel Documents
Fuente:
arXiv
Saved in:
| Main Authors: | García-Romero, Cristian, Esplà-Gomis, Miquel, Sánchez-Martínez, Felipe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
by: García-Romero, Cristian, et al.
Published: (2025)
by: García-Romero, Cristian, et al.
Published: (2025)
A Simple Approach to Use Bilingual Information Sources for Word Alignment
by: Miquel Esplà-Gomis
Published: (2012)
by: Miquel Esplà-Gomis
Published: (2012)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
ThreatCrawl: A BERT-based Focused Crawler for the Cybersecurity Domain
by: Kuehn, Philipp, et al.
Published: (2023)
by: Kuehn, Philipp, et al.
Published: (2023)
Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation
by: Esplà-Gomis, Miquel, et al.
Published: (2024)
by: Esplà-Gomis, Miquel, et al.
Published: (2024)
Non-Fluent Synthetic Target-Language Data Improve Neural Machine Translation
by: Sánchez-Cartagena, Víctor M., et al.
Published: (2024)
by: Sánchez-Cartagena, Víctor M., et al.
Published: (2024)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders
by: Inaba, Tatsuro, et al.
Published: (2025)
by: Inaba, Tatsuro, et al.
Published: (2025)
Semi-Supervised Learning for Bilingual Lexicon Induction
by: Garnier, Paul, et al.
Published: (2024)
by: Garnier, Paul, et al.
Published: (2024)
Kanana: Compute-efficient Bilingual Language Models
by: Kanana LLM Team, et al.
Published: (2025)
by: Kanana LLM Team, et al.
Published: (2025)
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024)
by: Seto, Skyler, et al.
Published: (2024)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
by: Ruder, Sebastian, et al.
Published: (2018)
by: Ruder, Sebastian, et al.
Published: (2018)
CroissantLLM: A Truly Bilingual French-English Language Model
by: Faysse, Manuel, et al.
Published: (2024)
by: Faysse, Manuel, et al.
Published: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding
by: Yi, Hanling, et al.
Published: (2024)
by: Yi, Hanling, et al.
Published: (2024)
Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space
by: Rom, Aviad, et al.
Published: (2024)
by: Rom, Aviad, et al.
Published: (2024)
Modeling Bilingual Sentence Processing: Evaluating RNN and Transformer Architectures for Cross-Language Structural Priming
by: Zhang, Demi, et al.
Published: (2024)
by: Zhang, Demi, et al.
Published: (2024)
The Impact of Language Mixing on Bilingual LLM Reasoning
by: Li, Yihao, et al.
Published: (2025)
by: Li, Yihao, et al.
Published: (2025)
Curated Datasets and Neural Models for Machine Translation of Informal Registers between Mayan and Spanish Vernaculars
by: Lou, Andrés, et al.
Published: (2024)
by: Lou, Andrés, et al.
Published: (2024)
Logit Reweighting for Topic-Focused Summarization
by: Braun, Joschka, et al.
Published: (2025)
by: Braun, Joschka, et al.
Published: (2025)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
Query-Focused Extractive Summarization for Sentiment Explanation
by: Moubtahij, Ahmed, et al.
Published: (2025)
by: Moubtahij, Ahmed, et al.
Published: (2025)
Gumbel Distillation for Parallel Text Generation
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
Parallel Scaling Law for Language Models
by: Chen, Mouxiang, et al.
Published: (2025)
by: Chen, Mouxiang, et al.
Published: (2025)
Parallel Token Prediction for Language Models
by: Draxler, Felix, et al.
Published: (2025)
by: Draxler, Felix, et al.
Published: (2025)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
by: Ram, Dhananjay, et al.
Published: (2025)
by: Ram, Dhananjay, et al.
Published: (2025)
Document Summarization with Conformal Importance Guarantees
by: Kuwahara, Bruce, et al.
Published: (2025)
by: Kuwahara, Bruce, et al.
Published: (2025)
Meta4XNLI: A Crosslingual Parallel Corpus for Metaphor Detection and Interpretation
by: Sanchez-Bayona, Elisa, et al.
Published: (2024)
by: Sanchez-Bayona, Elisa, et al.
Published: (2024)
Plan Optimization to Bilingual Dictionary Induction for Low-Resource Language Families
by: Nasution, Arbi Haza, et al.
Published: (2020)
by: Nasution, Arbi Haza, et al.
Published: (2020)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
by: Badger, Benjamin L.
Published: (2026)
by: Badger, Benjamin L.
Published: (2026)
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
by: Buzeta, Nicolas, et al.
Published: (2026)
by: Buzeta, Nicolas, et al.
Published: (2026)
MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series
by: Zhang, Ge, et al.
Published: (2024)
by: Zhang, Ge, et al.
Published: (2024)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025)
by: Rodionov, Gleb, et al.
Published: (2025)
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
PaPaformer: Language Model from Pre-trained Parallel Paths
by: Tapaninaho, Joonas, et al.
Published: (2025)
by: Tapaninaho, Joonas, et al.
Published: (2025)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
by: Sivtsov, Danil, et al.
Published: (2025)
by: Sivtsov, Danil, et al.
Published: (2025)
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
Teaching Large Language Models Number-Focused Headline Generation With Key Element Rationales
by: Qian, Zhen, et al.
Published: (2025)
by: Qian, Zhen, et al.
Published: (2025)
On Bilingual Lexicon Induction with Large Language Models
by: Li, Yaoyiran, et al.
Published: (2023)
by: Li, Yaoyiran, et al.
Published: (2023)
Similar Items
-
Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
by: García-Romero, Cristian, et al.
Published: (2025) -
A Simple Approach to Use Bilingual Information Sources for Word Alignment
by: Miquel Esplà-Gomis
Published: (2012) -
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024) -
ThreatCrawl: A BERT-based Focused Crawler for the Cybersecurity Domain
by: Kuehn, Philipp, et al.
Published: (2023) -
Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation
by: Esplà-Gomis, Miquel, et al.
Published: (2024)