Saved in:
| Main Authors: | Lazzaroni, Ruggero Marino, Lasser, Jana, Solovev, Kirill |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.18337 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
by: Lazzaroni, Ruggero Marino, et al.
Published: (2025)
by: Lazzaroni, Ruggero Marino, et al.
Published: (2025)
Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO
by: Pappone, Francesco, et al.
Published: (2025)
by: Pappone, Francesco, et al.
Published: (2025)
Quantifying Geospatial in the Common Crawl Corpus
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
by: Tessema, Bethel Melesse, et al.
Published: (2024)
by: Tessema, Bethel Melesse, et al.
Published: (2024)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024)
by: Kim, Minsoo, et al.
Published: (2024)
CC-GPX: Extracting High-Quality Annotated Geospatial Data from Common Crawl
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
by: Su, Dan, et al.
Published: (2024)
by: Su, Dan, et al.
Published: (2024)
SODIUM: From Open Web Data to Queryable Databases
by: Hu, Chuxuan, et al.
Published: (2026)
by: Hu, Chuxuan, et al.
Published: (2026)
Facts are Harder Than Opinions -- A Multilingual, Comparative Analysis of LLM-Based Fact-Checking Reliability
by: Saju, Lorraine, et al.
Published: (2025)
by: Saju, Lorraine, et al.
Published: (2025)
Decoding News Bias: Multi Bias Detection in News Articles
by: Shah, Bhushan Santosh, et al.
Published: (2025)
by: Shah, Bhushan Santosh, et al.
Published: (2025)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
Craw4LLM: Efficient Web Crawling for LLM Pretraining
by: Yu, Shi, et al.
Published: (2025)
by: Yu, Shi, et al.
Published: (2025)
Scalable Detection of Salient Entities in News Articles
by: Asgarieh, Eliyar, et al.
Published: (2024)
by: Asgarieh, Eliyar, et al.
Published: (2024)
Dataset of Quotation Attribution in German News Articles
by: Petersen-Frey, Fynn, et al.
Published: (2024)
by: Petersen-Frey, Fynn, et al.
Published: (2024)
Explaining Mixtures of Sources in News Articles
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
by: Gienapp, Lukas, et al.
Published: (2025)
by: Gienapp, Lukas, et al.
Published: (2025)
The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation Project
by: Aquino, Angelina A., et al.
Published: (2025)
by: Aquino, Angelina A., et al.
Published: (2025)
A Multilingual Similarity Dataset for News Article Frame
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
EMONA: Event-level Moral Opinions in News Articles
by: Lei, Yuanyuan, et al.
Published: (2024)
by: Lei, Yuanyuan, et al.
Published: (2024)
Headline-Guided Extractive Summarization for Thai News Articles
by: Kositcharoensuk, Pimpitchaya, et al.
Published: (2024)
by: Kositcharoensuk, Pimpitchaya, et al.
Published: (2024)
Urdu News Article Recommendation Model using Natural Language Processing Techniques
by: Abbas, Syed Zain, et al.
Published: (2022)
by: Abbas, Syed Zain, et al.
Published: (2022)
Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
by: Vaidya, Omatharv Bharat, et al.
Published: (2026)
Neutralizing the Narrative: AI-Powered Debiasing of Online News Articles
by: Kuo, Chen Wei, et al.
Published: (2025)
by: Kuo, Chen Wei, et al.
Published: (2025)
Fine-grained Narrative Classification in Biased News Articles
by: Afroz, Zeba, et al.
Published: (2025)
by: Afroz, Zeba, et al.
Published: (2025)
MuSaRoNews: A Multidomain, Multimodal Satire Dataset from Romanian News Articles
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Free Access to World News: Reconstructing Full-Text Articles from GDELT
by: Colladon, A. Fronzetti, et al.
Published: (2025)
by: Colladon, A. Fronzetti, et al.
Published: (2025)
ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
by: Campos, Ricardo, et al.
Published: (2026)
by: Campos, Ricardo, et al.
Published: (2026)
20min-XD: A Comparable Corpus of Swiss News Articles
by: Wastl, Michelle, et al.
Published: (2025)
by: Wastl, Michelle, et al.
Published: (2025)
A Corpus for Sentence-level Subjectivity Detection on English News Articles
by: Antici, Francesco, et al.
Published: (2023)
by: Antici, Francesco, et al.
Published: (2023)
InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model
by: Ouyang, Siqi, et al.
Published: (2025)
by: Ouyang, Siqi, et al.
Published: (2025)
A Regularized LSTM Method for Detecting Fake News Articles
by: Camelia, Tanjina Sultana, et al.
Published: (2024)
by: Camelia, Tanjina Sultana, et al.
Published: (2024)
Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
The 2021 Tokyo Olympics Multilingual News Article Dataset
by: Novak, Erik, et al.
Published: (2025)
by: Novak, Erik, et al.
Published: (2025)
Smart Bilingual Focused Crawling of Parallel Documents
by: García-Romero, Cristian, et al.
Published: (2024)
by: García-Romero, Cristian, et al.
Published: (2024)
A Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
by: Ibrahim, Hazem, et al.
Published: (2024)
by: Ibrahim, Hazem, et al.
Published: (2024)
Dataset of News Articles with Provenance Metadata for Media Relevance Assessment
by: Peterka, Tomas, et al.
Published: (2025)
by: Peterka, Tomas, et al.
Published: (2025)
Leveraging Web-Crawled Data for High-Quality Fine-Tuning
by: Zhou, Jing, et al.
Published: (2024)
by: Zhou, Jing, et al.
Published: (2024)
Similar Items
-
MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
by: Lazzaroni, Ruggero Marino, et al.
Published: (2025) -
Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO
by: Pappone, Francesco, et al.
Published: (2025) -
Quantifying Geospatial in the Common Crawl Corpus
by: Ilyankou, Ilya, et al.
Published: (2024) -
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
by: Tessema, Bethel Melesse, et al.
Published: (2024) -
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024)