Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Dongyang, Sabolčec, Vinko, Ansaripour, Matin, Tarun, Ayush Kumar, Jaggi, Martin, Bosselut, Antoine, Schlag, Imanol |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
by: Xu, Yixuan, et al.
Published: (2025)
by: Xu, Yixuan, et al.
Published: (2025)
Towards Fully FP8 GEMM LLM Training at Scale
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
by: Allard, Marc-Antoine, et al.
Published: (2024)
by: Allard, Marc-Antoine, et al.
Published: (2024)
Gabow's Cardinality Matching Algorithm in General Graphs: Implementation and Experiments
by: Ansaripour, Matin, et al.
Published: (2024)
by: Ansaripour, Matin, et al.
Published: (2024)
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
by: Penedo, Guilherme, et al.
Published: (2025)
by: Penedo, Guilherme, et al.
Published: (2025)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
by: Talokar, Nivya, et al.
Published: (2026)
by: Talokar, Nivya, et al.
Published: (2026)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Deriving Activation Functions Using Integration
by: Huang, Allen Hao, et al.
Published: (2024)
by: Huang, Allen Hao, et al.
Published: (2024)
Movie Recommendation using Web Crawling
by: Raj, Pronit, et al.
Published: (2024)
by: Raj, Pronit, et al.
Published: (2024)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
by: Ansaripour, Matin, et al.
Published: (2022)
by: Ansaripour, Matin, et al.
Published: (2022)
Approximate EFX and Exact tEFX Allocations for Indivisible Chores: Improved Algorithms
by: Afshinmehr, Mahyar, et al.
Published: (2024)
by: Afshinmehr, Mahyar, et al.
Published: (2024)
IDSIA/recurrent-fwp: v1.0.0 - Public Code Release
by: Kazuki Irie, et al.
Published: (2025)
by: Kazuki Irie, et al.
Published: (2025)
Efficient Crawling for Scalable Web Data Acquisition (Extended Version)
by: Gauquier, Antoine, et al.
Published: (2026)
by: Gauquier, Antoine, et al.
Published: (2026)
Web Page Classification using LLMs for Crawling Support
by: Sasazawa, Yuichi, et al.
Published: (2025)
by: Sasazawa, Yuichi, et al.
Published: (2025)
TiMoE: Time-Aware Mixture of Language Experts
by: Faro, Robin, et al.
Published: (2025)
by: Faro, Robin, et al.
Published: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Quantifying Geospatial in the Common Crawl Corpus
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026)
by: Kim, Kyuhee, et al.
Published: (2026)
Neural Prioritisation for Web Crawling
by: Pezzuti, Francesca, et al.
Published: (2025)
by: Pezzuti, Francesca, et al.
Published: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
by: Zheng, Qiaoyuan, et al.
Published: (2026)
by: Zheng, Qiaoyuan, et al.
Published: (2026)
On the Effect of (Near) Duplicate Subwords in Language Modelling
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
Kollektivierung und Opt-Out
by: Hohlefelder, Olaf
Published: (2020)
by: Hohlefelder, Olaf
Published: (2020)
Document Quality Scoring for Web Crawling
by: Pezzuti, Francesca, et al.
Published: (2025)
by: Pezzuti, Francesca, et al.
Published: (2025)
L'«ermeneutica della riforma nella continuità» A proposito dei libri di Ernst-Wolfgang Böckenförde, Martin Rhonheimer e Rudolf Uertz: tre tentativi di trovare un posto alla fede nella modernità
by: Martin Schlag
Published: (2010)
by: Martin Schlag
Published: (2010)
LLMs Are In-Context Bandit Reinforcement Learners
by: Monea, Giovanni, et al.
Published: (2024)
by: Monea, Giovanni, et al.
Published: (2024)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
by: Tessema, Bethel Melesse, et al.
Published: (2024)
by: Tessema, Bethel Melesse, et al.
Published: (2024)
Instruction-tuning Aligns LLMs to the Human Brain
by: Aw, Khai Loong, et al.
Published: (2023)
by: Aw, Khai Loong, et al.
Published: (2023)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
by: Limozin, Alexis, et al.
Published: (2026)
by: Limozin, Alexis, et al.
Published: (2026)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Understanding and Minimising Outlier Features in Neural Network Training
by: He, Bobby, et al.
Published: (2024)
by: He, Bobby, et al.
Published: (2024)
Similar Items
-
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
by: Fan, Dongyang, et al.
Published: (2025) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025) -
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026) -
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
by: Xu, Yixuan, et al.
Published: (2025) -
Towards Fully FP8 GEMM LLM Training at Scale
by: Hernández-Cano, Alejandro, et al.
Published: (2025)