Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Dongyang, Sabolčec, Vinko, Ansaripour, Matin, Tarun, Ayush Kumar, Jaggi, Martin, Bosselut, Antoine, Schlag, Imanol |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
von: Turki, Yassine, et al.
Veröffentlicht: (2026)
von: Turki, Yassine, et al.
Veröffentlicht: (2026)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
von: Xu, Yixuan, et al.
Veröffentlicht: (2025)
von: Xu, Yixuan, et al.
Veröffentlicht: (2025)
Towards Fully FP8 GEMM LLM Training at Scale
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025)
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
von: Badanin, Ilia, et al.
Veröffentlicht: (2026)
von: Badanin, Ilia, et al.
Veröffentlicht: (2026)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
von: Allard, Marc-Antoine, et al.
Veröffentlicht: (2024)
von: Allard, Marc-Antoine, et al.
Veröffentlicht: (2024)
Gabow's Cardinality Matching Algorithm in General Graphs: Implementation and Experiments
von: Ansaripour, Matin, et al.
Veröffentlicht: (2024)
von: Ansaripour, Matin, et al.
Veröffentlicht: (2024)
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
von: Penedo, Guilherme, et al.
Veröffentlicht: (2025)
von: Penedo, Guilherme, et al.
Veröffentlicht: (2025)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
von: Talokar, Nivya, et al.
Veröffentlicht: (2026)
von: Talokar, Nivya, et al.
Veröffentlicht: (2026)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
Towards an empirical understanding of MoE design choices
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
Deriving Activation Functions Using Integration
von: Huang, Allen Hao, et al.
Veröffentlicht: (2024)
von: Huang, Allen Hao, et al.
Veröffentlicht: (2024)
Movie Recommendation using Web Crawling
von: Raj, Pronit, et al.
Veröffentlicht: (2024)
von: Raj, Pronit, et al.
Veröffentlicht: (2024)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
Approximate EFX and Exact tEFX Allocations for Indivisible Chores: Improved Algorithms
von: Afshinmehr, Mahyar, et al.
Veröffentlicht: (2024)
von: Afshinmehr, Mahyar, et al.
Veröffentlicht: (2024)
IDSIA/recurrent-fwp: v1.0.0 - Public Code Release
von: Kazuki Irie, et al.
Veröffentlicht: (2025)
von: Kazuki Irie, et al.
Veröffentlicht: (2025)
Efficient Crawling for Scalable Web Data Acquisition (Extended Version)
von: Gauquier, Antoine, et al.
Veröffentlicht: (2026)
von: Gauquier, Antoine, et al.
Veröffentlicht: (2026)
Web Page Classification using LLMs for Crawling Support
von: Sasazawa, Yuichi, et al.
Veröffentlicht: (2025)
von: Sasazawa, Yuichi, et al.
Veröffentlicht: (2025)
TiMoE: Time-Aware Mixture of Language Experts
von: Faro, Robin, et al.
Veröffentlicht: (2025)
von: Faro, Robin, et al.
Veröffentlicht: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
Quantifying Geospatial in the Common Crawl Corpus
von: Ilyankou, Ilya, et al.
Veröffentlicht: (2024)
von: Ilyankou, Ilya, et al.
Veröffentlicht: (2024)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
von: Kim, Kyuhee, et al.
Veröffentlicht: (2026)
von: Kim, Kyuhee, et al.
Veröffentlicht: (2026)
Neural Prioritisation for Web Crawling
von: Pezzuti, Francesca, et al.
Veröffentlicht: (2025)
von: Pezzuti, Francesca, et al.
Veröffentlicht: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
von: Zheng, Qiaoyuan, et al.
Veröffentlicht: (2026)
von: Zheng, Qiaoyuan, et al.
Veröffentlicht: (2026)
On the Effect of (Near) Duplicate Subwords in Language Modelling
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
von: Anagnostidis, Sotiris, et al.
Veröffentlicht: (2023)
von: Anagnostidis, Sotiris, et al.
Veröffentlicht: (2023)
Kollektivierung und Opt-Out
von: Hohlefelder, Olaf
Veröffentlicht: (2020)
von: Hohlefelder, Olaf
Veröffentlicht: (2020)
Document Quality Scoring for Web Crawling
von: Pezzuti, Francesca, et al.
Veröffentlicht: (2025)
von: Pezzuti, Francesca, et al.
Veröffentlicht: (2025)
L'«ermeneutica della riforma nella continuità» A proposito dei libri di Ernst-Wolfgang Böckenförde, Martin Rhonheimer e Rudolf Uertz: tre tentativi di trovare un posto alla fede nella modernità
von: Martin Schlag
Veröffentlicht: (2010)
von: Martin Schlag
Veröffentlicht: (2010)
LLMs Are In-Context Bandit Reinforcement Learners
von: Monea, Giovanni, et al.
Veröffentlicht: (2024)
von: Monea, Giovanni, et al.
Veröffentlicht: (2024)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
von: Gao, Silin, et al.
Veröffentlicht: (2025)
von: Gao, Silin, et al.
Veröffentlicht: (2025)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
von: Tessema, Bethel Melesse, et al.
Veröffentlicht: (2024)
von: Tessema, Bethel Melesse, et al.
Veröffentlicht: (2024)
Instruction-tuning Aligns LLMs to the Human Brain
von: Aw, Khai Loong, et al.
Veröffentlicht: (2023)
von: Aw, Khai Loong, et al.
Veröffentlicht: (2023)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
Understanding and Minimising Outlier Features in Neural Network Training
von: He, Bobby, et al.
Veröffentlicht: (2024)
von: He, Bobby, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
von: Fan, Dongyang, et al.
Veröffentlicht: (2025) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025) -
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
von: Turki, Yassine, et al.
Veröffentlicht: (2026) -
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
von: Xu, Yixuan, et al.
Veröffentlicht: (2025) -
Towards Fully FP8 GEMM LLM Training at Scale
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025)