Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
Fuente:
arXiv
Salvato in:
| Autori principali: | Fan, Dongyang, Sabolčec, Vinko, Ansaripour, Matin, Tarun, Ayush Kumar, Jaggi, Martin, Bosselut, Antoine, Schlag, Imanol |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
di: Turki, Yassine, et al.
Pubblicazione: (2026)
di: Turki, Yassine, et al.
Pubblicazione: (2026)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
di: Xu, Yixuan, et al.
Pubblicazione: (2025)
di: Xu, Yixuan, et al.
Pubblicazione: (2025)
Towards Fully FP8 GEMM LLM Training at Scale
di: Hernández-Cano, Alejandro, et al.
Pubblicazione: (2025)
di: Hernández-Cano, Alejandro, et al.
Pubblicazione: (2025)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
di: Badanin, Ilia, et al.
Pubblicazione: (2026)
di: Badanin, Ilia, et al.
Pubblicazione: (2026)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
di: Allard, Marc-Antoine, et al.
Pubblicazione: (2024)
di: Allard, Marc-Antoine, et al.
Pubblicazione: (2024)
Gabow's Cardinality Matching Algorithm in General Graphs: Implementation and Experiments
di: Ansaripour, Matin, et al.
Pubblicazione: (2024)
di: Ansaripour, Matin, et al.
Pubblicazione: (2024)
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
di: Penedo, Guilherme, et al.
Pubblicazione: (2025)
di: Penedo, Guilherme, et al.
Pubblicazione: (2025)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
di: Talokar, Nivya, et al.
Pubblicazione: (2026)
di: Talokar, Nivya, et al.
Pubblicazione: (2026)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
di: Wagner, Nicolas, et al.
Pubblicazione: (2024)
di: Wagner, Nicolas, et al.
Pubblicazione: (2024)
Towards an empirical understanding of MoE design choices
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
Deriving Activation Functions Using Integration
di: Huang, Allen Hao, et al.
Pubblicazione: (2024)
di: Huang, Allen Hao, et al.
Pubblicazione: (2024)
Movie Recommendation using Web Crawling
di: Raj, Pronit, et al.
Pubblicazione: (2024)
di: Raj, Pronit, et al.
Pubblicazione: (2024)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
di: Ansaripour, Matin, et al.
Pubblicazione: (2022)
di: Ansaripour, Matin, et al.
Pubblicazione: (2022)
Approximate EFX and Exact tEFX Allocations for Indivisible Chores: Improved Algorithms
di: Afshinmehr, Mahyar, et al.
Pubblicazione: (2024)
di: Afshinmehr, Mahyar, et al.
Pubblicazione: (2024)
IDSIA/recurrent-fwp: v1.0.0 - Public Code Release
di: Kazuki Irie, et al.
Pubblicazione: (2025)
di: Kazuki Irie, et al.
Pubblicazione: (2025)
Efficient Crawling for Scalable Web Data Acquisition (Extended Version)
di: Gauquier, Antoine, et al.
Pubblicazione: (2026)
di: Gauquier, Antoine, et al.
Pubblicazione: (2026)
Web Page Classification using LLMs for Crawling Support
di: Sasazawa, Yuichi, et al.
Pubblicazione: (2025)
di: Sasazawa, Yuichi, et al.
Pubblicazione: (2025)
TiMoE: Time-Aware Mixture of Language Experts
di: Faro, Robin, et al.
Pubblicazione: (2025)
di: Faro, Robin, et al.
Pubblicazione: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
Quantifying Geospatial in the Common Crawl Corpus
di: Ilyankou, Ilya, et al.
Pubblicazione: (2024)
di: Ilyankou, Ilya, et al.
Pubblicazione: (2024)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
di: Kim, Kyuhee, et al.
Pubblicazione: (2026)
di: Kim, Kyuhee, et al.
Pubblicazione: (2026)
Neural Prioritisation for Web Crawling
di: Pezzuti, Francesca, et al.
Pubblicazione: (2025)
di: Pezzuti, Francesca, et al.
Pubblicazione: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
di: Zheng, Qiaoyuan, et al.
Pubblicazione: (2026)
di: Zheng, Qiaoyuan, et al.
Pubblicazione: (2026)
On the Effect of (Near) Duplicate Subwords in Language Modelling
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2023)
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2023)
Kollektivierung und Opt-Out
di: Hohlefelder, Olaf
Pubblicazione: (2020)
di: Hohlefelder, Olaf
Pubblicazione: (2020)
Document Quality Scoring for Web Crawling
di: Pezzuti, Francesca, et al.
Pubblicazione: (2025)
di: Pezzuti, Francesca, et al.
Pubblicazione: (2025)
L'«ermeneutica della riforma nella continuità» A proposito dei libri di Ernst-Wolfgang Böckenförde, Martin Rhonheimer e Rudolf Uertz: tre tentativi di trovare un posto alla fede nella modernità
di: Martin Schlag
Pubblicazione: (2010)
di: Martin Schlag
Pubblicazione: (2010)
LLMs Are In-Context Bandit Reinforcement Learners
di: Monea, Giovanni, et al.
Pubblicazione: (2024)
di: Monea, Giovanni, et al.
Pubblicazione: (2024)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
di: Gao, Silin, et al.
Pubblicazione: (2025)
di: Gao, Silin, et al.
Pubblicazione: (2025)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
di: Tessema, Bethel Melesse, et al.
Pubblicazione: (2024)
di: Tessema, Bethel Melesse, et al.
Pubblicazione: (2024)
Instruction-tuning Aligns LLMs to the Human Brain
di: Aw, Khai Loong, et al.
Pubblicazione: (2023)
di: Aw, Khai Loong, et al.
Pubblicazione: (2023)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
di: Limozin, Alexis, et al.
Pubblicazione: (2026)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
Understanding and Minimising Outlier Features in Neural Network Training
di: He, Bobby, et al.
Pubblicazione: (2024)
di: He, Bobby, et al.
Pubblicazione: (2024)
Documenti analoghi
-
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
di: Fan, Dongyang, et al.
Pubblicazione: (2025) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
di: Messmer, Bettina, et al.
Pubblicazione: (2025) -
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
di: Turki, Yassine, et al.
Pubblicazione: (2026) -
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
di: Xu, Yixuan, et al.
Pubblicazione: (2025) -
Towards Fully FP8 GEMM LLM Training at Scale
di: Hernández-Cano, Alejandro, et al.
Pubblicazione: (2025)