AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
Fuente:
arXiv
Salvato in:
| Autori principali: | Lucy, Li, Gururangan, Suchin, Soldaini, Luca, Strubell, Emma, Bamman, David, Klein, Lauren F., Dodge, Jesse |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Time is Encoded in the Weights of Finetuned Language Models
di: Nylund, Kai, et al.
Pubblicazione: (2023)
di: Nylund, Kai, et al.
Pubblicazione: (2023)
LESS: Selecting Influential Data for Targeted Instruction Tuning
di: Xia, Mengzhou, et al.
Pubblicazione: (2024)
di: Xia, Mengzhou, et al.
Pubblicazione: (2024)
Gradient Localization Improves Lifelong Pretraining of Language Models
di: Fernandez, Jared, et al.
Pubblicazione: (2024)
di: Fernandez, Jared, et al.
Pubblicazione: (2024)
Holistically Evaluating the Environmental Impact of Creating Language Models
di: Morrison, Jacob, et al.
Pubblicazione: (2025)
di: Morrison, Jacob, et al.
Pubblicazione: (2025)
Scalable Data Ablation Approximations for Language Models through Modular Training and Merging
di: Na, Clara, et al.
Pubblicazione: (2024)
di: Na, Clara, et al.
Pubblicazione: (2024)
DataDecide: How to Predict Best Pretraining Data with Small Experiments
di: Magnusson, Ian, et al.
Pubblicazione: (2025)
di: Magnusson, Ian, et al.
Pubblicazione: (2025)
The Hidden Cost of Thinking: Energy Use and Environmental Impact of LMs Beyond Pretraining
di: Morrison, Jacob, et al.
Pubblicazione: (2026)
di: Morrison, Jacob, et al.
Pubblicazione: (2026)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
di: Yang, Xianjun, et al.
Pubblicazione: (2025)
di: Yang, Xianjun, et al.
Pubblicazione: (2025)
olmOCR 2: Unit Test Rewards for Document OCR
di: Poznanski, Jake, et al.
Pubblicazione: (2025)
di: Poznanski, Jake, et al.
Pubblicazione: (2025)
Beyond Text: Characterizing Domain Expert Needs in Document Research
di: Gururaja, Sireesh, et al.
Pubblicazione: (2025)
di: Gururaja, Sireesh, et al.
Pubblicazione: (2025)
From Efficiency Gains to Rebound Effects: The Problem of Jevons' Paradox in AI's Polarized Environmental Debate
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2025)
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2025)
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
di: Min, Sewon, et al.
Pubblicazione: (2023)
di: Min, Sewon, et al.
Pubblicazione: (2023)
Breaking the Curse of Multilinguality with Cross-lingual Expert Language Models
di: Blevins, Terra, et al.
Pubblicazione: (2024)
di: Blevins, Terra, et al.
Pubblicazione: (2024)
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
di: Jayalath, Dulhan, et al.
Pubblicazione: (2025)
di: Jayalath, Dulhan, et al.
Pubblicazione: (2025)
On Classification with Large Language Models in Cultural Analytics
di: Bamman, David, et al.
Pubblicazione: (2024)
di: Bamman, David, et al.
Pubblicazione: (2024)
Automated Phishing Detection Using URLs and Webpages
di: Wang, Huilin, et al.
Pubblicazione: (2024)
di: Wang, Huilin, et al.
Pubblicazione: (2024)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
di: Saada, Thiziri Nait, et al.
Pubblicazione: (2025)
di: Saada, Thiziri Nait, et al.
Pubblicazione: (2025)
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
di: Soldaini, Luca, et al.
Pubblicazione: (2024)
di: Soldaini, Luca, et al.
Pubblicazione: (2024)
Once More, With Feeling: Measuring Emotion of Acting Performances in Contemporary American Film
di: Zhou, Naitian, et al.
Pubblicazione: (2024)
di: Zhou, Naitian, et al.
Pubblicazione: (2024)
What's In My Big Data?
di: Elazar, Yanai, et al.
Pubblicazione: (2023)
di: Elazar, Yanai, et al.
Pubblicazione: (2023)
El projecte Atlantis: recursos digitals per a les llengües minoritzades de la UE, recursos per a l´ensenyament del català
di: Miquel Strubell
Pubblicazione: (2003)
di: Miquel Strubell
Pubblicazione: (2003)
Rethinking Thinking Tokens: LLMs as Improvement Operators
di: Madaan, Lovish, et al.
Pubblicazione: (2025)
di: Madaan, Lovish, et al.
Pubblicazione: (2025)
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2023)
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2023)
"Are Adversarial Phishing Webpages a Threat in Reality?" Understanding the Users' Perception of Adversarial Webpages
di: Yuan, Ying, et al.
Pubblicazione: (2024)
di: Yuan, Ying, et al.
Pubblicazione: (2024)
Representation Fidelity:Auditing Algorithmic Decisions About Humans Using Self-Descriptions
di: Elstner, Theresa, et al.
Pubblicazione: (2026)
di: Elstner, Theresa, et al.
Pubblicazione: (2026)
Hypertext Entity Extraction in Webpage
di: Yang, Yifei, et al.
Pubblicazione: (2024)
di: Yang, Yifei, et al.
Pubblicazione: (2024)
Self-Generated Critiques Boost Reward Modeling for Language Models
di: Yu, Yue, et al.
Pubblicazione: (2024)
di: Yu, Yue, et al.
Pubblicazione: (2024)
Filter Like You Test: Data-Driven Data Filtering for CLIP Pretraining
di: Shechter, Mikey, et al.
Pubblicazione: (2025)
di: Shechter, Mikey, et al.
Pubblicazione: (2025)
Teaching Models to Understand (but not Generate) High-risk Data
di: Wang, Ryan, et al.
Pubblicazione: (2025)
di: Wang, Ryan, et al.
Pubblicazione: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
di: Ponnock, Jesse
Pubblicazione: (2025)
di: Ponnock, Jesse
Pubblicazione: (2025)
On-device Streaming Discrete Speech Units
di: Choi, Kwanghee, et al.
Pubblicazione: (2025)
di: Choi, Kwanghee, et al.
Pubblicazione: (2025)
Mathfish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula
di: Lucy, Li, et al.
Pubblicazione: (2024)
di: Lucy, Li, et al.
Pubblicazione: (2024)
Self-Directed Synthetic Dialogues and Revisions Technical Report
di: Lambert, Nathan, et al.
Pubblicazione: (2024)
di: Lambert, Nathan, et al.
Pubblicazione: (2024)
Hold Me Down
di: Lauren, Ben
Pubblicazione: (2025)
di: Lauren, Ben
Pubblicazione: (2025)
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
di: Baral, Sami, et al.
Pubblicazione: (2025)
di: Baral, Sami, et al.
Pubblicazione: (2025)
FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction
di: Johnson, Natasha, et al.
Pubblicazione: (2025)
di: Johnson, Natasha, et al.
Pubblicazione: (2025)
Tell, Don't Show: Leveraging Language Models' Abstractive Retellings to Model Literary Themes
di: Lucy, Li, et al.
Pubblicazione: (2025)
di: Lucy, Li, et al.
Pubblicazione: (2025)
Towards Fine-Grained Webpage Fingerprinting at Scale
di: Zhao, Xiyuan, et al.
Pubblicazione: (2024)
di: Zhao, Xiyuan, et al.
Pubblicazione: (2024)
Cloenda
di: Miquel Strubell i Trueta
Pubblicazione: (2004)
di: Miquel Strubell i Trueta
Pubblicazione: (2004)
Information Flow Control in Machine Learning through Modular Model Architecture
di: Tiwari, Trishita, et al.
Pubblicazione: (2023)
di: Tiwari, Trishita, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Time is Encoded in the Weights of Finetuned Language Models
di: Nylund, Kai, et al.
Pubblicazione: (2023) -
LESS: Selecting Influential Data for Targeted Instruction Tuning
di: Xia, Mengzhou, et al.
Pubblicazione: (2024) -
Gradient Localization Improves Lifelong Pretraining of Language Models
di: Fernandez, Jared, et al.
Pubblicazione: (2024) -
Holistically Evaluating the Environmental Impact of Creating Language Models
di: Morrison, Jacob, et al.
Pubblicazione: (2025) -
Scalable Data Ablation Approximations for Language Models through Modular Training and Merging
di: Na, Clara, et al.
Pubblicazione: (2024)