URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Dongyang, Sabolčec, Vinko, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
by: Penedo, Guilherme, et al.
Published: (2025)
by: Penedo, Guilherme, et al.
Published: (2025)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
TiMoE: Time-Aware Mixture of Language Experts
by: Faro, Robin, et al.
Published: (2025)
by: Faro, Robin, et al.
Published: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Semantic uncertainty in advanced decoding methods for LLM generation
by: Foodeei, Darius, et al.
Published: (2025)
by: Foodeei, Darius, et al.
Published: (2025)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
DomURLs_BERT: Pre-trained BERT-based Model for Malicious Domains and URLs Detection and Classification
by: Mahdaouy, Abdelkader El, et al.
Published: (2024)
by: Mahdaouy, Abdelkader El, et al.
Published: (2024)
Safety Training Persists Through Helpfulness Optimization in LLM Agents
by: Plaut, Benjamin
Published: (2026)
by: Plaut, Benjamin
Published: (2026)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
SciTopic: Enhancing Topic Discovery in Scientific Literature through Advanced LLM
by: Li, Pengjiang, et al.
Published: (2025)
by: Li, Pengjiang, et al.
Published: (2025)
Self-Critique-Guided Curiosity Refinement: Enhancing Honesty and Helpfulness in Large Language Models via In-Context Learning
by: Ho, Duc Hieu, et al.
Published: (2025)
by: Ho, Duc Hieu, et al.
Published: (2025)
TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries
by: Lv, Jinze, et al.
Published: (2025)
by: Lv, Jinze, et al.
Published: (2025)
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes
by: Loakman, Tyler, et al.
Published: (2025)
by: Loakman, Tyler, et al.
Published: (2025)
With a Little Help from my (Linguistic) Friends: Topic Segmentation of Multi-party Casual Conversations
by: Decker, Amandine, et al.
Published: (2024)
by: Decker, Amandine, et al.
Published: (2024)
Multi-Prompting Decoder Helps Better Language Understanding
by: Cheng, Zifeng, et al.
Published: (2024)
by: Cheng, Zifeng, et al.
Published: (2024)
Towards Transparency: Exploring LLM Trainings Datasets through Visual Topic Modeling and Semantic Frame
by: de Dampierre, Charles, et al.
Published: (2024)
by: de Dampierre, Charles, et al.
Published: (2024)
Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Understanding Cross-Domain Adaptation in Low-Resource Topic Modeling
by: Akash, Pritom Saha, et al.
Published: (2025)
by: Akash, Pritom Saha, et al.
Published: (2025)
Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
by: Fan, Qihui, et al.
Published: (2026)
by: Fan, Qihui, et al.
Published: (2026)
Semantic-Augmented Latent Topic Modeling with LLM-in-the-Loop
by: Hong, Mengze, et al.
Published: (2025)
by: Hong, Mengze, et al.
Published: (2025)
Does Reasoning Help LLM Agents Play Dungeons and Dragons? A Prompt Engineering Experiment
by: Delafuente, Patricia, et al.
Published: (2025)
by: Delafuente, Patricia, et al.
Published: (2025)
Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
by: Liu, Qiang, et al.
Published: (2025)
by: Liu, Qiang, et al.
Published: (2025)
Parsing Millions of URLs per Second
by: Nizipli, Yagiz, et al.
Published: (2023)
by: Nizipli, Yagiz, et al.
Published: (2023)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
ReMiT: RL-Guided Mid-Training for Iterative LLM Evolution
by: Huang, Junjie, et al.
Published: (2026)
by: Huang, Junjie, et al.
Published: (2026)
Personalized Topic Selection Model for Topic-Grounded Dialogue
by: Fan, Shixuan, et al.
Published: (2024)
by: Fan, Shixuan, et al.
Published: (2024)
Understanding the Progression of Educational Topics via Semantic Matching
by: Alkhidir, Tamador, et al.
Published: (2024)
by: Alkhidir, Tamador, et al.
Published: (2024)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
by: Li, Chuyuan, et al.
Published: (2025)
by: Li, Chuyuan, et al.
Published: (2025)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
by: Rawat, Ankit Singh, et al.
Published: (2024)
by: Rawat, Ankit Singh, et al.
Published: (2024)
CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
by: Kostelník, Martin, et al.
Published: (2026)
by: Kostelník, Martin, et al.
Published: (2026)
Mined Prompting and Metadata-Guided Generation for Wound Care Visual Question Answering
by: Durgapraveen, Bavana, et al.
Published: (2025)
by: Durgapraveen, Bavana, et al.
Published: (2025)
Service, Solidarity, and Self-Help: A Comparative Topic Modeling Analysis of Community Unionism in the Boot and Shoe Union and Unite Community
by: Compton, Thomas
Published: (2025)
by: Compton, Thomas
Published: (2025)
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
by: Suh, Joseph, et al.
Published: (2026)
by: Suh, Joseph, et al.
Published: (2026)
Knowledge Hierarchy Guided Biological-Medical Dataset Distillation for Domain LLM Training
by: Cai, Xunxin, et al.
Published: (2025)
by: Cai, Xunxin, et al.
Published: (2025)
Similar Items
-
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025) -
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026) -
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025) -
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
by: Penedo, Guilherme, et al.
Published: (2025)