RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
Fuente:
arXiv
Salvato in:
| Autori principali: | Bi, Baolong, Liu, Shenghua, Ren, Xingzhang, Liu, Dayiheng, Lin, Junyang, Wang, Yiwei, Mei, Lingrui, Fang, Junfeng, Guo, Jiafeng, Cheng, Xueqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SLANG: New Concept Comprehension of Large Language Models
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
LPNL: Scalable Link Prediction with Large Language Models
di: Bi, Baolong, et al.
Pubblicazione: (2024)
di: Bi, Baolong, et al.
Pubblicazione: (2024)
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
di: Bi, Baolong, et al.
Pubblicazione: (2025)
di: Bi, Baolong, et al.
Pubblicazione: (2025)
StruEdit: Structured Outputs Enable the Fast and Accurate Knowledge Editing for Large Language Models
di: Bi, Baolong, et al.
Pubblicazione: (2024)
di: Bi, Baolong, et al.
Pubblicazione: (2024)
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
Decoding by Contrasting Knowledge: Enhancing LLMs' Confidence on Edited Facts
di: Bi, Baolong, et al.
Pubblicazione: (2024)
di: Bi, Baolong, et al.
Pubblicazione: (2024)
HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
di: Mei, Lingrui, et al.
Pubblicazione: (2024)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
di: Bi, Baolong, et al.
Pubblicazione: (2024)
di: Bi, Baolong, et al.
Pubblicazione: (2024)
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
di: Ge, Yuyao, et al.
Pubblicazione: (2026)
di: Ge, Yuyao, et al.
Pubblicazione: (2026)
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
Adaptive Token Biaser: Knowledge Editing via Biasing Key Entities
di: Bi, Baolong, et al.
Pubblicazione: (2024)
di: Bi, Baolong, et al.
Pubblicazione: (2024)
a1: Steep Test-time Scaling Law via Environment Augmented Generation
di: Mei, Lingrui, et al.
Pubblicazione: (2025)
di: Mei, Lingrui, et al.
Pubblicazione: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
di: Bi, Baolong, et al.
Pubblicazione: (2025)
di: Bi, Baolong, et al.
Pubblicazione: (2025)
Not in Sync: Unveiling Temporal Bias in Audio Chat Models
di: Yao, Jiayu, et al.
Pubblicazione: (2025)
di: Yao, Jiayu, et al.
Pubblicazione: (2025)
Can Graph Descriptive Order Affect Solving Graph Problems with LLMs?
di: Ge, Yuyao, et al.
Pubblicazione: (2024)
di: Ge, Yuyao, et al.
Pubblicazione: (2024)
Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
di: Yao, Jiayu, et al.
Pubblicazione: (2025)
di: Yao, Jiayu, et al.
Pubblicazione: (2025)
Gated Differentiable Working Memory for Long-Context Language Modeling
di: Mei, Lingrui, et al.
Pubblicazione: (2026)
di: Mei, Lingrui, et al.
Pubblicazione: (2026)
The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates
di: Wang, Shaobo, et al.
Pubblicazione: (2026)
di: Wang, Shaobo, et al.
Pubblicazione: (2026)
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
di: Tian, Yuxin, et al.
Pubblicazione: (2024)
di: Tian, Yuxin, et al.
Pubblicazione: (2024)
A Survey of Vibe Coding with Large Language Models
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
di: Bi, Baolong, et al.
Pubblicazione: (2026)
di: Bi, Baolong, et al.
Pubblicazione: (2026)
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
di: Ni, Shiyu, et al.
Pubblicazione: (2025)
di: Ni, Shiyu, et al.
Pubblicazione: (2025)
Rethinking All Evidence: Enhancing Trustworthy Retrieval-Augmented Generation via Conflict-Driven Summarization
di: Chen, Juan, et al.
Pubblicazione: (2025)
di: Chen, Juan, et al.
Pubblicazione: (2025)
DataMan: Data Manager for Pre-training Large Language Models
di: Peng, Ru, et al.
Pubblicazione: (2025)
di: Peng, Ru, et al.
Pubblicazione: (2025)
Context-DPO: Aligning Language Models for Context-Faithfulness
di: Bi, Baolong, et al.
Pubblicazione: (2024)
di: Bi, Baolong, et al.
Pubblicazione: (2024)
A Survey of Context Engineering for Large Language Models
di: Mei, Lingrui, et al.
Pubblicazione: (2025)
di: Mei, Lingrui, et al.
Pubblicazione: (2025)
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
di: Wang, Shaobo, et al.
Pubblicazione: (2026)
di: Wang, Shaobo, et al.
Pubblicazione: (2026)
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
di: Wang, Lexin, et al.
Pubblicazione: (2026)
di: Wang, Lexin, et al.
Pubblicazione: (2026)
Bootstrapped Pre-training with Dynamic Identifier Prediction for Generative Retrieval
di: Tang, Yubao, et al.
Pubblicazione: (2024)
di: Tang, Yubao, et al.
Pubblicazione: (2024)
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
di: Zhou, Fan, et al.
Pubblicazione: (2024)
di: Zhou, Fan, et al.
Pubblicazione: (2024)
RefineStyle: Dynamic Convolution Refinement for StyleGAN
di: Xia, Siwei, et al.
Pubblicazione: (2024)
di: Xia, Siwei, et al.
Pubblicazione: (2024)
Bridging Queries and Tables through Entities in Table Retrieval
di: Li, Da, et al.
Pubblicazione: (2025)
di: Li, Da, et al.
Pubblicazione: (2025)
MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
Reproducibility Analysis and Enhancements for Multi-Aspect Dense Retriever with Aspect Learning
di: Bi, Keping, et al.
Pubblicazione: (2024)
di: Bi, Keping, et al.
Pubblicazione: (2024)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
di: Zhang, Hengran, et al.
Pubblicazione: (2024)
di: Zhang, Hengran, et al.
Pubblicazione: (2024)
Tailoring Table Retrieval from a Field-aware Hybrid Matching Perspective
di: Li, Da, et al.
Pubblicazione: (2025)
di: Li, Da, et al.
Pubblicazione: (2025)
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
di: Ni, Shiyu, et al.
Pubblicazione: (2025)
di: Ni, Shiyu, et al.
Pubblicazione: (2025)
CIR at the NTCIR-17 ULTRE-2 Task
di: Yu, Lulu, et al.
Pubblicazione: (2023)
di: Yu, Lulu, et al.
Pubblicazione: (2023)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
di: Ni, Shiyu, et al.
Pubblicazione: (2024)
di: Ni, Shiyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SLANG: New Concept Comprehension of Large Language Models
di: Mei, Lingrui, et al.
Pubblicazione: (2024) -
LPNL: Scalable Link Prediction with Large Language Models
di: Bi, Baolong, et al.
Pubblicazione: (2024) -
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
di: Bi, Baolong, et al.
Pubblicazione: (2025) -
StruEdit: Structured Outputs Enable the Fast and Accurate Knowledge Editing for Large Language Models
di: Bi, Baolong, et al.
Pubblicazione: (2024) -
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
di: Mei, Lingrui, et al.
Pubblicazione: (2024)