Enhancing LLMs via High-Knowledge Data Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Feiyu, Zhang, Xuemiao, Wang, Sirui, Que, Haoran, Liu, Yuqi, Rong, Wenge, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation
by: Wang, Keheng, et al.
Published: (2024)
by: Wang, Keheng, et al.
Published: (2024)
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
Large-Scale Diverse Synthesis for Mid-Training
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation
by: Pan, Ruotong, et al.
Published: (2024)
by: Pan, Ruotong, et al.
Published: (2024)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
by: Zhang, Jianfei, et al.
Published: (2025)
by: Zhang, Jianfei, et al.
Published: (2025)
Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
A Survey on LLM Mid-Training
by: Tu, Chengying, et al.
Published: (2025)
by: Tu, Chengying, et al.
Published: (2025)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024)
by: Liu, Jiahao, et al.
Published: (2024)
Prejudge-Before-Think: Enhancing Large Language Models at Test-Time by Process Prejudge Reasoning
by: Wang, Jianing, et al.
Published: (2025)
by: Wang, Jianing, et al.
Published: (2025)
HyperG: Hypergraph-Enhanced LLMs for Structured Knowledge
by: Huang, Sirui, et al.
Published: (2025)
by: Huang, Sirui, et al.
Published: (2025)
Explainable Few-shot Knowledge Tracing
by: Li, Haoxuan, et al.
Published: (2024)
by: Li, Haoxuan, et al.
Published: (2024)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
by: Hong, Hanhua, et al.
Published: (2025)
by: Hong, Hanhua, et al.
Published: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
by: Ding, Peng, et al.
Published: (2025)
by: Ding, Peng, et al.
Published: (2025)
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
by: Cang, Yueyang, et al.
Published: (2026)
by: Cang, Yueyang, et al.
Published: (2026)
medIKAL: Integrating Knowledge Graphs as Assistants of LLMs for Enhanced Clinical Diagnosis on EMRs
by: Jia, Mingyi, et al.
Published: (2024)
by: Jia, Mingyi, et al.
Published: (2024)
Dynamic Fisher-weighted Model Merging via Bayesian Optimization
by: Lee, Sanwoo, et al.
Published: (2025)
by: Lee, Sanwoo, et al.
Published: (2025)
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
by: Tao, Xingjian, et al.
Published: (2024)
by: Tao, Xingjian, et al.
Published: (2024)
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
by: Que, Shumin, et al.
Published: (2025)
by: Que, Shumin, et al.
Published: (2025)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
by: Tang, Hongyin, et al.
Published: (2024)
by: Tang, Hongyin, et al.
Published: (2024)
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
KSOD: Knowledge Supplement for LLMs On Demand
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Knowledge Graph-Enhanced Large Language Models via Path Selection
by: Liu, Haochen, et al.
Published: (2024)
by: Liu, Haochen, et al.
Published: (2024)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
by: Cao, Maosong, et al.
Published: (2025)
by: Cao, Maosong, et al.
Published: (2025)
Beyond the Known: Investigating LLMs Performance on Out-of-Domain Intent Detection
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
by: Zhao, Wanru, et al.
Published: (2025)
by: Zhao, Wanru, et al.
Published: (2025)
Decoding by Contrasting Knowledge: Enhancing LLMs' Confidence on Edited Facts
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
by: Duan, Feiyu, et al.
Published: (2026)
by: Duan, Feiyu, et al.
Published: (2026)
PositionID: LLMs can Control Lengths, Copy and Paste with Explicit Positional Awareness
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
DKE-Research at SemEval-2024 Task 2: Incorporating Data Augmentation with Generative Models and Biomedical Knowledge to Enhance Inference Robustness
by: Wang, Yuqi, et al.
Published: (2024)
by: Wang, Yuqi, et al.
Published: (2024)
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
by: Wang, Sizhe, et al.
Published: (2026)
by: Wang, Sizhe, et al.
Published: (2026)
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
by: Ding, Peng, et al.
Published: (2025)
by: Ding, Peng, et al.
Published: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026)
by: Yang, Xin, et al.
Published: (2026)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
by: Li, Yunxin, et al.
Published: (2023)
by: Li, Yunxin, et al.
Published: (2023)
Similar Items
-
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025) -
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025) -
FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy
by: Zhang, Xuemiao, et al.
Published: (2025) -
LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation
by: Wang, Keheng, et al.
Published: (2024) -
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
by: Zhang, Xuemiao, et al.
Published: (2025)