Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yu, Shang, Xiaoran, Pei, Qizhi, Zhu, Yun, Gao, Xin, Lin, Honglin, Zhong, Zhanping, Pan, Zhuoshi, Liu, Zheng, Wang, Xiaoyang, He, Conghui, Lin, Dahua, Zhao, Feng, Wu, Lijun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
by: Cai, Mengzhang, et al.
Published: (2025)
by: Cai, Mengzhang, et al.
Published: (2025)
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
by: Liu, Zheng, et al.
Published: (2026)
by: Liu, Zheng, et al.
Published: (2026)
MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer
by: Lin, Honglin, et al.
Published: (2025)
by: Lin, Honglin, et al.
Published: (2025)
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
by: Lin, Honglin, et al.
Published: (2025)
by: Lin, Honglin, et al.
Published: (2025)
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
by: Pei, Qizhi, et al.
Published: (2025)
by: Pei, Qizhi, et al.
Published: (2025)
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
by: Pei, Qizhi, et al.
Published: (2025)
by: Pei, Qizhi, et al.
Published: (2025)
A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
by: Pan, Zhuoshi, et al.
Published: (2025)
by: Pan, Zhuoshi, et al.
Published: (2025)
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
by: Lin, Honglin, et al.
Published: (2026)
by: Lin, Honglin, et al.
Published: (2026)
Unlocking Data Value in Finance: A Study on Distillation and Difficulty-Aware Training
by: Cao, Chuxue, et al.
Published: (2026)
by: Cao, Chuxue, et al.
Published: (2026)
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
by: Tang, Zinan, et al.
Published: (2025)
by: Tang, Zinan, et al.
Published: (2025)
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
by: Pan, Zhuoshi, et al.
Published: (2025)
by: Pan, Zhuoshi, et al.
Published: (2025)
MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
by: Lin, Honglin, et al.
Published: (2026)
by: Lin, Honglin, et al.
Published: (2026)
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
by: Ming, Chenlin, et al.
Published: (2025)
by: Ming, Chenlin, et al.
Published: (2025)
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Bidirectional Curriculum Generation: A Multi-Agent Framework for Data-Efficient Mathematical Reasoning
by: Hu, Boren, et al.
Published: (2026)
by: Hu, Boren, et al.
Published: (2026)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
by: Lv, Kai, et al.
Published: (2024)
by: Lv, Kai, et al.
Published: (2024)
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
by: Lin, Juekai, et al.
Published: (2026)
by: Lin, Juekai, et al.
Published: (2026)
Surface Engineering on Bacteria for Tumor Immunotherapy: Strategies and Perspectives
by: Lijun Fu, et al.
Published: (2024)
by: Lijun Fu, et al.
Published: (2024)
Uncovering Convergent Pattern Recognition Receptors Recognising Phytophthora Across Plant Lineages
by: Yong Pei, et al.
Published: (2025)
by: Yong Pei, et al.
Published: (2025)
3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text Modeling
by: Pei, Qizhi, et al.
Published: (2024)
by: Pei, Qizhi, et al.
Published: (2024)
Data Lineage Inference: Uncovering Privacy Vulnerabilities of Dataset Pruning
by: Li, Qi, et al.
Published: (2024)
by: Li, Qi, et al.
Published: (2024)
Post‐Synthetic Sulfonation of Metal–Organic Frameworks for Efficient Aziridine Ring‐Opening Functionalization
by: Ailing Gu, et al.
Published: (2026)
by: Ailing Gu, et al.
Published: (2026)
Scaling Laws of RoPE-based Extrapolation
by: Liu, Xiaoran, et al.
Published: (2023)
by: Liu, Xiaoran, et al.
Published: (2023)
Estimation and Inference for the Additive Mean Residual Life Model With Change Point in the Continuous Covariate
by: Yun Lin, et al.
Published: (2025)
by: Yun Lin, et al.
Published: (2025)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
by: Zhu, Runchuan, et al.
Published: (2025)
by: Zhu, Runchuan, et al.
Published: (2025)
OpenDataLab: Empowering General Artificial Intelligence with Open Datasets
by: He, Conghui, et al.
Published: (2024)
by: He, Conghui, et al.
Published: (2024)
MemLineage: Lineage-Guided Enforcement for LLM Agent Memory
by: Ouyang, Ciyan, et al.
Published: (2026)
by: Ouyang, Ciyan, et al.
Published: (2026)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
by: Li, Weijia, et al.
Published: (2024)
by: Li, Weijia, et al.
Published: (2024)
Prospective Results of the Minimally Invasive Laser Enucleation of the Prostate (MiLEP)
by: Fu Feng, et al.
Published: (2024)
by: Fu Feng, et al.
Published: (2024)
E2EAI: End-to-End Deep Learning Framework for Active Investing
by: Wei, Zikai, et al.
Published: (2023)
by: Wei, Zikai, et al.
Published: (2023)
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
Track and Trace: Automatically Uncovering Cross-chain Transactions in the Multi-blockchain Ecosystems
by: Lin, Dan, et al.
Published: (2025)
by: Lin, Dan, et al.
Published: (2025)
Training-Free Tunnel Defect Inspection and Engineering Interpretation via Visual Recalibration and Entity Reconstruction
by: Liu, Shipeng, et al.
Published: (2026)
by: Liu, Shipeng, et al.
Published: (2026)
A Two-Field-Scan Harmonic Hall Voltage Analysis For Fast, Accurate Quantification Of Spin-Orbit Torques In Magnetic Heterostructures
by: Lin, Xin, et al.
Published: (2024)
by: Lin, Xin, et al.
Published: (2024)
UniPTS: A Unified Framework for Proficient Post-Training Sparsity
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
Efficient Row-Level Lineage Leveraging Predicate Pushdown
by: Lin, Yin, et al.
Published: (2024)
by: Lin, Yin, et al.
Published: (2024)
FABind+: Enhancing Molecular Docking through Improved Pocket Prediction and Pose Generation
by: Gao, Kaiyuan, et al.
Published: (2024)
by: Gao, Kaiyuan, et al.
Published: (2024)
Similar Items
-
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
by: Cai, Mengzhang, et al.
Published: (2025) -
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
by: Liu, Zheng, et al.
Published: (2026) -
MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer
by: Lin, Honglin, et al.
Published: (2025) -
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
by: Lin, Honglin, et al.
Published: (2025) -
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
by: Pei, Qizhi, et al.
Published: (2025)