DataSculpt: Crafting Data Landscapes for Long-Context LLMs through Multi-Objective Partitioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Keer, Nie, Xiaonan, Liang, Zheng, Pan, Da, Zhang, Shusen, Zhao, Keshi, Chen, Weipeng, Zhou, Zenan, Dong, Guosheng, Cui, Bin, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs
von: Lu, Keer, et al.
Veröffentlicht: (2024)
von: Lu, Keer, et al.
Veröffentlicht: (2024)
Med-R$^2$: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine
von: Lu, Keer, et al.
Veröffentlicht: (2025)
von: Lu, Keer, et al.
Veröffentlicht: (2025)
Data Proportion Detection for Optimized Data Management for Large Language Models
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning
von: Lu, Keer, et al.
Veröffentlicht: (2025)
von: Lu, Keer, et al.
Veröffentlicht: (2025)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
von: Li, Youquan, et al.
Veröffentlicht: (2024)
von: Li, Youquan, et al.
Veröffentlicht: (2024)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline
von: Dong, Guosheng, et al.
Veröffentlicht: (2024)
von: Dong, Guosheng, et al.
Veröffentlicht: (2024)
PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization
von: Zhang, Jiajun, et al.
Veröffentlicht: (2025)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2025)
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System
von: Zheng, Miao, et al.
Veröffentlicht: (2024)
von: Zheng, Miao, et al.
Veröffentlicht: (2024)
FedSC: Provable Federated Self-supervised Learning with Spectral Contrastive Objective over Non-i.i.d. Data
von: Jing, Shusen, et al.
Veröffentlicht: (2024)
von: Jing, Shusen, et al.
Veröffentlicht: (2024)
PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
von: Lu, Keer, et al.
Veröffentlicht: (2025)
von: Lu, Keer, et al.
Veröffentlicht: (2025)
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Diet‐Induced Developmental and Morphological Plasticity in a Thelytokous Predatory Mite Amblyseius herbicolus (Chant) (Acari: Phytoseiidae)
von: Keshi Zhang, et al.
Veröffentlicht: (2025)
von: Keshi Zhang, et al.
Veröffentlicht: (2025)
MatWheel: Addressing Data Scarcity in Materials Science Through Synthetic Data
von: Li, Wentao, et al.
Veröffentlicht: (2025)
von: Li, Wentao, et al.
Veröffentlicht: (2025)
LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
von: Jia, Junlong, et al.
Veröffentlicht: (2025)
von: Jia, Junlong, et al.
Veröffentlicht: (2025)
Mastering the Craft of Data Synthesis for CodeLLMs
von: Chen, Meng, et al.
Veröffentlicht: (2024)
von: Chen, Meng, et al.
Veröffentlicht: (2024)
PLATE 11 in Ontogenetic stages of Czenspinskia transversostriata (Oudemans) (Acari: Winterschmidtiidae)
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
Czenspinskia transversostriata
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
PLATE 10 in Ontogenetic stages of Czenspinskia transversostriata (Oudemans) (Acari: Winterschmidtiidae)
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
PLATE 8 in Ontogenetic stages of Czenspinskia transversostriata (Oudemans) (Acari: Winterschmidtiidae)
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
von: Zhang, Keshi, et al.
Veröffentlicht: (2025)
Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting
von: Schaldenbrand, Peter, et al.
Veröffentlicht: (2026)
von: Schaldenbrand, Peter, et al.
Veröffentlicht: (2026)
SysBench: Can Large Language Models Follow System Messages?
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning
von: Lai, Peichao, et al.
Veröffentlicht: (2024)
von: Lai, Peichao, et al.
Veröffentlicht: (2024)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
von: Tian, Junfeng, et al.
Veröffentlicht: (2024)
von: Tian, Junfeng, et al.
Veröffentlicht: (2024)
Alzheimer's Disease Common Data Element Ontology for Clinical Trials (AD-CDO)
von: Sun, Zenan
Veröffentlicht: (2026)
von: Sun, Zenan
Veröffentlicht: (2026)
Phototactic behavior and oviposition of seven species of Phytoseiidae (Acari: Mesostigmata)
von: Zhenguo Liu, et al.
Veröffentlicht: (2024)
von: Zhenguo Liu, et al.
Veröffentlicht: (2024)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
TopoSculpt: Betti-Steered Topological Sculpting of 3D Fine-grained Tubular Shapes
von: Zhang, Minghui, et al.
Veröffentlicht: (2025)
von: Zhang, Minghui, et al.
Veröffentlicht: (2025)
Sculpting Spin-Wave Landscapes via Curvature of 2D Magnonic Crystals
von: Wojewoda, Ondřej, et al.
Veröffentlicht: (2026)
von: Wojewoda, Ondřej, et al.
Veröffentlicht: (2026)
Rethinking Layer-wise Gaussian Noise Injection: Bridging Implicit Objectives and Privacy Budget Allocation
von: Tan, Qifeng, et al.
Veröffentlicht: (2025)
von: Tan, Qifeng, et al.
Veröffentlicht: (2025)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
von: Zhang, Chuyifei, et al.
Veröffentlicht: (2026)
von: Zhang, Chuyifei, et al.
Veröffentlicht: (2026)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Baichuan-Omni Technical Report
von: Li, Yadong, et al.
Veröffentlicht: (2024)
von: Li, Yadong, et al.
Veröffentlicht: (2024)
Sculpting Mechanical Properties of Hydrogels by Patterning Seamlessly Interlocked Stiff Skeleton
von: Bin Zhu, et al.
Veröffentlicht: (2024)
von: Bin Zhu, et al.
Veröffentlicht: (2024)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic Data
von: Men, Yifang, et al.
Veröffentlicht: (2024)
von: Men, Yifang, et al.
Veröffentlicht: (2024)
Effective In-Context Example Selection through Data Compression
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs
von: Lu, Keer, et al.
Veröffentlicht: (2024) -
Med-R$^2$: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine
von: Lu, Keer, et al.
Veröffentlicht: (2025) -
Data Proportion Detection for Optimized Data Management for Large Language Models
von: Liang, Hao, et al.
Veröffentlicht: (2024) -
Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning
von: Lu, Keer, et al.
Veröffentlicht: (2025) -
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
von: Li, Youquan, et al.
Veröffentlicht: (2024)