Data-Centric AI in the Age of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xinyi, Wu, Zhaoxuan, Qiao, Rui, Verma, Arun, Shu, Yao, Wang, Jingtan, Niu, Xinyuan, He, Zhenfeng, Chen, Jiangwei, Zhou, Zijian, Lau, Gregory Kang Ruey, Dao, Hieu, Agussurja, Lucas, Sim, Rachael Hwee Ling, Lin, Xiaoqiang, Hu, Wenyang, Dai, Zhongxiang, Koh, Pang Wei, Low, Bryan Kian Hsiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
by: Verma, Arun, et al.
Published: (2025)
by: Verma, Arun, et al.
Published: (2025)
Incentivizing Time-Aware Fairness in Data Sharing
by: Chen, Jiangwei, et al.
Published: (2025)
by: Chen, Jiangwei, et al.
Published: (2025)
Group-robust Sample Reweighting for Subpopulation Shifts via Influence Functions
by: Qiao, Rui, et al.
Published: (2025)
by: Qiao, Rui, et al.
Published: (2025)
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
by: Chen, Jiangwei, et al.
Published: (2026)
by: Chen, Jiangwei, et al.
Published: (2026)
Active Human Feedback Collection via Neural Contextual Dueling Bandits
by: Verma, Arun, et al.
Published: (2025)
by: Verma, Arun, et al.
Published: (2025)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
Global-to-Local Support Spectrums for Language Model Explainability
by: Agussurja, Lucas, et al.
Published: (2024)
by: Agussurja, Lucas, et al.
Published: (2024)
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
by: Lau, Gregory Kang Ruey, et al.
Published: (2026)
by: Lau, Gregory Kang Ruey, et al.
Published: (2026)
DeRDaVa: Deletion-Robust Data Valuation for Machine Learning
by: Tian, Xiao, et al.
Published: (2023)
by: Tian, Xiao, et al.
Published: (2023)
INO-SGD: Addressing Utility Imbalance under Individualized Differential Privacy
by: Tian, Xiao, et al.
Published: (2026)
by: Tian, Xiao, et al.
Published: (2026)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
by: Wu, Zhaoxuan, et al.
Published: (2025)
by: Wu, Zhaoxuan, et al.
Published: (2025)
Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars
by: Wu, Zhaoxuan, et al.
Published: (2024)
by: Wu, Zhaoxuan, et al.
Published: (2024)
Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers
by: Lin, Xiaoqiang, et al.
Published: (2023)
by: Lin, Xiaoqiang, et al.
Published: (2023)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Prompt Optimization with Human Feedback
by: Lin, Xiaoqiang, et al.
Published: (2024)
by: Lin, Xiaoqiang, et al.
Published: (2024)
Robustifying and Boosting Training-Free Neural Architecture Search
by: He, Zhenfeng, et al.
Published: (2024)
by: He, Zhenfeng, et al.
Published: (2024)
Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning
by: Sim, Rachael Hwee Ling, et al.
Published: (2026)
by: Sim, Rachael Hwee Ling, et al.
Published: (2026)
Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model Predictions
by: Wang, Jingtan, et al.
Published: (2024)
by: Wang, Jingtan, et al.
Published: (2024)
MineDraft: A Framework for Batch Parallel Speculative Decoding
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Incentives in Private Collaborative Machine Learning
by: Sim, Rachael Hwee Ling, et al.
Published: (2024)
by: Sim, Rachael Hwee Ling, et al.
Published: (2024)
On Newton's Method to Unlearn Neural Networks
by: Bui, Nhung, et al.
Published: (2024)
by: Bui, Nhung, et al.
Published: (2024)
Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Localized Zeroth-Order Prompt Optimization
by: Hu, Wenyang, et al.
Published: (2024)
by: Hu, Wenyang, et al.
Published: (2024)
DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
WaterDrum: Watermarking for Data-centric Unlearning Metric
by: Lu, Xinyang, et al.
Published: (2025)
by: Lu, Xinyang, et al.
Published: (2025)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
by: Tian, Xiao, et al.
Published: (2026)
by: Tian, Xiao, et al.
Published: (2026)
Self-Interested Agents in Collaborative Machine Learning: An Incentivized Adaptive Data-Centric Framework
by: Vijayan, Nithia, et al.
Published: (2024)
by: Vijayan, Nithia, et al.
Published: (2024)
Source Attribution for Large Language Model-Generated Data
by: Wang, Jingtan, et al.
Published: (2023)
by: Wang, Jingtan, et al.
Published: (2023)
PIED: Physics-Informed Experimental Design for Inverse Problems
by: Hemachandra, Apivich, et al.
Published: (2025)
by: Hemachandra, Apivich, et al.
Published: (2025)
DUET: Optimizing Training Data Mixtures via Feedback from Unseen Evaluation Tasks
by: Chen, Zhiliang, et al.
Published: (2025)
by: Chen, Zhiliang, et al.
Published: (2025)
PINNACLE: PINN Adaptive ColLocation and Experimental points selection
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Data value estimation on private gradients
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents
by: Verma, Arun, et al.
Published: (2025)
by: Verma, Arun, et al.
Published: (2025)
DUPRE: Data Utility Prediction for Efficient Data Valuation
by: Pham, Kieu Thao Nguyen, et al.
Published: (2025)
by: Pham, Kieu Thao Nguyen, et al.
Published: (2025)
Paid with Models: Optimal Contract Design for Collaborative Machine Learning
by: Wang, Bingchen, et al.
Published: (2024)
by: Wang, Bingchen, et al.
Published: (2024)
Understanding the Relationship between Prompts and Response Uncertainty in Large Language Models
by: Zhang, Ze Yu, et al.
Published: (2024)
by: Zhang, Ze Yu, et al.
Published: (2024)
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
by: Chen, Zhiliang, et al.
Published: (2025)
by: Chen, Zhiliang, et al.
Published: (2025)
Adjusted Expected Improvement for Cumulative Regret Minimization in Noisy Bayesian Optimization
by: Hu, Shouri, et al.
Published: (2022)
by: Hu, Shouri, et al.
Published: (2022)
Similar Items
-
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
by: Lau, Gregory Kang Ruey, et al.
Published: (2024) -
Uncovering Scaling Laws for Large Language Models via Inverse Problems
by: Verma, Arun, et al.
Published: (2025) -
Incentivizing Time-Aware Fairness in Data Sharing
by: Chen, Jiangwei, et al.
Published: (2025) -
Group-robust Sample Reweighting for Subpopulation Shifts via Influence Functions
by: Qiao, Rui, et al.
Published: (2025) -
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
by: Chen, Jiangwei, et al.
Published: (2026)