DUET: Optimizing Training Data Mixtures via Feedback from Unseen Evaluation Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhiliang, Lau, Gregory Kang Ruey, Foo, Chuan-Sheng, Low, Bryan Kian Hsiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
The Chicken and Egg Dilemma: Co-optimizing Data and Model Configurations for LLMs
von: Chen, Zhiliang, et al.
Veröffentlicht: (2026)
von: Chen, Zhiliang, et al.
Veröffentlicht: (2026)
PIED: Physics-Informed Experimental Design for Inverse Problems
von: Hemachandra, Apivich, et al.
Veröffentlicht: (2025)
von: Hemachandra, Apivich, et al.
Veröffentlicht: (2025)
PINNACLE: PINN Adaptive ColLocation and Experimental points selection
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model Predictions
von: Wang, Jingtan, et al.
Veröffentlicht: (2024)
von: Wang, Jingtan, et al.
Veröffentlicht: (2024)
BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks
von: Chew, Ruth Wan Theng, et al.
Veröffentlicht: (2026)
von: Chew, Ruth Wan Theng, et al.
Veröffentlicht: (2026)
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2026)
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2026)
Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)
Source Attribution for Large Language Model-Generated Data
von: Wang, Jingtan, et al.
Veröffentlicht: (2023)
von: Wang, Jingtan, et al.
Veröffentlicht: (2023)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
von: Verma, Arun, et al.
Veröffentlicht: (2024)
von: Verma, Arun, et al.
Veröffentlicht: (2024)
Prompt Optimization with Human Feedback
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2024)
Dependency Structure Search Bayesian Optimization for Decision Making Models
von: Rajpal, Mohit, et al.
Veröffentlicht: (2023)
von: Rajpal, Mohit, et al.
Veröffentlicht: (2023)
Data value estimation on private gradients
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
Fine-tuning Language Models with Generative Adversarial Reward Modelling
von: Yu, Zhang Ze, et al.
Veröffentlicht: (2023)
von: Yu, Zhang Ze, et al.
Veröffentlicht: (2023)
DeRDaVa: Deletion-Robust Data Valuation for Machine Learning
von: Tian, Xiao, et al.
Veröffentlicht: (2023)
von: Tian, Xiao, et al.
Veröffentlicht: (2023)
Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies
von: Verma, Arun, et al.
Veröffentlicht: (2024)
von: Verma, Arun, et al.
Veröffentlicht: (2024)
COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
von: Zhou, Zijian, et al.
Veröffentlicht: (2024)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
Data Distribution Valuation
von: Xu, Xinyi, et al.
Veröffentlicht: (2024)
von: Xu, Xinyi, et al.
Veröffentlicht: (2024)
INO-SGD: Addressing Utility Imbalance under Individualized Differential Privacy
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Localized Zeroth-Order Prompt Optimization
von: Hu, Wenyang, et al.
Veröffentlicht: (2024)
von: Hu, Wenyang, et al.
Veröffentlicht: (2024)
BarrierSteer: LLM Safety via Learning Barrier Steering
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
Paid with Models: Optimal Contract Design for Collaborative Machine Learning
von: Wang, Bingchen, et al.
Veröffentlicht: (2024)
von: Wang, Bingchen, et al.
Veröffentlicht: (2024)
Group-robust Sample Reweighting for Subpopulation Shifts via Influence Functions
von: Qiao, Rui, et al.
Veröffentlicht: (2025)
von: Qiao, Rui, et al.
Veröffentlicht: (2025)
REFRAG: Rethinking RAG based Decoding
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2025)
ReasonIR: Training Retrievers for Reasoning Tasks
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models
von: Shu, Yao, et al.
Veröffentlicht: (2024)
von: Shu, Yao, et al.
Veröffentlicht: (2024)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
von: Hu, Haoyu, et al.
Veröffentlicht: (2026)
von: Hu, Haoyu, et al.
Veröffentlicht: (2026)
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
von: Chen, Jiangwei, et al.
Veröffentlicht: (2026)
von: Chen, Jiangwei, et al.
Veröffentlicht: (2026)
On Newton's Method to Unlearn Neural Networks
von: Bui, Nhung, et al.
Veröffentlicht: (2024)
von: Bui, Nhung, et al.
Veröffentlicht: (2024)
DUET: Agentic Design Understanding via Experimentation and Testing
von: Smith, Gus Henry, et al.
Veröffentlicht: (2025)
von: Smith, Gus Henry, et al.
Veröffentlicht: (2025)
Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2024)
Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning
von: Sim, Rachael Hwee Ling, et al.
Veröffentlicht: (2026)
von: Sim, Rachael Hwee Ling, et al.
Veröffentlicht: (2026)
PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs
von: Zhang, Ze Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Ze Yu, et al.
Veröffentlicht: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
von: Chen, Zhiliang, et al.
Veröffentlicht: (2025) -
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024) -
The Chicken and Egg Dilemma: Co-optimizing Data and Model Configurations for LLMs
von: Chen, Zhiliang, et al.
Veröffentlicht: (2026) -
PIED: Physics-Informed Experimental Design for Inverse Problems
von: Hemachandra, Apivich, et al.
Veröffentlicht: (2025) -
PINNACLE: PINN Adaptive ColLocation and Experimental points selection
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2024)