Automatic Dataset Construction (ADC): Sample Collection, Data Curation, and Beyond
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Minghao, Di, Zonglin, Wei, Jiaheng, Wang, Zhongruo, Zhang, Hengxiang, Xiao, Ruixuan, Wang, Haoyu, Pang, Jinlong, Chen, Hao, Shah, Ankit, Wei, Hongxin, He, Xinlei, Zhao, Zhaowei, Wang, Haobo, Feng, Lei, Wang, Jindong, Davis, James, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Data Efficiency via Curating LLM-Driven Rating Systems
by: Pang, Jinlong, et al.
Published: (2024)
by: Pang, Jinlong, et al.
Published: (2024)
LLM Unlearning via Loss Adjustment with Only Forget Data
by: Wang, Yaxuan, et al.
Published: (2024)
by: Wang, Yaxuan, et al.
Published: (2024)
Defending Membership Inference Attacks via Privacy-aware Sparsity Tuning
by: Hu, Qiang, et al.
Published: (2024)
by: Hu, Qiang, et al.
Published: (2024)
Inference Time Optimization with Confidence Dynamics
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
Human and AI Perceptual Differences in Image Classification Errors
by: Liu, Minghao, et al.
Published: (2023)
by: Liu, Minghao, et al.
Published: (2023)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
by: Long, Lin, et al.
Published: (2024)
by: Long, Lin, et al.
Published: (2024)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
by: Zhang, Hengxiang, et al.
Published: (2024)
by: Zhang, Hengxiang, et al.
Published: (2024)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
by: Chen, Hao, et al.
Published: (2023)
by: Chen, Hao, et al.
Published: (2023)
Detecting Distillation Data from Reasoning Models
by: Zhang, Hengxiang, et al.
Published: (2025)
by: Zhang, Hengxiang, et al.
Published: (2025)
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
by: Deng, Zhijie, et al.
Published: (2025)
by: Deng, Zhijie, et al.
Published: (2025)
Incentivizing High-quality Participation From Federated Learning Agents
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
Open-Vocabulary Calibration for Fine-tuned CLIP
by: Wang, Shuoyuan, et al.
Published: (2024)
by: Wang, Shuoyuan, et al.
Published: (2024)
Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Fairness Without Harm: An Influence-Guided Active Sampling Approach
by: Pang, Jinlong, et al.
Published: (2024)
by: Pang, Jinlong, et al.
Published: (2024)
AFPR-CIM: An Analog-Domain Floating-Point RRAM-based Compute-In-Memory Architecture with Dynamic Range Adaptive FP-ADC
by: Liu, Haobo, et al.
Published: (2024)
by: Liu, Haobo, et al.
Published: (2024)
Optimization-Free Test-Time Adaptation for Cross-Person Activity Recognition
by: Wang, Shuoyuan, et al.
Published: (2023)
by: Wang, Shuoyuan, et al.
Published: (2023)
Small-Margin Preferences Still Matter-If You Train Them Right
by: Pang, Jinlong, et al.
Published: (2026)
by: Pang, Jinlong, et al.
Published: (2026)
MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars
by: Wang, Shuoyuan, et al.
Published: (2026)
by: Wang, Shuoyuan, et al.
Published: (2026)
Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring
by: Wei, Siyuan, et al.
Published: (2025)
by: Wei, Siyuan, et al.
Published: (2025)
PAGE: Parametric Generative Explainer for Graph Neural Network
by: Qiu, Yang, et al.
Published: (2024)
by: Qiu, Yang, et al.
Published: (2024)
Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables
by: Geng, Xuzhao, et al.
Published: (2025)
by: Geng, Xuzhao, et al.
Published: (2025)
ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
An Efficient Wireless iBCI Headstage with Adaptive ADC Sample Rate
by: Liu, Hongyao, et al.
Published: (2026)
by: Liu, Hongyao, et al.
Published: (2026)
Impact of Noisy Supervision in Foundation Model Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
by: Yang, Haobo, et al.
Published: (2025)
by: Yang, Haobo, et al.
Published: (2025)
SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing
by: Liu, Qiuhao, et al.
Published: (2025)
by: Liu, Qiuhao, et al.
Published: (2025)
Understanding and Mitigating Miscalibration in Prompt Tuning for Vision-Language Models
by: Wang, Shuoyuan, et al.
Published: (2024)
by: Wang, Shuoyuan, et al.
Published: (2024)
LM-mixup: Text Data Augmentation via Language Model based Mixup
by: Deng, Zhijie, et al.
Published: (2025)
by: Deng, Zhijie, et al.
Published: (2025)
CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecasting
by: Li, Haobo, et al.
Published: (2024)
by: Li, Haobo, et al.
Published: (2024)
ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models
by: Zhang, Hengxiang, et al.
Published: (2024)
by: Zhang, Hengxiang, et al.
Published: (2024)
Human Detection in Realistic Through-the-Wall Environments using Raw Radar ADC Data and Parametric Neural Networks
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
The Impact of Bitcoin ETF Approval on Bitcoin's Hedging Properties Against Traditional Assets
by: Hong, Yihan, et al.
Published: (2025)
by: Hong, Yihan, et al.
Published: (2025)
A Euclidean Distance Matrix Model for Convex Clustering
by: Wang, Zhaowei, et al.
Published: (2021)
by: Wang, Zhaowei, et al.
Published: (2021)
Self-Ensemble Post Learning for Noisy Domain Generalization
by: Lu, Wang, et al.
Published: (2025)
by: Lu, Wang, et al.
Published: (2025)
H2O2 Self‐Supplying CaO2/CuO2/Fe3O4 Nanoplatform for Enhanced Chemodynamic Therapy of Cancer Cells
by: Bojian Liu, et al.
Published: (2024)
by: Bojian Liu, et al.
Published: (2024)
An efficient evolutionary structural optimization method for multi-resolution designs
by: Wang, Hongxin, et al.
Published: (2019)
by: Wang, Hongxin, et al.
Published: (2019)
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
DAMBench: A Multi-Modal Benchmark for Deep Learning-based Atmospheric Data Assimilation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
RePST: Language Model Empowered Spatio-Temporal Forecasting via Semantic-Oriented Reprogramming
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Similar Items
-
Improving Data Efficiency via Curating LLM-Driven Rating Systems
by: Pang, Jinlong, et al.
Published: (2024) -
LLM Unlearning via Loss Adjustment with Only Forget Data
by: Wang, Yaxuan, et al.
Published: (2024) -
Defending Membership Inference Attacks via Privacy-aware Sparsity Tuning
by: Hu, Qiang, et al.
Published: (2024) -
Inference Time Optimization with Confidence Dynamics
by: Wang, Yu, et al.
Published: (2026) -
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)