A Pre-trained Data Deduplication Model based on Active Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Haochen, Liu, Xinyao, Lv, Fengmao, Xue, Hongtao, Hu, Jie, Du, Shengdong, Li, Tianrui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Contrastive Feature Representations for Facial Action Unit Detection
von: Shang, Ziqiao, et al.
Veröffentlicht: (2024)
von: Shang, Ziqiao, et al.
Veröffentlicht: (2024)
Machine Unlearning of Pre-trained Large Language Models
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
Feature Alignment: Rethinking Efficient Active Learning via Proxy in the Context of Pre-trained Models
von: Wen, Ziting, et al.
Veröffentlicht: (2024)
von: Wen, Ziting, et al.
Veröffentlicht: (2024)
Privacy-Preserving Data Deduplication for Enhancing Federated Learning of Language Models (Extended Version)
von: Abadi, Aydin, et al.
Veröffentlicht: (2024)
von: Abadi, Aydin, et al.
Veröffentlicht: (2024)
Dataset Ownership Verification in Contrastive Pre-trained Models
von: Xie, Yuechen, et al.
Veröffentlicht: (2025)
von: Xie, Yuechen, et al.
Veröffentlicht: (2025)
Pre-training with Synthetic Data Helps Offline Reinforcement Learning
von: Wang, Zecheng, et al.
Veröffentlicht: (2023)
von: Wang, Zecheng, et al.
Veröffentlicht: (2023)
Graph Generative Pre-trained Transformer
von: Chen, Xiaohui, et al.
Veröffentlicht: (2025)
von: Chen, Xiaohui, et al.
Veröffentlicht: (2025)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
Transfer Learning with Pre-trained Conditional Generative Models
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2022)
von: Yamaguchi, Shin'ya, et al.
Veröffentlicht: (2022)
DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
von: Guo, Ke, et al.
Veröffentlicht: (2025)
von: Guo, Ke, et al.
Veröffentlicht: (2025)
Unified View Imputation and Feature Selection Learning for Incomplete Multi-view Data
von: Huang, Yanyong, et al.
Veröffentlicht: (2024)
von: Huang, Yanyong, et al.
Veröffentlicht: (2024)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
von: Chen, Tong, et al.
Veröffentlicht: (2025)
von: Chen, Tong, et al.
Veröffentlicht: (2025)
GraphSculptor: Sculpting Pre-training Coreset for Graph Self-supervised Learning
von: Liu, Chuang, et al.
Veröffentlicht: (2026)
von: Liu, Chuang, et al.
Veröffentlicht: (2026)
An Efficient Replay for Class-Incremental Learning with Pre-trained Models
von: Yin, Weimin, et al.
Veröffentlicht: (2024)
von: Yin, Weimin, et al.
Veröffentlicht: (2024)
MOTO: Offline Pre-training to Online Fine-tuning for Model-based Robot Learning
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
Causally Sufficient and Necessary Feature Expansion for Class-Incremental Learning
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
Relational In-Context Learning via Synthetic Pre-training with Structural Prior
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
Enhanced Atrial Fibrillation Prediction in ESUS Patients with Hypergraph-based Pre-training
von: Xie, Yuzhang, et al.
Veröffentlicht: (2026)
von: Xie, Yuzhang, et al.
Veröffentlicht: (2026)
Cross-domain Random Pre-training with Prototypes for Reinforcement Learning
von: Liu, Xin, et al.
Veröffentlicht: (2023)
von: Liu, Xin, et al.
Veröffentlicht: (2023)
Pre-trained Large Language Models Learn Hidden Markov Models In-context
von: Dai, Yijia, et al.
Veröffentlicht: (2025)
von: Dai, Yijia, et al.
Veröffentlicht: (2025)
Parallel Structures in Pre-training Data Yield In-Context Learning
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
Utilizing Strategic Pre-training to Reduce Overfitting: Baguan -- A Pre-trained Weather Forecasting Model
von: Niu, Peisong, et al.
Veröffentlicht: (2025)
von: Niu, Peisong, et al.
Veröffentlicht: (2025)
Measuring Pre-training Data Quality without Labels for Time Series Foundation Models
von: Wen, Songkang, et al.
Veröffentlicht: (2024)
von: Wen, Songkang, et al.
Veröffentlicht: (2024)
Gradient-based Fine-Tuning through Pre-trained Model Regularization
von: Liu, Xuanbo, et al.
Veröffentlicht: (2025)
von: Liu, Xuanbo, et al.
Veröffentlicht: (2025)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
von: Song, Weixi, et al.
Veröffentlicht: (2023)
von: Song, Weixi, et al.
Veröffentlicht: (2023)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
KG-TREAT: Pre-training for Treatment Effect Estimation by Synergizing Patient Data with Knowledge Graphs
von: Liu, Ruoqi, et al.
Veröffentlicht: (2024)
von: Liu, Ruoqi, et al.
Veröffentlicht: (2024)
OSF: On Pre-training and Scaling of Sleep Foundation Models
von: Shuai, Zitao, et al.
Veröffentlicht: (2026)
von: Shuai, Zitao, et al.
Veröffentlicht: (2026)
Scaling Laws for Pre-training Agents and World Models
von: Pearce, Tim, et al.
Veröffentlicht: (2024)
von: Pearce, Tim, et al.
Veröffentlicht: (2024)
A Pre-training Framework for Relational Data with Information-theoretic Principles
von: Truong, Quang, et al.
Veröffentlicht: (2025)
von: Truong, Quang, et al.
Veröffentlicht: (2025)
Stack Trace Deduplication: Faster, More Accurately, and in More Realistic Scenarios
von: Shibaev, Egor, et al.
Veröffentlicht: (2024)
von: Shibaev, Egor, et al.
Veröffentlicht: (2024)
Subequivariant Reinforcement Learning in 3D Multi-Entity Physical Environments
von: Chen, Runfa, et al.
Veröffentlicht: (2024)
von: Chen, Runfa, et al.
Veröffentlicht: (2024)
Endowing Pre-trained Graph Models with Provable Fairness
von: Zhang, Zhongjian, et al.
Veröffentlicht: (2024)
von: Zhang, Zhongjian, et al.
Veröffentlicht: (2024)
Urban Region Pre-training and Prompting: A Graph-based Approach
von: Jin, Jiahui, et al.
Veröffentlicht: (2024)
von: Jin, Jiahui, et al.
Veröffentlicht: (2024)
Mochi: Aligning Pre-training and Inference for Efficient Graph Foundation Models via Meta-Learning
von: Mattos, João, et al.
Veröffentlicht: (2026)
von: Mattos, João, et al.
Veröffentlicht: (2026)
GPD-1: Generative Pre-training for Driving
von: Xie, Zixun, et al.
Veröffentlicht: (2024)
von: Xie, Zixun, et al.
Veröffentlicht: (2024)
3D Interaction Geometric Pre-training for Molecular Relational Learning
von: Lee, Namkyeong, et al.
Veröffentlicht: (2024)
von: Lee, Namkyeong, et al.
Veröffentlicht: (2024)
Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization
von: Hu, Shengchao, et al.
Veröffentlicht: (2024)
von: Hu, Shengchao, et al.
Veröffentlicht: (2024)
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
von: Zhou, Fan, et al.
Veröffentlicht: (2024)
von: Zhou, Fan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Contrastive Feature Representations for Facial Action Unit Detection
von: Shang, Ziqiao, et al.
Veröffentlicht: (2024) -
Machine Unlearning of Pre-trained Large Language Models
von: Yao, Jin, et al.
Veröffentlicht: (2024) -
Feature Alignment: Rethinking Efficient Active Learning via Proxy in the Context of Pre-trained Models
von: Wen, Ziting, et al.
Veröffentlicht: (2024) -
Privacy-Preserving Data Deduplication for Enhancing Federated Learning of Language Models (Extended Version)
von: Abadi, Aydin, et al.
Veröffentlicht: (2024) -
Dataset Ownership Verification in Contrastive Pre-trained Models
von: Xie, Yuechen, et al.
Veröffentlicht: (2025)