Real-Fake: Effective Training Data Synthesis Through Distribution Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Jianhao, Zhang, Jie, Sun, Shuyang, Torr, Philip, Zhao, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MatchDiffusion: Training-free Generation of Match-cuts
by: Pardo, Alejandro, et al.
Published: (2024)
by: Pardo, Alejandro, et al.
Published: (2024)
Gradient Residual Connections
by: Pan, Yangchen, et al.
Published: (2026)
by: Pan, Yangchen, et al.
Published: (2026)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
by: Aljaafari, Tala, et al.
Published: (2025)
by: Aljaafari, Tala, et al.
Published: (2025)
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
by: Yang, Puning, et al.
Published: (2026)
by: Yang, Puning, et al.
Published: (2026)
Fake Advertisements Detection Using Automated Multimodal Learning: A Case Study for Vietnamese Real Estate Data
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Fairness Through Matching
by: Kim, Kunwoong, et al.
Published: (2025)
by: Kim, Kunwoong, et al.
Published: (2025)
An MRP Formulation for Supervised Learning: Generalized Temporal Difference Learning Models
by: Pan, Yangchen, et al.
Published: (2024)
by: Pan, Yangchen, et al.
Published: (2024)
MDM: Advancing Multi-Domain Distribution Matching for Automatic Modulation Recognition Dataset Synthesis
by: Xu, Dongwei, et al.
Published: (2024)
by: Xu, Dongwei, et al.
Published: (2024)
DistDD: Distributed Data Distillation Aggregation through Gradient Matching
by: Wang, Peiran, et al.
Published: (2024)
by: Wang, Peiran, et al.
Published: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Integrating Distribution Matching into Semi-Supervised Contrastive Learning for Labeled and Unlabeled Data
by: Nakayama, Shogo, et al.
Published: (2026)
by: Nakayama, Shogo, et al.
Published: (2026)
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
by: Rose, Aaron, et al.
Published: (2026)
by: Rose, Aaron, et al.
Published: (2026)
kNN-CLIP: Retrieval Enables Training-Free Segmentation on Continually Expanding Large Vocabularies
by: Gui, Zhongrui, et al.
Published: (2024)
by: Gui, Zhongrui, et al.
Published: (2024)
Regurgitative Training: The Value of Real Data in Training Large Language Models
by: Zhang, Jinghui, et al.
Published: (2024)
by: Zhang, Jinghui, et al.
Published: (2024)
Counterfactual Explanations for Continuous Action Reinforcement Learning
by: Dong, Shuyang, et al.
Published: (2025)
by: Dong, Shuyang, et al.
Published: (2025)
Enhancing Stability in Training Conditional Generative Adversarial Networks via Selective Data Matching
by: Kong, Kyeongbo, et al.
Published: (2021)
by: Kong, Kyeongbo, et al.
Published: (2021)
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
by: Liu, Tong, et al.
Published: (2026)
by: Liu, Tong, et al.
Published: (2026)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Distribution Matching for Self-Supervised Transfer Learning
by: Jiao, Yuling, et al.
Published: (2025)
by: Jiao, Yuling, et al.
Published: (2025)
MALT: Improving Reasoning with Multi-Agent LLM Training
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
by: Huang, Jianhao, et al.
Published: (2026)
by: Huang, Jianhao, et al.
Published: (2026)
Federated Generative Learning with Foundation Models
by: Zhang, Jie, et al.
Published: (2023)
by: Zhang, Jie, et al.
Published: (2023)
Prompting a Pretrained Transformer Can Be a Universal Approximator
by: Petrov, Aleksandar, et al.
Published: (2024)
by: Petrov, Aleksandar, et al.
Published: (2024)
Learning to Explore: Policy-Guided Outlier Synthesis for Graph Out-of-Distribution Detection
by: Sun, Li, et al.
Published: (2026)
by: Sun, Li, et al.
Published: (2026)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
by: Zhang, Wenxuan, et al.
Published: (2024)
by: Zhang, Wenxuan, et al.
Published: (2024)
FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
by: Cui, Jin, et al.
Published: (2025)
by: Cui, Jin, et al.
Published: (2025)
Imitation Learning as Return Distribution Matching
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
EDGE: Efficient Data Selection for LLM Agents via Guideline Effectiveness
by: Zhang, Yunxiao, et al.
Published: (2025)
by: Zhang, Yunxiao, et al.
Published: (2025)
Fast-DataShapley: Neural Modeling for Training Data Valuation
by: Sun, Haifeng, et al.
Published: (2025)
by: Sun, Haifeng, et al.
Published: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
by: Zhu, Xuekai, et al.
Published: (2025)
by: Zhu, Xuekai, et al.
Published: (2025)
Distributional Training Data Attribution: What do Influence Functions Sample?
by: Mlodozeniec, Bruno, et al.
Published: (2025)
by: Mlodozeniec, Bruno, et al.
Published: (2025)
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach
by: Chen, Xiaobing, et al.
Published: (2025)
by: Chen, Xiaobing, et al.
Published: (2025)
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
by: Tan, Zelin, et al.
Published: (2025)
by: Tan, Zelin, et al.
Published: (2025)
Distribution Matching via Generalized Consistency Models
by: Shrestha, Sagar, et al.
Published: (2025)
by: Shrestha, Sagar, et al.
Published: (2025)
EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models
by: Kim, Jinhee, et al.
Published: (2024)
by: Kim, Jinhee, et al.
Published: (2024)
Training Free Guided Flow Matching with Optimal Control
by: Wang, Luran, et al.
Published: (2024)
by: Wang, Luran, et al.
Published: (2024)
Human in the Latent Loop (HILL): Interactively Guiding Model Training Through Human Intuition
by: Geissler, Daniel, et al.
Published: (2025)
by: Geissler, Daniel, et al.
Published: (2025)
Similar Items
-
MatchDiffusion: Training-free Generation of Match-cuts
by: Pardo, Alejandro, et al.
Published: (2024) -
Gradient Residual Connections
by: Pan, Yangchen, et al.
Published: (2026) -
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
by: Aljaafari, Tala, et al.
Published: (2025) -
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
by: Yang, Puning, et al.
Published: (2026) -
Fake Advertisements Detection Using Automated Multimodal Learning: A Case Study for Vietnamese Real Estate Data
by: Nguyen, Duy, et al.
Published: (2025)