Task Oriented In-Domain Data Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Xiao, Hu, Xinyu, Zuo, Simiao, Gong, Yeyun, Lou, Qiang, Liu, Yi, Huang, Shao-Lun, Jiao, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2026)
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2026)
DeepThink: Aligning Language Models with Domain-Specific User Intents
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues
von: Xue, Liang, et al.
Veröffentlicht: (2025)
von: Xue, Liang, et al.
Veröffentlicht: (2025)
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
von: Park, Sungjin, et al.
Veröffentlicht: (2024)
von: Park, Sungjin, et al.
Veröffentlicht: (2024)
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
von: Deng, Haoran, et al.
Veröffentlicht: (2025)
von: Deng, Haoran, et al.
Veröffentlicht: (2025)
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
Improving Data and Reward Design for Scientific Reasoning in Large Language Models
von: Chen, Zijie, et al.
Veröffentlicht: (2026)
von: Chen, Zijie, et al.
Veröffentlicht: (2026)
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
HiCaM: A Hierarchical-Causal Modification Framework for Long-Form Text Modification
von: Shi, Yuntao, et al.
Veröffentlicht: (2025)
von: Shi, Yuntao, et al.
Veröffentlicht: (2025)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Data Augmentation for Low-Resource Domain Tasks
von: Seo, Minju, et al.
Veröffentlicht: (2024)
von: Seo, Minju, et al.
Veröffentlicht: (2024)
Comparing Data Augmentation Methods for End-to-End Task-Oriented Dialog Systems
von: Vlachos, Christos, et al.
Veröffentlicht: (2024)
von: Vlachos, Christos, et al.
Veröffentlicht: (2024)
Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
von: He, Hongyi, et al.
Veröffentlicht: (2025)
von: He, Hongyi, et al.
Veröffentlicht: (2025)
Process-based Self-Rewarding Language Models
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
Boosting Biomedical Concept Extraction by Rule-Based Data Augmentation
von: Shao, Qiwei, et al.
Veröffentlicht: (2024)
von: Shao, Qiwei, et al.
Veröffentlicht: (2024)
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
von: Kim, Junu, et al.
Veröffentlicht: (2025)
von: Kim, Junu, et al.
Veröffentlicht: (2025)
Exploring the Mystery of Influential Data for Mathematical Reasoning
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
von: Ni, Xinzhe, et al.
Veröffentlicht: (2024)
CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2025)
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2025)
Thought-Path Contrastive Learning via Premise-Oriented Data Augmentation for Logical Reading Comprehension
von: Wang, Chenxu, et al.
Veröffentlicht: (2024)
von: Wang, Chenxu, et al.
Veröffentlicht: (2024)
Rho-1: Not All Tokens Are What You Need
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
Integrative Decoding: Improve Factuality via Implicit Self-consistency
von: Cheng, Yi, et al.
Veröffentlicht: (2024)
von: Cheng, Yi, et al.
Veröffentlicht: (2024)
APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning
von: Sun, Jiashuo, et al.
Veröffentlicht: (2022)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2022)
PROM: A Phrase-level Copying Mechanism with Pre-training for Abstractive Summarization
von: Ma, Xinbei, et al.
Veröffentlicht: (2023)
von: Ma, Xinbei, et al.
Veröffentlicht: (2023)
DOS: Dependency-Oriented Sampler for Masked Diffusion Language Models
von: Zhou, Xueyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xueyu, et al.
Veröffentlicht: (2026)
STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media
von: Xue, Liang, et al.
Veröffentlicht: (2026)
von: Xue, Liang, et al.
Veröffentlicht: (2026)
Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
von: Wan, Zhongwei, et al.
Veröffentlicht: (2022)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2022)
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
von: Wu, Yang, et al.
Veröffentlicht: (2024)
von: Wu, Yang, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model Performance with Gradient-Based Parameter Selection
von: Li, Haoling, et al.
Veröffentlicht: (2024)
von: Li, Haoling, et al.
Veröffentlicht: (2024)
Trust-Oriented Adaptive Guardrails for Large Language Models
von: Hu, Jinwei, et al.
Veröffentlicht: (2024)
von: Hu, Jinwei, et al.
Veröffentlicht: (2024)
Evaluating and Enhancing Out-of-Domain Generalization of Task-Oriented Dialog Systems for Task Completion without Turn-level Dialog Annotations
von: Mosharrof, Adib, et al.
Veröffentlicht: (2025)
von: Mosharrof, Adib, et al.
Veröffentlicht: (2025)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
von: Gou, Zhibin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder
von: Wang, Yaoxiang, et al.
Veröffentlicht: (2026) -
DeepThink: Aligning Language Models with Domain-Specific User Intents
von: Li, Yang, et al.
Veröffentlicht: (2025) -
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
von: Yang, Kailai, et al.
Veröffentlicht: (2025) -
MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues
von: Xue, Liang, et al.
Veröffentlicht: (2025) -
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
von: Park, Sungjin, et al.
Veröffentlicht: (2024)