SoftEDA: Rethinking Rule-Based Data Augmentation with Soft Labels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Juhwan, Jin, Kyohoon, Lee, Junho, Song, Sangmin, Kim, Youngbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoAugment Is What You Need: Enhancing Rule-based Augmentation Methods in Low-resource Regimes
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
von: Jin, Kyohoon, et al.
Veröffentlicht: (2024)
von: Jin, Kyohoon, et al.
Veröffentlicht: (2024)
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
von: Jin, Kyohoon, et al.
Veröffentlicht: (2025)
von: Jin, Kyohoon, et al.
Veröffentlicht: (2025)
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
GPTs Are Multilingual Annotators for Sequence Generation Tasks
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
von: Kim, Kyeonghyun, et al.
Veröffentlicht: (2025)
von: Kim, Kyeonghyun, et al.
Veröffentlicht: (2025)
Adverb Is the Key: Simple Text Data Augmentation with Adverb Deletion
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Models
von: Song, Sangmin, et al.
Veröffentlicht: (2025)
von: Song, Sangmin, et al.
Veröffentlicht: (2025)
Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
von: Choi, Eunseong, et al.
Veröffentlicht: (2025)
von: Choi, Eunseong, et al.
Veröffentlicht: (2025)
Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning
von: Kim, Jisu, et al.
Veröffentlicht: (2024)
von: Kim, Jisu, et al.
Veröffentlicht: (2024)
SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluation
von: Lee, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jeongsoo, et al.
Veröffentlicht: (2025)
Control Token with Dense Passage Retrieval
von: Lee, Juhwan, et al.
Veröffentlicht: (2024)
von: Lee, Juhwan, et al.
Veröffentlicht: (2024)
Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective
von: Liu, Yunhao, et al.
Veröffentlicht: (2026)
von: Liu, Yunhao, et al.
Veröffentlicht: (2026)
Don't be a Fool: Pooling Strategies in Offensive Language Detection from User-Intended Adversarial Attacks
von: Yu, Seunguk, et al.
Veröffentlicht: (2024)
von: Yu, Seunguk, et al.
Veröffentlicht: (2024)
Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language Models
von: Yu, Seunguk, et al.
Veröffentlicht: (2025)
von: Yu, Seunguk, et al.
Veröffentlicht: (2025)
Geometric-Averaged Preference Optimization for Soft Preference Labels
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method
von: Kim, Mihyeon, et al.
Veröffentlicht: (2025)
von: Kim, Mihyeon, et al.
Veröffentlicht: (2025)
Improving Neural Topic Modeling with Semantically-Grounded Soft Label Distributions
von: Li, Raymond, et al.
Veröffentlicht: (2026)
von: Li, Raymond, et al.
Veröffentlicht: (2026)
Colorful Cutout: Enhancing Image Data Augmentation with Curriculum Learning
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Embracing Diversity: A Multi-Perspective Approach with Soft Labels
von: Muscato, Benedetta, et al.
Veröffentlicht: (2025)
von: Muscato, Benedetta, et al.
Veröffentlicht: (2025)
Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
von: Lee, Hojae, et al.
Veröffentlicht: (2024)
von: Lee, Hojae, et al.
Veröffentlicht: (2024)
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
von: Song, Hwanjun, et al.
Veröffentlicht: (2025)
von: Song, Hwanjun, et al.
Veröffentlicht: (2025)
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
von: Pavlovic, Maja, et al.
Veröffentlicht: (2026)
von: Pavlovic, Maja, et al.
Veröffentlicht: (2026)
FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation
von: Jang, Jinhee, et al.
Veröffentlicht: (2026)
von: Jang, Jinhee, et al.
Veröffentlicht: (2026)
SUMMPILOT: Bridging Efficiency and Customization for Interactive Summarization System
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling
von: Zheng, Guangmin, et al.
Veröffentlicht: (2024)
von: Zheng, Guangmin, et al.
Veröffentlicht: (2024)
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
von: Jin, Jiho, et al.
Veröffentlicht: (2025)
von: Jin, Jiho, et al.
Veröffentlicht: (2025)
Rethinking Retrieval-Augmented Generation as a Cooperative Decision-Making Problem
von: Song, Lichang, et al.
Veröffentlicht: (2026)
von: Song, Lichang, et al.
Veröffentlicht: (2026)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2026)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2026)
Topology-Aware Reasoning over Incomplete Knowledge Graph with Graph-Based Soft Prompting
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
von: He, Nan, et al.
Veröffentlicht: (2024)
von: He, Nan, et al.
Veröffentlicht: (2024)
ADEPT: Adaptive Dynamic Early-Exit Process for Transformers
von: Yoo, Sangmin, et al.
Veröffentlicht: (2026)
von: Yoo, Sangmin, et al.
Veröffentlicht: (2026)
Soft Adaptive Policy Optimization
von: Gao, Chang, et al.
Veröffentlicht: (2025)
von: Gao, Chang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AutoAugment Is What You Need: Enhancing Rule-based Augmentation Methods in Low-resource Regimes
von: Choi, Juhwan, et al.
Veröffentlicht: (2024) -
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
von: Jin, Kyohoon, et al.
Veröffentlicht: (2024) -
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
von: Jin, Kyohoon, et al.
Veröffentlicht: (2025) -
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
von: Choi, Juhwan, et al.
Veröffentlicht: (2024) -
GPTs Are Multilingual Annotators for Sequence Generation Tasks
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)