Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Juhwan, Yun, Jungmin, Jin, Kyohoon, Kim, YoungBin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GPTs Are Multilingual Annotators for Sequence Generation Tasks
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Adverb Is the Key: Simple Text Data Augmentation with Adverb Deletion
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
von: Kim, Kyeonghyun, et al.
Veröffentlicht: (2025)
von: Kim, Kyeonghyun, et al.
Veröffentlicht: (2025)
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
von: Jin, Kyohoon, et al.
Veröffentlicht: (2025)
von: Jin, Kyohoon, et al.
Veröffentlicht: (2025)
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Models
von: Song, Sangmin, et al.
Veröffentlicht: (2025)
von: Song, Sangmin, et al.
Veröffentlicht: (2025)
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Colorful Cutout: Enhancing Image Data Augmentation with Curriculum Learning
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
SUMMPILOT: Bridging Efficiency and Customization for Interactive Summarization System
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
SoftEDA: Rethinking Rule-Based Data Augmentation with Soft Labels
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
von: Jin, Kyohoon, et al.
Veröffentlicht: (2024)
von: Jin, Kyohoon, et al.
Veröffentlicht: (2024)
AutoAugment Is What You Need: Enhancing Rule-based Augmentation Methods in Low-resource Regimes
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning
von: Kim, Jisu, et al.
Veröffentlicht: (2024)
von: Kim, Jisu, et al.
Veröffentlicht: (2024)
GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluation
von: Lee, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jeongsoo, et al.
Veröffentlicht: (2025)
Control Token with Dense Passage Retrieval
von: Lee, Juhwan, et al.
Veröffentlicht: (2024)
von: Lee, Juhwan, et al.
Veröffentlicht: (2024)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
Cleanse: Uncertainty Estimation Approach Using Clustering-based Semantic Consistency in LLMs
von: Joo, Minsuh, et al.
Veröffentlicht: (2025)
von: Joo, Minsuh, et al.
Veröffentlicht: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
Relation-based Counterfactual Data Augmentation and Contrastive Learning for Robustifying Natural Language Inference Models
von: Yang, Heerin, et al.
Veröffentlicht: (2024)
von: Yang, Heerin, et al.
Veröffentlicht: (2024)
Position on LLM-Assisted Peer Review: Addressing Reviewer Gap through Mentoring and Feedback
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
Investigating Low-Cost LLM Annotation for~Spoken Dialogue Understanding Datasets
von: Druart, Lucas, et al.
Veröffentlicht: (2024)
von: Druart, Lucas, et al.
Veröffentlicht: (2024)
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
Cost-efficient Crowdsourcing for Span-based Sequence Labeling: Worker Selection and Data Augmentation
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
Emergent Convergence in Multi-Agent LLM Annotation
von: Parfenova, Angelina, et al.
Veröffentlicht: (2025)
von: Parfenova, Angelina, et al.
Veröffentlicht: (2025)
Cleansing the Artificial Mind: A Self-Reflective Detoxification Framework for Large Language Models
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
von: Zhang, Kaituo, et al.
Veröffentlicht: (2026)
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
von: Jin, Song, et al.
Veröffentlicht: (2025)
von: Jin, Song, et al.
Veröffentlicht: (2025)
Sina at FigNews 2024: Multilingual Datasets Annotated with Bias and Propaganda
von: Duaibes, Lina, et al.
Veröffentlicht: (2024)
von: Duaibes, Lina, et al.
Veröffentlicht: (2024)
Enhancing LLM-Based Data Annotation with Error Decomposition
von: Xu, Zhen, et al.
Veröffentlicht: (2026)
von: Xu, Zhen, et al.
Veröffentlicht: (2026)
Leveraging KV Similarity for Online Structured Pruning in LLMs
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
von: Liu, Maoqi, et al.
Veröffentlicht: (2025)
von: Liu, Maoqi, et al.
Veröffentlicht: (2025)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
von: Lee, Nayeon, et al.
Veröffentlicht: (2023)
von: Lee, Nayeon, et al.
Veröffentlicht: (2023)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
von: Ban, Minjeong, et al.
Veröffentlicht: (2026)
von: Ban, Minjeong, et al.
Veröffentlicht: (2026)
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
von: Liu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Liu, Wenxuan, et al.
Veröffentlicht: (2025)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
von: Kim, Taesu, et al.
Veröffentlicht: (2024)
von: Kim, Taesu, et al.
Veröffentlicht: (2024)
Cost-efficient Knowledge-based Question Answering with Large Language Models
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
MCFEND: A Multi-source Benchmark Dataset for Chinese Fake News Detection
von: Li, Yupeng, et al.
Veröffentlicht: (2024)
von: Li, Yupeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GPTs Are Multilingual Annotators for Sequence Generation Tasks
von: Choi, Juhwan, et al.
Veröffentlicht: (2024) -
Adverb Is the Key: Simple Text Data Augmentation with Adverb Deletion
von: Choi, Juhwan, et al.
Veröffentlicht: (2024) -
Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
von: Kim, Kyeonghyun, et al.
Veröffentlicht: (2025) -
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
von: Jin, Kyohoon, et al.
Veröffentlicht: (2025) -
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)