Generating Synthetic Oracle Datasets to Analyze Noise Impact: A Study on Building Function Classification Using Tweets
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Shanshan, Kruspe, Anna, Zhu, Xiaoxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harmonic Reasoning in Large Language Models
by: Kruspe, Anna
Published: (2024)
by: Kruspe, Anna
Published: (2024)
Towards detecting unanticipated bias in Large Language Models
by: Kruspe, Anna
Published: (2024)
by: Kruspe, Anna
Published: (2024)
Musical ethnocentrism in Large Language Models
by: Kruspe, Anna
Published: (2025)
by: Kruspe, Anna
Published: (2025)
More than words: Advancements and challenges in speech recognition for singing
by: Kruspe, Anna
Published: (2024)
by: Kruspe, Anna
Published: (2024)
Joint sentiment analysis of lyrics and audio in music
by: Schaab, Lea, et al.
Published: (2024)
by: Schaab, Lea, et al.
Published: (2024)
OSINT or BULLSHINT? Exploring Open-Source Intelligence tweets about the Russo-Ukrainian War
by: Niu, Johannes, et al.
Published: (2025)
by: Niu, Johannes, et al.
Published: (2025)
A Comparative Analysis of Machine Learning and Deep Learning Models for Tweet Sentiment Classification: A Case Study on the Sentiment140 Dataset
by: Anggraini, Vita, et al.
Published: (2026)
by: Anggraini, Vita, et al.
Published: (2026)
Semi-Synthetic Parallel Data for Translation Quality Estimation: A Case Study of Dataset Building for an Under-Resourced Language Pair
by: Siani, Assaf, et al.
Published: (2026)
by: Siani, Assaf, et al.
Published: (2026)
Zero-Shot Classification of Crisis Tweets Using Instruction-Finetuned Large Language Models
by: McDaniel, Emma, et al.
Published: (2024)
by: McDaniel, Emma, et al.
Published: (2024)
Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
by: Bagci, Mevlüt, et al.
Published: (2025)
by: Bagci, Mevlüt, et al.
Published: (2025)
Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
by: Zisad, Sharif Noor, et al.
Published: (2025)
by: Zisad, Sharif Noor, et al.
Published: (2025)
BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories
by: Ouyang, Yuxuan, et al.
Published: (2026)
by: Ouyang, Yuxuan, et al.
Published: (2026)
ADSumm: Annotated Ground-truth Summary Datasets for Disaster Tweet Summarization
by: Garg, Piyush Kumar, et al.
Published: (2024)
by: Garg, Piyush Kumar, et al.
Published: (2024)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
by: Liu, Peipei, et al.
Published: (2023)
by: Liu, Peipei, et al.
Published: (2023)
Synthetic News Generation for Fake News Classification
by: Sittar, Abdul, et al.
Published: (2025)
by: Sittar, Abdul, et al.
Published: (2025)
SenWave: A Fine-Grained Multi-Language Sentiment Analysis Dataset Sourced from COVID-19 Tweets
by: Yang, Qiang, et al.
Published: (2025)
by: Yang, Qiang, et al.
Published: (2025)
LT4SG@SMM4H24: Tweets Classification for Digital Epidemiology of Childhood Health Outcomes Using Pre-Trained Language Models
by: Athukoralage, Dasun, et al.
Published: (2024)
by: Athukoralage, Dasun, et al.
Published: (2024)
A dataset of Open Source Intelligence (OSINT) Tweets about the Russo-Ukrainian war
by: Niu, Johannes, et al.
Published: (2024)
by: Niu, Johannes, et al.
Published: (2024)
Utilising Large Language Models for Generating Effective Counter Arguments to Anti-Vaccine Tweets
by: Dhanuka, Utsav, et al.
Published: (2025)
by: Dhanuka, Utsav, et al.
Published: (2025)
Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated Tweets
by: Tuck, Bryan E., et al.
Published: (2024)
by: Tuck, Bryan E., et al.
Published: (2024)
A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers
by: Barriere, Valentin, et al.
Published: (2024)
by: Barriere, Valentin, et al.
Published: (2024)
MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets
by: Aboutalebi, Hossein, et al.
Published: (2024)
by: Aboutalebi, Hossein, et al.
Published: (2024)
The Synthetic Imputation Approach: Generating Optimal Synthetic Texts For Underrepresented Categories In Supervised Classification Tasks
by: Timoneda, Joan C.
Published: (2025)
by: Timoneda, Joan C.
Published: (2025)
An Annotated Corpus of Arabic Tweets for Hate Speech Analysis
by: Zaghouani, Wajdi, et al.
Published: (2025)
by: Zaghouani, Wajdi, et al.
Published: (2025)
Measuring Diversity in Synthetic Datasets
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
IsraParlTweet: The Israeli Parliamentary and Twitter Resource
by: Mor-Lan, Guy, et al.
Published: (2024)
by: Mor-Lan, Guy, et al.
Published: (2024)
MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection
by: Toraman, Cagri, et al.
Published: (2022)
by: Toraman, Cagri, et al.
Published: (2022)
Harnessing LLMs for API Interactions: A Framework for Classification and Synthetic Data Generation
by: Tao, Chunliang, et al.
Published: (2024)
by: Tao, Chunliang, et al.
Published: (2024)
Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties
by: Pham, Nhi, et al.
Published: (2024)
by: Pham, Nhi, et al.
Published: (2024)
Cyber Risks of Machine Translation Critical Errors : Arabic Mental Health Tweets as a Case Study
by: Saadany, Hadeel, et al.
Published: (2024)
by: Saadany, Hadeel, et al.
Published: (2024)
Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset
by: Amalvy, Arthur, et al.
Published: (2023)
by: Amalvy, Arthur, et al.
Published: (2023)
Analyzing Dataset Annotation Quality Management in the Wild
by: Klie, Jan-Christoph, et al.
Published: (2023)
by: Klie, Jan-Christoph, et al.
Published: (2023)
Who Attacks, and Why? Using LLMs to Identify Negative Campaigning in 18M Tweets across 19 Countries
by: Hartman, Victor, et al.
Published: (2025)
by: Hartman, Victor, et al.
Published: (2025)
ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
by: M., Yashwanth, et al.
Published: (2025)
by: M., Yashwanth, et al.
Published: (2025)
Synthetic Dialogue Dataset Generation using LLM Agents
by: Abdullin, Yelaman, et al.
Published: (2024)
by: Abdullin, Yelaman, et al.
Published: (2024)
Generating Synthetic Datasets for Few-shot Prompt Tuning
by: Guo, Xu, et al.
Published: (2024)
by: Guo, Xu, et al.
Published: (2024)
Deciphering Oracle Bone Language with Diffusion Models
by: Guan, Haisu, et al.
Published: (2024)
by: Guan, Haisu, et al.
Published: (2024)
ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media
by: Huang, Kung-Hsiang, et al.
Published: (2023)
by: Huang, Kung-Hsiang, et al.
Published: (2023)
EquiSumm : A Gender Bias-Aware Framework for Inclusive Tweet Summarization
by: Wanjari, Chaitanya, et al.
Published: (2026)
by: Wanjari, Chaitanya, et al.
Published: (2026)
Impact of Preference Noise on the Alignment Performance of Generative Language Models
by: Gao, Yang, et al.
Published: (2024)
by: Gao, Yang, et al.
Published: (2024)
Similar Items
-
Harmonic Reasoning in Large Language Models
by: Kruspe, Anna
Published: (2024) -
Towards detecting unanticipated bias in Large Language Models
by: Kruspe, Anna
Published: (2024) -
Musical ethnocentrism in Large Language Models
by: Kruspe, Anna
Published: (2025) -
More than words: Advancements and challenges in speech recognition for singing
by: Kruspe, Anna
Published: (2024) -
Joint sentiment analysis of lyrics and audio in music
by: Schaab, Lea, et al.
Published: (2024)