Assessing the Impact of the Quality of Textual Data on Feature Representation and Machine Learning Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sarwar, Tabinda, Yepes, Antonio Jose Jimeno, Cavedon, Lawrence |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Deep Contrastive Unlearning for Language Models
par: He, Estrid, et autres
Publié: (2025)
par: He, Estrid, et autres
Publié: (2025)
RADS: Reinforcement Learning-Based Sample Selection Improves Transfer Learning in Low-resource and Imbalanced Clinical Settings
par: Han, Wei, et autres
Publié: (2026)
par: Han, Wei, et autres
Publié: (2026)
Financial Report Chunking for Effective Retrieval Augmented Generation
par: Yepes, Antonio Jimeno, et autres
Publié: (2024)
par: Yepes, Antonio Jimeno, et autres
Publié: (2024)
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
par: Hirak, Vitalii, et autres
Publié: (2026)
par: Hirak, Vitalii, et autres
Publié: (2026)
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
par: Shajalal, Md, et autres
Publié: (2022)
par: Shajalal, Md, et autres
Publié: (2022)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
par: Sun, Kun, et autres
Publié: (2024)
par: Sun, Kun, et autres
Publié: (2024)
Disentangling Textual and Acoustic Features of Neural Speech Representations
par: Mohebbi, Hosein, et autres
Publié: (2024)
par: Mohebbi, Hosein, et autres
Publié: (2024)
SCORE: A Semantic Evaluation Framework for Generative Document Parsing
par: Li, Renyu, et autres
Publié: (2025)
par: Li, Renyu, et autres
Publié: (2025)
Assessing the Role of Data Quality in Training Bilingual Language Models
par: Seto, Skyler, et autres
Publié: (2025)
par: Seto, Skyler, et autres
Publié: (2025)
CLARITY: A Framework and Benchmark for Conversational Language Ambiguity and Unanswerability in Interactive NL2SQL Systems
par: Sarwar, Tabinda, et autres
Publié: (2026)
par: Sarwar, Tabinda, et autres
Publié: (2026)
Cross-lingual Transfer or Machine Translation? On Data Augmentation for Monolingual Semantic Textual Similarity
par: Hoshino, Sho, et autres
Publié: (2024)
par: Hoshino, Sho, et autres
Publié: (2024)
Exploration of Attention Mechanism-Enhanced Deep Learning Models in the Mining of Medical Textual Data
par: Xiao, Lingxi, et autres
Publié: (2024)
par: Xiao, Lingxi, et autres
Publié: (2024)
Eliciting Textual Descriptions from Representations of Continuous Prompts
par: Ramati, Dana, et autres
Publié: (2024)
par: Ramati, Dana, et autres
Publié: (2024)
VI-OOD: A Unified Representation Learning Framework for Textual Out-of-distribution Detection
par: Zhan, Li-Ming, et autres
Publié: (2024)
par: Zhan, Li-Ming, et autres
Publié: (2024)
Evolutionary thoughts: integration of large language models and evolutionary algorithms
par: Yepes, Antonio Jimeno, et autres
Publié: (2025)
par: Yepes, Antonio Jimeno, et autres
Publié: (2025)
Empowering Large Language Models for Textual Data Augmentation
par: Li, Yichuan, et autres
Publié: (2024)
par: Li, Yichuan, et autres
Publié: (2024)
Can Automatic Metrics Assess High-Quality Translations?
par: Agrawal, Sweta, et autres
Publié: (2024)
par: Agrawal, Sweta, et autres
Publié: (2024)
A Survey of Machine Learning Models and Datasets for the Multi-label Classification of Textual Hate Speech in English
par: Bäumler, Julian, et autres
Publié: (2025)
par: Bäumler, Julian, et autres
Publié: (2025)
Fooling the Textual Fooler via Randomizing Latent Representations
par: Hoang, Duy C., et autres
Publié: (2023)
par: Hoang, Duy C., et autres
Publié: (2023)
Integrating Large Language Models and Knowledge Graphs for Extraction and Validation of Textual Test Data
par: De Santis, Antonio, et autres
Publié: (2024)
par: De Santis, Antonio, et autres
Publié: (2024)
ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
par: Li, Ruihang, et autres
Publié: (2024)
par: Li, Ruihang, et autres
Publié: (2024)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
par: Sarwar, Nobin
Publié: (2025)
par: Sarwar, Nobin
Publié: (2025)
GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training
par: Nacar, Omer, et autres
Publié: (2025)
par: Nacar, Omer, et autres
Publié: (2025)
From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation
par: Zhao, Haofei, et autres
Publié: (2024)
par: Zhao, Haofei, et autres
Publié: (2024)
Potential and Perils of Large Language Models as Judges of Unstructured Textual Data
par: Bedemariam, Rewina, et autres
Publié: (2025)
par: Bedemariam, Rewina, et autres
Publié: (2025)
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
par: Hossain, Eftekhar, et autres
Publié: (2024)
par: Hossain, Eftekhar, et autres
Publié: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
par: Hua, Jiacheng, et autres
Publié: (2026)
par: Hua, Jiacheng, et autres
Publié: (2026)
Learning High-Quality and General-Purpose Phrase Representations
par: Chen, Lihu, et autres
Publié: (2024)
par: Chen, Lihu, et autres
Publié: (2024)
Enhancing Traffic Prediction with Textual Data Using Large Language Models
par: Huang, Xiannan
Publié: (2024)
par: Huang, Xiannan
Publié: (2024)
An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data
par: Yoffe, Gideon, et autres
Publié: (2025)
par: Yoffe, Gideon, et autres
Publié: (2025)
CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
par: Zhang, Bowen, et autres
Publié: (2025)
par: Zhang, Bowen, et autres
Publié: (2025)
Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models
par: Zhang, Gaifan, et autres
Publié: (2025)
par: Zhang, Gaifan, et autres
Publié: (2025)
Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
par: Lin, Tzu-Ling, et autres
Publié: (2025)
par: Lin, Tzu-Ling, et autres
Publié: (2025)
Tell Me What's Next: Textual Foresight for Generic UI Representations
par: Burns, Andrea, et autres
Publié: (2024)
par: Burns, Andrea, et autres
Publié: (2024)
New Textual Corpora for Serbian Language Modeling
par: Škorić, Mihailo, et autres
Publié: (2024)
par: Škorić, Mihailo, et autres
Publié: (2024)
CraftRTL: High-quality Synthetic Data Generation for Verilog Code Models with Correct-by-Construction Non-Textual Representations and Targeted Code Repair
par: Liu, Mingjie, et autres
Publié: (2024)
par: Liu, Mingjie, et autres
Publié: (2024)
Mashee at SemEval-2024 Task 8: The Impact of Samples Quality on the Performance of In-Context Learning for Machine Text Classification
par: Rasheed, Areeg Fahad, et autres
Publié: (2024)
par: Rasheed, Areeg Fahad, et autres
Publié: (2024)
Ordered Semantically Diverse Sampling for Textual Data
par: Tiwari, Ashish, et autres
Publié: (2025)
par: Tiwari, Ashish, et autres
Publié: (2025)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
par: Agrawal, Sweta, et autres
Publié: (2024)
par: Agrawal, Sweta, et autres
Publié: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
par: Hashemi, Mohammad Abuzar, et autres
Publié: (2021)
par: Hashemi, Mohammad Abuzar, et autres
Publié: (2021)
Documents similaires
-
Deep Contrastive Unlearning for Language Models
par: He, Estrid, et autres
Publié: (2025) -
RADS: Reinforcement Learning-Based Sample Selection Improves Transfer Learning in Low-resource and Imbalanced Clinical Settings
par: Han, Wei, et autres
Publié: (2026) -
Financial Report Chunking for Effective Retrieval Augmented Generation
par: Yepes, Antonio Jimeno, et autres
Publié: (2024) -
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
par: Hirak, Vitalii, et autres
Publié: (2026) -
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
par: Shajalal, Md, et autres
Publié: (2022)