Assessing the Impact of the Quality of Textual Data on Feature Representation and Machine Learning Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Sarwar, Tabinda, Yepes, Antonio Jose Jimeno, Cavedon, Lawrence |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Deep Contrastive Unlearning for Language Models
por: He, Estrid, et al.
Publicado: (2025)
por: He, Estrid, et al.
Publicado: (2025)
RADS: Reinforcement Learning-Based Sample Selection Improves Transfer Learning in Low-resource and Imbalanced Clinical Settings
por: Han, Wei, et al.
Publicado: (2026)
por: Han, Wei, et al.
Publicado: (2026)
Financial Report Chunking for Effective Retrieval Augmented Generation
por: Yepes, Antonio Jimeno, et al.
Publicado: (2024)
por: Yepes, Antonio Jimeno, et al.
Publicado: (2024)
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
por: Hirak, Vitalii, et al.
Publicado: (2026)
por: Hirak, Vitalii, et al.
Publicado: (2026)
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
por: Shajalal, Md, et al.
Publicado: (2022)
por: Shajalal, Md, et al.
Publicado: (2022)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
por: Sun, Kun, et al.
Publicado: (2024)
por: Sun, Kun, et al.
Publicado: (2024)
Disentangling Textual and Acoustic Features of Neural Speech Representations
por: Mohebbi, Hosein, et al.
Publicado: (2024)
por: Mohebbi, Hosein, et al.
Publicado: (2024)
SCORE: A Semantic Evaluation Framework for Generative Document Parsing
por: Li, Renyu, et al.
Publicado: (2025)
por: Li, Renyu, et al.
Publicado: (2025)
Assessing the Role of Data Quality in Training Bilingual Language Models
por: Seto, Skyler, et al.
Publicado: (2025)
por: Seto, Skyler, et al.
Publicado: (2025)
CLARITY: A Framework and Benchmark for Conversational Language Ambiguity and Unanswerability in Interactive NL2SQL Systems
por: Sarwar, Tabinda, et al.
Publicado: (2026)
por: Sarwar, Tabinda, et al.
Publicado: (2026)
Cross-lingual Transfer or Machine Translation? On Data Augmentation for Monolingual Semantic Textual Similarity
por: Hoshino, Sho, et al.
Publicado: (2024)
por: Hoshino, Sho, et al.
Publicado: (2024)
Exploration of Attention Mechanism-Enhanced Deep Learning Models in the Mining of Medical Textual Data
por: Xiao, Lingxi, et al.
Publicado: (2024)
por: Xiao, Lingxi, et al.
Publicado: (2024)
Eliciting Textual Descriptions from Representations of Continuous Prompts
por: Ramati, Dana, et al.
Publicado: (2024)
por: Ramati, Dana, et al.
Publicado: (2024)
VI-OOD: A Unified Representation Learning Framework for Textual Out-of-distribution Detection
por: Zhan, Li-Ming, et al.
Publicado: (2024)
por: Zhan, Li-Ming, et al.
Publicado: (2024)
Evolutionary thoughts: integration of large language models and evolutionary algorithms
por: Yepes, Antonio Jimeno, et al.
Publicado: (2025)
por: Yepes, Antonio Jimeno, et al.
Publicado: (2025)
Empowering Large Language Models for Textual Data Augmentation
por: Li, Yichuan, et al.
Publicado: (2024)
por: Li, Yichuan, et al.
Publicado: (2024)
Can Automatic Metrics Assess High-Quality Translations?
por: Agrawal, Sweta, et al.
Publicado: (2024)
por: Agrawal, Sweta, et al.
Publicado: (2024)
A Survey of Machine Learning Models and Datasets for the Multi-label Classification of Textual Hate Speech in English
por: Bäumler, Julian, et al.
Publicado: (2025)
por: Bäumler, Julian, et al.
Publicado: (2025)
Fooling the Textual Fooler via Randomizing Latent Representations
por: Hoang, Duy C., et al.
Publicado: (2023)
por: Hoang, Duy C., et al.
Publicado: (2023)
Integrating Large Language Models and Knowledge Graphs for Extraction and Validation of Textual Test Data
por: De Santis, Antonio, et al.
Publicado: (2024)
por: De Santis, Antonio, et al.
Publicado: (2024)
ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
por: Li, Ruihang, et al.
Publicado: (2024)
por: Li, Ruihang, et al.
Publicado: (2024)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
por: Sarwar, Nobin
Publicado: (2025)
por: Sarwar, Nobin
Publicado: (2025)
GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training
por: Nacar, Omer, et al.
Publicado: (2025)
por: Nacar, Omer, et al.
Publicado: (2025)
From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation
por: Zhao, Haofei, et al.
Publicado: (2024)
por: Zhao, Haofei, et al.
Publicado: (2024)
Potential and Perils of Large Language Models as Judges of Unstructured Textual Data
por: Bedemariam, Rewina, et al.
Publicado: (2025)
por: Bedemariam, Rewina, et al.
Publicado: (2025)
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
por: Hossain, Eftekhar, et al.
Publicado: (2024)
por: Hossain, Eftekhar, et al.
Publicado: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
por: Hua, Jiacheng, et al.
Publicado: (2026)
por: Hua, Jiacheng, et al.
Publicado: (2026)
Learning High-Quality and General-Purpose Phrase Representations
por: Chen, Lihu, et al.
Publicado: (2024)
por: Chen, Lihu, et al.
Publicado: (2024)
Enhancing Traffic Prediction with Textual Data Using Large Language Models
por: Huang, Xiannan
Publicado: (2024)
por: Huang, Xiannan
Publicado: (2024)
An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data
por: Yoffe, Gideon, et al.
Publicado: (2025)
por: Yoffe, Gideon, et al.
Publicado: (2025)
CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models
por: Zhang, Gaifan, et al.
Publicado: (2025)
por: Zhang, Gaifan, et al.
Publicado: (2025)
Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
por: Lin, Tzu-Ling, et al.
Publicado: (2025)
por: Lin, Tzu-Ling, et al.
Publicado: (2025)
Tell Me What's Next: Textual Foresight for Generic UI Representations
por: Burns, Andrea, et al.
Publicado: (2024)
por: Burns, Andrea, et al.
Publicado: (2024)
New Textual Corpora for Serbian Language Modeling
por: Škorić, Mihailo, et al.
Publicado: (2024)
por: Škorić, Mihailo, et al.
Publicado: (2024)
CraftRTL: High-quality Synthetic Data Generation for Verilog Code Models with Correct-by-Construction Non-Textual Representations and Targeted Code Repair
por: Liu, Mingjie, et al.
Publicado: (2024)
por: Liu, Mingjie, et al.
Publicado: (2024)
Mashee at SemEval-2024 Task 8: The Impact of Samples Quality on the Performance of In-Context Learning for Machine Text Classification
por: Rasheed, Areeg Fahad, et al.
Publicado: (2024)
por: Rasheed, Areeg Fahad, et al.
Publicado: (2024)
Ordered Semantically Diverse Sampling for Textual Data
por: Tiwari, Ashish, et al.
Publicado: (2025)
por: Tiwari, Ashish, et al.
Publicado: (2025)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
por: Agrawal, Sweta, et al.
Publicado: (2024)
por: Agrawal, Sweta, et al.
Publicado: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
Ejemplares similares
-
Deep Contrastive Unlearning for Language Models
por: He, Estrid, et al.
Publicado: (2025) -
RADS: Reinforcement Learning-Based Sample Selection Improves Transfer Learning in Low-resource and Imbalanced Clinical Settings
por: Han, Wei, et al.
Publicado: (2026) -
Financial Report Chunking for Effective Retrieval Augmented Generation
por: Yepes, Antonio Jimeno, et al.
Publicado: (2024) -
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
por: Hirak, Vitalii, et al.
Publicado: (2026) -
Textual Entailment Recognition with Semantic Features from Empirical Text Representation
por: Shajalal, Md, et al.
Publicado: (2022)