Saved in:
| Main Authors: | Cao, Renjie, Hu, Miaoyan, Wei, Jiahan, Ihnaini, Baha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.09612 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022)
by: Trager, Jackson, et al.
Published: (2022)
On the Convergence of Moral Self-Correction in Large Language Models
by: Liu, Guangliang, et al.
Published: (2025)
by: Liu, Guangliang, et al.
Published: (2025)
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
by: Feng, Zhaopeng, et al.
Published: (2025)
by: Feng, Zhaopeng, et al.
Published: (2025)
MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning
by: Feng, Zhaopeng, et al.
Published: (2025)
by: Feng, Zhaopeng, et al.
Published: (2025)
EuroSpeech: A Multilingual Speech Corpus
by: Pfisterer, Samuel, et al.
Published: (2025)
by: Pfisterer, Samuel, et al.
Published: (2025)
Determinants of Training Corpus Size for Clinical Text Classification
by: Chaturvedi, Jaya, et al.
Published: (2026)
by: Chaturvedi, Jaya, et al.
Published: (2026)
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models
by: Xue, Yida, et al.
Published: (2026)
by: Xue, Yida, et al.
Published: (2026)
MONOVAB : An Annotated Corpus for Bangla Multi-label Emotion Detection
by: Banshal, Sumit Kumar, et al.
Published: (2023)
by: Banshal, Sumit Kumar, et al.
Published: (2023)
GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data
by: Qi, Cong, et al.
Published: (2025)
by: Qi, Cong, et al.
Published: (2025)
ANUBHUTI: A Comprehensive Corpus For Sentiment Analysis In Bangla Regional Languages
by: Kundu, Swastika, et al.
Published: (2025)
by: Kundu, Swastika, et al.
Published: (2025)
NSINA: A News Corpus for Sinhala
by: Hettiarachchi, Hansi, et al.
Published: (2024)
by: Hettiarachchi, Hansi, et al.
Published: (2024)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
BanglaSTEM: A Parallel Corpus for Technical Domain Bangla-English Translation
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
by: Thorat, Shantanu, et al.
Published: (2025)
by: Thorat, Shantanu, et al.
Published: (2025)
MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models
by: Jadhav, Suramya, et al.
Published: (2025)
by: Jadhav, Suramya, et al.
Published: (2025)
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
by: Lu, Qiuyang, et al.
Published: (2025)
by: Lu, Qiuyang, et al.
Published: (2025)
Measuring Moral Inconsistencies in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages
by: Mutisya, Hillary, et al.
Published: (2026)
by: Mutisya, Hillary, et al.
Published: (2026)
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
by: Joshi, Raviraj, et al.
Published: (2024)
by: Joshi, Raviraj, et al.
Published: (2024)
Investigating Affect Mining Techniques for Annotation Sample Selection in the Creation of Finnish Affective Speech Corpus
by: Lahtinen, Kalle, et al.
Published: (2025)
by: Lahtinen, Kalle, et al.
Published: (2025)
Cross-lingual Named Entity Corpus for Slavic Languages
by: Piskorski, Jakub, et al.
Published: (2024)
by: Piskorski, Jakub, et al.
Published: (2024)
Do Language Models Understand Morality? Towards a Robust Detection of Moral Content
by: Bulla, Luana, et al.
Published: (2024)
by: Bulla, Luana, et al.
Published: (2024)
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
by: Teixeira, Francisco, et al.
Published: (2026)
by: Teixeira, Francisco, et al.
Published: (2026)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
BeliN: A Novel Corpus for Bengali Religious News Headline Generation using Contextual Feature Fusion
by: Osama, Md, et al.
Published: (2025)
by: Osama, Md, et al.
Published: (2025)
AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
by: Çano, Erion, et al.
Published: (2024)
by: Çano, Erion, et al.
Published: (2024)
Representation Learning with Conditional Information Flow Maximization
by: Hu, Dou, et al.
Published: (2024)
by: Hu, Dou, et al.
Published: (2024)
A Hybrid Protocol for Large-Scale Semantic Dataset Generation in Low-Resource Languages: The Turkish Semantic Relations Corpus
by: Tosun, Ebubekir, et al.
Published: (2026)
by: Tosun, Ebubekir, et al.
Published: (2026)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
UQA: Corpus for Urdu Question Answering
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models
by: Lu, Cong, et al.
Published: (2024)
by: Lu, Cong, et al.
Published: (2024)
Automated Capability Discovery via Foundation Model Self-Exploration
by: Lu, Cong, et al.
Published: (2025)
by: Lu, Cong, et al.
Published: (2025)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
Retrieve, Then Classify: Corpus-Grounded Automation of Clinical Value Set Authoring
by: Mukherjee, Sumit, et al.
Published: (2026)
by: Mukherjee, Sumit, et al.
Published: (2026)
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
by: Morlat, Geoffroy, et al.
Published: (2025)
by: Morlat, Geoffroy, et al.
Published: (2025)
Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT
by: Nguyen, Duy Anh
Published: (2026)
by: Nguyen, Duy Anh
Published: (2026)
Structured Probabilistic Coding
by: Hu, Dou, et al.
Published: (2023)
by: Hu, Dou, et al.
Published: (2023)
Meta4XNLI: A Crosslingual Parallel Corpus for Metaphor Detection and Interpretation
by: Sanchez-Bayona, Elisa, et al.
Published: (2024)
by: Sanchez-Bayona, Elisa, et al.
Published: (2024)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
by: Kesgin, H. Toprak, et al.
Published: (2024)
by: Kesgin, H. Toprak, et al.
Published: (2024)
Similar Items
-
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022) -
On the Convergence of Moral Self-Correction in Large Language Models
by: Liu, Guangliang, et al.
Published: (2025) -
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
by: Feng, Zhaopeng, et al.
Published: (2025) -
MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning
by: Feng, Zhaopeng, et al.
Published: (2025) -
EuroSpeech: A Multilingual Speech Corpus
by: Pfisterer, Samuel, et al.
Published: (2025)