CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Guixian, Su, Zeli, Zhang, Ziyin, Liu, Jianing, Han, XU, Zhang, Ting, Dong, Yushuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages
by: Su, Zeli, et al.
Published: (2025)
by: Su, Zeli, et al.
Published: (2025)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
by: Xu, Guixian, et al.
Published: (2026)
by: Xu, Guixian, et al.
Published: (2026)
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
by: Su, Zeli, et al.
Published: (2026)
by: Su, Zeli, et al.
Published: (2026)
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation
by: Su, Zeli, et al.
Published: (2026)
by: Su, Zeli, et al.
Published: (2026)
LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
by: Babalola, Olusola, et al.
Published: (2025)
by: Babalola, Olusola, et al.
Published: (2025)
Teaching Large Language Models Number-Focused Headline Generation With Key Element Rationales
by: Qian, Zhen, et al.
Published: (2025)
by: Qian, Zhen, et al.
Published: (2025)
L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages
by: Tanksale, Nishant, et al.
Published: (2025)
by: Tanksale, Nishant, et al.
Published: (2025)
SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
by: Su, Zeli, et al.
Published: (2025)
by: Su, Zeli, et al.
Published: (2025)
MiLiC-Eval: Benchmarking Multilingual LLMs for China's Minority Languages
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Beyond Quality: Unlocking Diversity in Ad Headline Generation with Large Language Models
by: Wang, Chang, et al.
Published: (2025)
by: Wang, Chang, et al.
Published: (2025)
TeClass: A Human-Annotated Relevance-based Headline Classification and Generation Dataset for Telugu
by: Kanumolu, Gopichand, et al.
Published: (2024)
by: Kanumolu, Gopichand, et al.
Published: (2024)
A diverse Multilingual News Headlines Dataset from around the World
by: Leeb, Felix, et al.
Published: (2024)
by: Leeb, Felix, et al.
Published: (2024)
KoWit-24: A Richly Annotated Dataset of Wordplay in News Headlines
by: Baranov, Alexander, et al.
Published: (2025)
by: Baranov, Alexander, et al.
Published: (2025)
RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
by: Avram, Andrei-Marius, et al.
Published: (2024)
by: Avram, Andrei-Marius, et al.
Published: (2024)
Panoramic Interests: Stylistic-Content Aware Personalized Headline Generation
by: Lian, Junhong, et al.
Published: (2025)
by: Lian, Junhong, et al.
Published: (2025)
CAP-LLM: Context-Augmented Personalized Large Language Models for News Headline Generation
by: Wilson, Raymond, et al.
Published: (2025)
by: Wilson, Raymond, et al.
Published: (2025)
Fact-Preserved Personalized News Headline Generation
by: Yang, Zhao, et al.
Published: (2025)
by: Yang, Zhao, et al.
Published: (2025)
The MediaSpin Dataset: Post-Publication News Headline Edits Annotated for Media Bias
by: Verma, Preetika, et al.
Published: (2024)
by: Verma, Preetika, et al.
Published: (2024)
Multilingual Fine-Grained News Headline Hallucination Detection
by: Shen, Jiaming, et al.
Published: (2024)
by: Shen, Jiaming, et al.
Published: (2024)
Modeling Unified Semantic Discourse Structure for High-quality Headline Generation
by: Xu, Minghui, et al.
Published: (2024)
by: Xu, Minghui, et al.
Published: (2024)
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025)
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025)
TrendFact: A Benchmark for Explainable Hotspot Perception in Fact-Checking with Natural Language Explanation
by: Zhang, Xiaocheng, et al.
Published: (2024)
by: Zhang, Xiaocheng, et al.
Published: (2024)
COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline Generation
by: Kumar, Raghvendra, et al.
Published: (2025)
by: Kumar, Raghvendra, et al.
Published: (2025)
Tug-of-War within A Decade: Conflict Resolution in Vulnerability Analysis via Teacher-Guided Retrieval-Augmented Generations
by: Zhou, Ziyin, et al.
Published: (2026)
by: Zhou, Ziyin, et al.
Published: (2026)
MC$^2$: Towards Transparent and Culturally-Aware NLP for Minority Languages in China
by: Zhang, Chen, et al.
Published: (2023)
by: Zhang, Chen, et al.
Published: (2023)
Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
by: Liu, Kejin, et al.
Published: (2025)
by: Liu, Kejin, et al.
Published: (2025)
Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian
by: Üyük, Cem, et al.
Published: (2024)
by: Üyük, Cem, et al.
Published: (2024)
Headline-Guided Extractive Summarization for Thai News Articles
by: Kositcharoensuk, Pimpitchaya, et al.
Published: (2024)
by: Kositcharoensuk, Pimpitchaya, et al.
Published: (2024)
From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms
by: Jiang, Zhaokun, et al.
Published: (2025)
by: Jiang, Zhaokun, et al.
Published: (2025)
NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in Norwegian
by: Liu, Peng, et al.
Published: (2023)
by: Liu, Peng, et al.
Published: (2023)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
by: Rony, Md Main Uddin, et al.
Published: (2024)
by: Rony, Md Main Uddin, et al.
Published: (2024)
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
What Makes You CLIC: Detection of Croatian Clickbait Headlines
by: Anđelić, Marija, et al.
Published: (2025)
by: Anđelić, Marija, et al.
Published: (2025)
Fine-Tuning Gemma-7B for Enhanced Sentiment Analysis of Financial News Headlines
by: Mo, Kangtong, et al.
Published: (2024)
by: Mo, Kangtong, et al.
Published: (2024)
Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization
by: Wu, Weiqi, et al.
Published: (2025)
by: Wu, Weiqi, et al.
Published: (2025)
Headlines You Won't Forget: Can Pronoun Insertion Increase Memorability?
by: Meyer, Selina, et al.
Published: (2026)
by: Meyer, Selina, et al.
Published: (2026)
LLMs for Targeted Sentiment in News Headlines: Exploring the Descriptive-Prescriptive Dilemma
by: Juroš, Jana, et al.
Published: (2024)
by: Juroš, Jana, et al.
Published: (2024)
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
PediaBench: A Comprehensive Chinese Pediatric Dataset for Benchmarking Large Language Models
by: Zhang, Qian, et al.
Published: (2024)
by: Zhang, Qian, et al.
Published: (2024)
Distinguishing Translations by Human, NMT, and ChatGPT: A Linguistic and Statistical Approach
by: Jiang, Zhaokun, et al.
Published: (2023)
by: Jiang, Zhaokun, et al.
Published: (2023)
Similar Items
-
Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages
by: Su, Zeli, et al.
Published: (2025) -
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
by: Xu, Guixian, et al.
Published: (2026) -
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
by: Su, Zeli, et al.
Published: (2026) -
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation
by: Su, Zeli, et al.
Published: (2026) -
LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
by: Babalola, Olusola, et al.
Published: (2025)