LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Kuzman, Taja, Ljubešić, Nikola |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
by: Vintar, Špela, et al.
Published: (2025)
by: Vintar, Špela, et al.
Published: (2025)
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
by: Ljubešić, Nikola, et al.
Published: (2025)
by: Ljubešić, Nikola, et al.
Published: (2025)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings
by: Mochtak, Michal, et al.
Published: (2023)
by: Mochtak, Michal, et al.
Published: (2023)
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
by: Xia, Mingxuan, et al.
Published: (2025)
by: Xia, Mingxuan, et al.
Published: (2025)
Mići Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect
by: Ljubešić, Nikola, et al.
Published: (2026)
by: Ljubešić, Nikola, et al.
Published: (2026)
Leveraging Annotator Disagreement for Text Classification
by: Xu, Jin, et al.
Published: (2024)
by: Xu, Jin, et al.
Published: (2024)
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
by: Ljubešić, Nikola, et al.
Published: (2025)
by: Ljubešić, Nikola, et al.
Published: (2025)
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification
by: Gao, Lingyu, et al.
Published: (2026)
by: Gao, Lingyu, et al.
Published: (2026)
Topic Bias in Emotion Classification
by: Wegge, Maximilian, et al.
Published: (2023)
by: Wegge, Maximilian, et al.
Published: (2023)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
by: Li, Yinheng, et al.
Published: (2024)
by: Li, Yinheng, et al.
Published: (2024)
Pushing The Limit of LLM Capacity for Text Classification
by: Zhang, Yazhou, et al.
Published: (2024)
by: Zhang, Yazhou, et al.
Published: (2024)
Hierarchical Text Classification with LLM-Refined Taxonomies
by: Golde, Jonas, et al.
Published: (2026)
by: Golde, Jonas, et al.
Published: (2026)
Manual Verbalizer Enrichment for Few-Shot Text Classification
by: Nguyen, Quang Anh, et al.
Published: (2024)
by: Nguyen, Quang Anh, et al.
Published: (2024)
Reasoning for Hierarchical Text Classification: The Case of Patents
by: Jiang, Lekang, et al.
Published: (2025)
by: Jiang, Lekang, et al.
Published: (2025)
Affective Polarization across European Parliaments
by: Evkoski, Bojan, et al.
Published: (2025)
by: Evkoski, Bojan, et al.
Published: (2025)
Geographic Adaptation of Pretrained Language Models
by: Hofmann, Valentin, et al.
Published: (2022)
by: Hofmann, Valentin, et al.
Published: (2022)
Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
by: Kuo, Hsun-Yu, et al.
Published: (2024)
by: Kuo, Hsun-Yu, et al.
Published: (2024)
Multilingual Topic Classification in X: Dataset and Analysis
by: Antypas, Dimosthenis, et al.
Published: (2024)
by: Antypas, Dimosthenis, et al.
Published: (2024)
Assessing In-context Learning and Fine-tuning for Topic Classification of German Web Data
by: Schelb, Julian, et al.
Published: (2024)
by: Schelb, Julian, et al.
Published: (2024)
Combining Language and Topic Models for Hierarchical Text Classification
by: Toit, Jaco du, et al.
Published: (2025)
by: Toit, Jaco du, et al.
Published: (2025)
Applying LLMs to Active Learning: Towards Cost-Efficient Cross-Task Text Classification without Manually Labeled Data
by: Zhang, Yejian, et al.
Published: (2025)
by: Zhang, Yejian, et al.
Published: (2025)
Large Language Models For Text Classification: Case Study And Comprehensive Review
by: Kostina, Arina, et al.
Published: (2025)
by: Kostina, Arina, et al.
Published: (2025)
Multilingual Power and Ideology Identification in the Parliament: a Reference Dataset and Simple Baselines
by: Çöltekin, Çağrı, et al.
Published: (2024)
by: Çöltekin, Çağrı, et al.
Published: (2024)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
by: Kim, Kon Woo, et al.
Published: (2025)
by: Kim, Kon Woo, et al.
Published: (2025)
Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation
by: Rouzegar, Hamidreza, et al.
Published: (2024)
by: Rouzegar, Hamidreza, et al.
Published: (2024)
LLMEmbed: Rethinking Lightweight LLM's Genuine Function in Text Classification
by: Liu, Chun, et al.
Published: (2024)
by: Liu, Chun, et al.
Published: (2024)
Text Classification in the LLM Era -- Where do we stand?
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
Token Prediction as Implicit Classification to Identify LLM-Generated Text
by: Chen, Yutian, et al.
Published: (2023)
by: Chen, Yutian, et al.
Published: (2023)
Towards Automating Text Annotation: A Case Study on Semantic Proximity Annotation using GPT-4
by: Yadav, Sachin, et al.
Published: (2024)
by: Yadav, Sachin, et al.
Published: (2024)
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
by: Xiong, Chenfei, et al.
Published: (2025)
by: Xiong, Chenfei, et al.
Published: (2025)
Similar Items
-
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024) -
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026) -
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
by: Pungeršek, Taja Kuzman, et al.
Published: (2025) -
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
by: Vintar, Špela, et al.
Published: (2025) -
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
by: Ljubešić, Nikola, et al.
Published: (2025)