Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Dahan, Noam, Kidron, Omer, Stanovsky, Gabriel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The State and Fate of Summarization Datasets: A Survey
by: Dahan, Noam, et al.
Published: (2024)
by: Dahan, Noam, et al.
Published: (2024)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
by: Lior, Gili, et al.
Published: (2024)
by: Lior, Gili, et al.
Published: (2024)
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
by: Goldstein, Ariel, et al.
Published: (2024)
by: Goldstein, Ariel, et al.
Published: (2024)
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
by: Lior, Gili, et al.
Published: (2023)
by: Lior, Gili, et al.
Published: (2023)
Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback
by: Luo, Chu Fei, et al.
Published: (2025)
by: Luo, Chu Fei, et al.
Published: (2025)
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
by: Gabay, Adi, et al.
Published: (2026)
by: Gabay, Adi, et al.
Published: (2026)
In-Context Learning on a Budget: A Case Study in Token Classification
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
Scaling Up Summarization: Leveraging Large Language Models for Long Text Extractive Summarization
by: Hemamou, Léo, et al.
Published: (2024)
by: Hemamou, Léo, et al.
Published: (2024)
A Cookbook for Community-driven Data Collection of Impaired Speech in LowResource Languages
by: Salihs, Sumaya Ahmed, et al.
Published: (2025)
by: Salihs, Sumaya Ahmed, et al.
Published: (2025)
Anticipatory Evaluation of Language Models
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
Low-Resource Cross-Lingual Summarization through Few-Shot Learning with Large Language Models
by: Park, Gyutae, et al.
Published: (2024)
by: Park, Gyutae, et al.
Published: (2024)
Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
by: Bandarupalli, Srihari, et al.
Published: (2025)
by: Bandarupalli, Srihari, et al.
Published: (2025)
Mixture of Experts for Low-Resource LLMs
by: Joseph, Ori Bar, et al.
Published: (2026)
by: Joseph, Ori Bar, et al.
Published: (2026)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
by: Iluz, Bar, et al.
Published: (2024)
by: Iluz, Bar, et al.
Published: (2024)
Efficient Extractive Summarization with MAMBA-Transformer Hybrids for Low-Resource Scenarios
by: Khayi, Nisrine Ait
Published: (2026)
by: Khayi, Nisrine Ait
Published: (2026)
Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games
by: Eckhaus, Niv, et al.
Published: (2025)
by: Eckhaus, Niv, et al.
Published: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
by: Itzhak, Itay, et al.
Published: (2025)
by: Itzhak, Itay, et al.
Published: (2025)
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
by: Levy, Shahar, et al.
Published: (2026)
by: Levy, Shahar, et al.
Published: (2026)
From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule Discovery
by: Chen, Yuhan, et al.
Published: (2023)
by: Chen, Yuhan, et al.
Published: (2023)
State Space Models for Extractive Summarization in Low Resource Scenarios
by: Khayi, Nisrine Ait
Published: (2025)
by: Khayi, Nisrine Ait
Published: (2025)
Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language
by: Zhukova, Anastasia, et al.
Published: (2024)
by: Zhukova, Anastasia, et al.
Published: (2024)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
by: Berger, Uri, et al.
Published: (2025)
by: Berger, Uri, et al.
Published: (2025)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
Looking Beyond The Top-1: Transformers Determine Top Tokens In Order
by: Lioubashevski, Daria, et al.
Published: (2024)
by: Lioubashevski, Daria, et al.
Published: (2024)
Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese Newspaper
by: Ishihara, Shotaro, et al.
Published: (2024)
by: Ishihara, Shotaro, et al.
Published: (2024)
Comparing Approaches to Automatic Summarization in Less-Resourced Languages
by: Palen-Michel, Chester, et al.
Published: (2025)
by: Palen-Michel, Chester, et al.
Published: (2025)
Reliability Gated Multi-Teacher Distillation for Low Resource Abstractive Summarization
by: Sumit, Dipto, et al.
Published: (2026)
by: Sumit, Dipto, et al.
Published: (2026)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
by: Levy, Shahar, et al.
Published: (2025)
by: Levy, Shahar, et al.
Published: (2025)
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
by: Prome, Ruhina Tabasshum, et al.
Published: (2025)
YouTube Comments Decoded: Leveraging LLMs for Low Resource Language Classification
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
Efficient Language Modeling for Low-Resource Settings with Hybrid RNN-Transformer Architectures
by: Lindenmaier, Gabriel, et al.
Published: (2025)
by: Lindenmaier, Gabriel, et al.
Published: (2025)
Bridging Gaps in Hate Speech Detection: Meta-Collections and Benchmarks for Low-Resource Iberian Languages
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Cascading Adaptors to Leverage English Data to Improve Performance of Question Answering for Low-Resource Languages
by: Pandya, Hariom A., et al.
Published: (2021)
by: Pandya, Hariom A., et al.
Published: (2021)
Leveraging Large Language Models for Comparative Literature Summarization with Reflective Incremental Mechanisms
by: Garcia, Fernando Gabriela, et al.
Published: (2024)
by: Garcia, Fernando Gabriela, et al.
Published: (2024)
Large Language Models' Detection of Political Orientation in Newspapers
by: Buscemi, Alessio, et al.
Published: (2024)
by: Buscemi, Alessio, et al.
Published: (2024)
Similar Items
-
The State and Fate of Summarization Datasets: A Survey
by: Dahan, Noam, et al.
Published: (2024) -
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025) -
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
by: Lior, Gili, et al.
Published: (2024) -
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
by: Goldstein, Ariel, et al.
Published: (2024) -
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
by: Lior, Gili, et al.
Published: (2023)