ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
Fuente:
arXiv
Saved in:
| Main Authors: | M., Yashwanth, Singh, Vaibhav, Maheshwari, Ayush, Krishna, Amrith, Ramakrishnan, Ganesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
by: Maheshwari, Ayush, et al.
Published: (2023)
by: Maheshwari, Ayush, et al.
Published: (2023)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022)
by: Maheshwari, Ayush, et al.
Published: (2022)
LexGen: Domain-aware Multilingual Lexicon Generation
by: Maheshwari, Ayush, et al.
Published: (2024)
by: Maheshwari, Ayush, et al.
Published: (2024)
RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical Rules
by: Li, Miaomiao, et al.
Published: (2024)
by: Li, Miaomiao, et al.
Published: (2024)
Understanding the Influence of Synthetic Data for Text Embedders
by: Springer, Jacob Mitchell, et al.
Published: (2025)
by: Springer, Jacob Mitchell, et al.
Published: (2025)
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
A2TTS: TTS for Low Resource Indian Languages
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
FAIR: Filtering of Automatically Induced Rules
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
by: Ray, Pretam, et al.
Published: (2024)
by: Ray, Pretam, et al.
Published: (2024)
Efficacy of Synthetic Data as a Benchmark
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
by: Gupta, Raavi, et al.
Published: (2025)
by: Gupta, Raavi, et al.
Published: (2025)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
by: Setlur, Amrith, et al.
Published: (2024)
by: Setlur, Amrith, et al.
Published: (2024)
The Synthetic Imputation Approach: Generating Optimal Synthetic Texts For Underrepresented Categories In Supervised Classification Tasks
by: Timoneda, Joan C.
Published: (2025)
by: Timoneda, Joan C.
Published: (2025)
PRIMO: Progressive Induction for Multi-hop Open Rule Generation
by: Liu, Jianyu, et al.
Published: (2024)
by: Liu, Jianyu, et al.
Published: (2024)
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
by: J, Karthika N, et al.
Published: (2024)
by: J, Karthika N, et al.
Published: (2024)
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
by: Ramesh, Krithika, et al.
Published: (2025)
by: Ramesh, Krithika, et al.
Published: (2025)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
WHERE and WHICH: Iterative Debate for Biomedical Synthetic Data Augmentation
by: Zhao, Zhengyi, et al.
Published: (2025)
by: Zhao, Zhengyi, et al.
Published: (2025)
Enhancing Low-Resource NMT with a Multilingual Encoder and Knowledge Distillation: A Case Study
by: Roy, Aniruddha, et al.
Published: (2024)
by: Roy, Aniruddha, et al.
Published: (2024)
AutoGeTS: Knowledge-based Automated Generation of Text Synthetics for Improving Text Classification
by: Xue, Chenhao, et al.
Published: (2025)
by: Xue, Chenhao, et al.
Published: (2025)
StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training
by: Ponkshe, Kaustubh, et al.
Published: (2024)
by: Ponkshe, Kaustubh, et al.
Published: (2024)
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
by: Pandey, Isha, et al.
Published: (2026)
by: Pandey, Isha, et al.
Published: (2026)
Innovations in Neural Data-to-text Generation: A Survey
by: Sharma, Mandar, et al.
Published: (2022)
by: Sharma, Mandar, et al.
Published: (2022)
Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
by: Sarkar, Sujoy, et al.
Published: (2025)
by: Sarkar, Sujoy, et al.
Published: (2025)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
Synthetic Data Generation Using Large Language Models: Advances in Text and Code
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
Harnessing LLMs for API Interactions: A Framework for Classification and Synthetic Data Generation
by: Tao, Chunliang, et al.
Published: (2024)
by: Tao, Chunliang, et al.
Published: (2024)
Token Prediction as Implicit Classification to Identify LLM-Generated Text
by: Chen, Yutian, et al.
Published: (2023)
by: Chen, Yutian, et al.
Published: (2023)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
by: Guliani, Keerat, et al.
Published: (2026)
by: Guliani, Keerat, et al.
Published: (2026)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
by: Devanathan, Rishikesh, et al.
Published: (2025)
by: Devanathan, Rishikesh, et al.
Published: (2025)
An Empirical Study of Validating Synthetic Data for Formula Generation
by: Singh, Usneek, et al.
Published: (2024)
by: Singh, Usneek, et al.
Published: (2024)
LLM-Guided Synthetic Augmentation (LGSA) for Mitigating Bias in AI Systems
by: Karri, Sai Suhruth Reddy, et al.
Published: (2025)
by: Karri, Sai Suhruth Reddy, et al.
Published: (2025)
Private Synthetic Text Generation with Diffusion Models
by: Ochs, Sebastian, et al.
Published: (2024)
by: Ochs, Sebastian, et al.
Published: (2024)
Synthetic News Generation for Fake News Classification
by: Sittar, Abdul, et al.
Published: (2025)
by: Sittar, Abdul, et al.
Published: (2025)
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
by: Maheshwari, Gaurav, et al.
Published: (2026)
by: Maheshwari, Gaurav, et al.
Published: (2026)
Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck
by: Ludan, Josh Magnus, et al.
Published: (2023)
by: Ludan, Josh Magnus, et al.
Published: (2023)
Similar Items
-
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
by: Singh, Vaibhav, et al.
Published: (2024) -
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
by: Maheshwari, Ayush, et al.
Published: (2023) -
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022) -
LexGen: Domain-aware Multilingual Lexicon Generation
by: Maheshwari, Ayush, et al.
Published: (2024) -
RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical Rules
by: Li, Miaomiao, et al.
Published: (2024)