ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
Fuente:
arXiv
Salvato in:
| Autori principali: | M., Yashwanth, Singh, Vaibhav, Maheshwari, Ayush, Krishna, Amrith, Ramakrishnan, Ganesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
di: Singh, Vaibhav, et al.
Pubblicazione: (2024)
di: Singh, Vaibhav, et al.
Pubblicazione: (2024)
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
di: Maheshwari, Ayush, et al.
Pubblicazione: (2023)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2023)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022)
LexGen: Domain-aware Multilingual Lexicon Generation
di: Maheshwari, Ayush, et al.
Pubblicazione: (2024)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2024)
RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical Rules
di: Li, Miaomiao, et al.
Pubblicazione: (2024)
di: Li, Miaomiao, et al.
Pubblicazione: (2024)
Understanding the Influence of Synthetic Data for Text Embedders
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
di: Jagadeeshan, Manoj Balaji, et al.
Pubblicazione: (2026)
di: Jagadeeshan, Manoj Balaji, et al.
Pubblicazione: (2026)
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
di: Maheshwari, Ayush, et al.
Pubblicazione: (2025)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2025)
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
di: Maheshwari, Ayush, et al.
Pubblicazione: (2025)
di: Maheshwari, Ayush, et al.
Pubblicazione: (2025)
A2TTS: TTS for Low Resource Indian Languages
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
FAIR: Filtering of Automatically Induced Rules
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2024)
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2024)
Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2024)
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2024)
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
di: Ray, Pretam, et al.
Pubblicazione: (2024)
di: Ray, Pretam, et al.
Pubblicazione: (2024)
Efficacy of Synthetic Data as a Benchmark
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2024)
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2024)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
di: Gupta, Raavi, et al.
Pubblicazione: (2025)
di: Gupta, Raavi, et al.
Pubblicazione: (2025)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
The Synthetic Imputation Approach: Generating Optimal Synthetic Texts For Underrepresented Categories In Supervised Classification Tasks
di: Timoneda, Joan C.
Pubblicazione: (2025)
di: Timoneda, Joan C.
Pubblicazione: (2025)
PRIMO: Progressive Induction for Multi-hop Open Rule Generation
di: Liu, Jianyu, et al.
Pubblicazione: (2024)
di: Liu, Jianyu, et al.
Pubblicazione: (2024)
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
di: J, Karthika N, et al.
Pubblicazione: (2024)
di: J, Karthika N, et al.
Pubblicazione: (2024)
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
di: Ramesh, Krithika, et al.
Pubblicazione: (2025)
di: Ramesh, Krithika, et al.
Pubblicazione: (2025)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
di: Yang, Yue, et al.
Pubblicazione: (2025)
di: Yang, Yue, et al.
Pubblicazione: (2025)
WHERE and WHICH: Iterative Debate for Biomedical Synthetic Data Augmentation
di: Zhao, Zhengyi, et al.
Pubblicazione: (2025)
di: Zhao, Zhengyi, et al.
Pubblicazione: (2025)
Enhancing Low-Resource NMT with a Multilingual Encoder and Knowledge Distillation: A Case Study
di: Roy, Aniruddha, et al.
Pubblicazione: (2024)
di: Roy, Aniruddha, et al.
Pubblicazione: (2024)
AutoGeTS: Knowledge-based Automated Generation of Text Synthetics for Improving Text Classification
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2024)
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2024)
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
di: Pandey, Isha, et al.
Pubblicazione: (2026)
di: Pandey, Isha, et al.
Pubblicazione: (2026)
Innovations in Neural Data-to-text Generation: A Survey
di: Sharma, Mandar, et al.
Pubblicazione: (2022)
di: Sharma, Mandar, et al.
Pubblicazione: (2022)
Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
di: Sarkar, Sujoy, et al.
Pubblicazione: (2025)
di: Sarkar, Sujoy, et al.
Pubblicazione: (2025)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2024)
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2024)
Synthetic Data Generation Using Large Language Models: Advances in Text and Code
di: Nadas, Mihai, et al.
Pubblicazione: (2025)
di: Nadas, Mihai, et al.
Pubblicazione: (2025)
Harnessing LLMs for API Interactions: A Framework for Classification and Synthetic Data Generation
di: Tao, Chunliang, et al.
Pubblicazione: (2024)
di: Tao, Chunliang, et al.
Pubblicazione: (2024)
Token Prediction as Implicit Classification to Identify LLM-Generated Text
di: Chen, Yutian, et al.
Pubblicazione: (2023)
di: Chen, Yutian, et al.
Pubblicazione: (2023)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
di: Guliani, Keerat, et al.
Pubblicazione: (2026)
di: Guliani, Keerat, et al.
Pubblicazione: (2026)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
di: Devanathan, Rishikesh, et al.
Pubblicazione: (2025)
di: Devanathan, Rishikesh, et al.
Pubblicazione: (2025)
An Empirical Study of Validating Synthetic Data for Formula Generation
di: Singh, Usneek, et al.
Pubblicazione: (2024)
di: Singh, Usneek, et al.
Pubblicazione: (2024)
LLM-Guided Synthetic Augmentation (LGSA) for Mitigating Bias in AI Systems
di: Karri, Sai Suhruth Reddy, et al.
Pubblicazione: (2025)
di: Karri, Sai Suhruth Reddy, et al.
Pubblicazione: (2025)
Private Synthetic Text Generation with Diffusion Models
di: Ochs, Sebastian, et al.
Pubblicazione: (2024)
di: Ochs, Sebastian, et al.
Pubblicazione: (2024)
Synthetic News Generation for Fake News Classification
di: Sittar, Abdul, et al.
Pubblicazione: (2025)
di: Sittar, Abdul, et al.
Pubblicazione: (2025)
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2026)
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2026)
Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck
di: Ludan, Josh Magnus, et al.
Pubblicazione: (2023)
di: Ludan, Josh Magnus, et al.
Pubblicazione: (2023)
Documenti analoghi
-
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
di: Singh, Vaibhav, et al.
Pubblicazione: (2024) -
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
di: Maheshwari, Ayush, et al.
Pubblicazione: (2023) -
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
di: Maheshwari, Ayush, et al.
Pubblicazione: (2022) -
LexGen: Domain-aware Multilingual Lexicon Generation
di: Maheshwari, Ayush, et al.
Pubblicazione: (2024) -
RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical Rules
di: Li, Miaomiao, et al.
Pubblicazione: (2024)