Fast, Not Fancy: Rethinking G2P with Rich Data and Rule-Based Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Qharabagh, Mahta Fetrat, Dehghanian, Zahra, Rabiee, Hamid R. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
por: Fetrat, Mahta, et al.
Publicado: (2025)
por: Fetrat, Mahta, et al.
Publicado: (2025)
CineLOG: A Training Free Approach for Cinematic Long Video Generation
por: Dehghanian, Zahra, et al.
Publicado: (2025)
por: Dehghanian, Zahra, et al.
Publicado: (2025)
Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect
por: Amanzadi, Amirtaha, et al.
Publicado: (2025)
por: Amanzadi, Amirtaha, et al.
Publicado: (2025)
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
por: Qharabagh, Muhammad Fetrat, et al.
Publicado: (2024)
por: Qharabagh, Muhammad Fetrat, et al.
Publicado: (2024)
Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions
por: Dehghanian, Zahra, et al.
Publicado: (2025)
por: Dehghanian, Zahra, et al.
Publicado: (2025)
ReDiF: Reinforced Distillation for Few Step Diffusion
por: Tighkhorshid, Amirhossein, et al.
Publicado: (2025)
por: Tighkhorshid, Amirhossein, et al.
Publicado: (2025)
LensCraft: Your Professional Virtual Cinematographer
por: Dehghanian, Zahra, et al.
Publicado: (2025)
por: Dehghanian, Zahra, et al.
Publicado: (2025)
UPL: Uncertainty-aware Pseudo-labeling for Imbalance Transductive Node Classification
por: Teimuri, Mohammad T., et al.
Publicado: (2025)
por: Teimuri, Mohammad T., et al.
Publicado: (2025)
Cueless EEG imagined speech for subject identification: dataset and benchmarks
por: Derakhshesh, Ali, et al.
Publicado: (2025)
por: Derakhshesh, Ali, et al.
Publicado: (2025)
Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
por: Atf, Zahra, et al.
Publicado: (2025)
por: Atf, Zahra, et al.
Publicado: (2025)
SoftEDA: Rethinking Rule-Based Data Augmentation with Soft Labels
por: Choi, Juhwan, et al.
Publicado: (2024)
por: Choi, Juhwan, et al.
Publicado: (2024)
RealDrag: The First Dragging Benchmark with Real Target Image
por: Zafarani, Ahmad, et al.
Publicado: (2025)
por: Zafarani, Ahmad, et al.
Publicado: (2025)
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
por: Goriely, Zébulon, et al.
Publicado: (2025)
por: Goriely, Zébulon, et al.
Publicado: (2025)
SuperPos-Prompt: Enhancing Soft Prompt Tuning of Language Models with Superposition of Multi Token Embeddings
por: SadraeiJavaeri, MohammadAli, et al.
Publicado: (2024)
por: SadraeiJavaeri, MohammadAli, et al.
Publicado: (2024)
3DLAND: 3D Lesion Abdominal Anomaly Localization Dataset
por: Advand, Mehran, et al.
Publicado: (2026)
por: Advand, Mehran, et al.
Publicado: (2026)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
por: Gallone, Anna, et al.
Publicado: (2025)
por: Gallone, Anna, et al.
Publicado: (2025)
Boosting Biomedical Concept Extraction by Rule-Based Data Augmentation
por: Shao, Qiwei, et al.
Publicado: (2024)
por: Shao, Qiwei, et al.
Publicado: (2024)
Data-adaptive Safety Rules for Training Reward Models
por: Li, Xiaomin, et al.
Publicado: (2025)
por: Li, Xiaomin, et al.
Publicado: (2025)
State of Abdominal CT Datasets: A Critical Review of Bias, Clinical Relevance, and Real-world Applicability
por: Danaei, Saeide, et al.
Publicado: (2025)
por: Danaei, Saeide, et al.
Publicado: (2025)
LLM as Graph Kernel: Rethinking Message Passing on Text-Rich Graphs
por: Zhang, Ying, et al.
Publicado: (2026)
por: Zhang, Ying, et al.
Publicado: (2026)
Learning to Execute Graph Algorithms Exactly with Graph Neural Networks
por: Qharabagh, Muhammad Fetrat, et al.
Publicado: (2026)
por: Qharabagh, Muhammad Fetrat, et al.
Publicado: (2026)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
por: Warczyński, Jędrzej, et al.
Publicado: (2025)
por: Warczyński, Jędrzej, et al.
Publicado: (2025)
Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
por: Wang, Danny, et al.
Publicado: (2025)
por: Wang, Danny, et al.
Publicado: (2025)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
por: Vemula, Saketh Reddy, et al.
Publicado: (2025)
por: Vemula, Saketh Reddy, et al.
Publicado: (2025)
Chain of Logic: Rule-Based Reasoning with Large Language Models
por: Servantez, Sergio, et al.
Publicado: (2024)
por: Servantez, Sergio, et al.
Publicado: (2024)
Rethinking Data Selection for Supervised Fine-Tuning
por: Shen, Ming
Publicado: (2024)
por: Shen, Ming
Publicado: (2024)
Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
por: Bugueño, Margarita, et al.
Publicado: (2025)
por: Bugueño, Margarita, et al.
Publicado: (2025)
Rule-Based Approaches to Atomic Sentence Extraction
por: Kamana, Lineesha, et al.
Publicado: (2026)
por: Kamana, Lineesha, et al.
Publicado: (2026)
Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives
por: Liu, Yajiao, et al.
Publicado: (2025)
por: Liu, Yajiao, et al.
Publicado: (2025)
Rule-Guided Feedback: Enhancing Reasoning by Enforcing Rule Adherence in Large Language Models
por: Diallo, Aissatou, et al.
Publicado: (2025)
por: Diallo, Aissatou, et al.
Publicado: (2025)
MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules
por: Chen, Kejie, et al.
Publicado: (2024)
por: Chen, Kejie, et al.
Publicado: (2024)
DANI: Fast Diffusion Aware Network Inference with Preserving Topological Structure Property
por: Ramezani, Maryam, et al.
Publicado: (2023)
por: Ramezani, Maryam, et al.
Publicado: (2023)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
por: Saada, Thiziri Nait, et al.
Publicado: (2025)
por: Saada, Thiziri Nait, et al.
Publicado: (2025)
Rule-Based Explanations for Retrieval-Augmented LLM Systems
por: Rorseth, Joel, et al.
Publicado: (2025)
por: Rorseth, Joel, et al.
Publicado: (2025)
Rethinking Data Synthesis: A Teacher Model Training Recipe with Interpretation
por: Chen, Yifang, et al.
Publicado: (2024)
por: Chen, Yifang, et al.
Publicado: (2024)
Distributed Automatic Generation Control subject to Ramp-Rate-Limits: Anytime Feasibility and Uniform Network-Connectivity
por: Doostmohammadian, Mohammadreza, et al.
Publicado: (2025)
por: Doostmohammadian, Mohammadreza, et al.
Publicado: (2025)
Exploring Group and Symmetry Principles in Large Language Models
por: Imani, Shima, et al.
Publicado: (2024)
por: Imani, Shima, et al.
Publicado: (2024)
Pensez: Less Data, Better Reasoning -- Rethinking French LLM
por: Ha, Huy Hoang
Publicado: (2025)
por: Ha, Huy Hoang
Publicado: (2025)
Ejemplares similares
-
LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024) -
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024) -
Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
por: Fetrat, Mahta, et al.
Publicado: (2025) -
CineLOG: A Training Free Approach for Cinematic Long Video Generation
por: Dehghanian, Zahra, et al.
Publicado: (2025) -
Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect
por: Amanzadi, Amirtaha, et al.
Publicado: (2025)