LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Lai, Junyu, Zhang, Jiakun, Xu, Shuo, Chen, Taolue, Wang, Zihang, Yang, Yao, Zhang, Jiarui, Cao, Chun, Xu, Jingwei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines
di: Lai, Junyu, et al.
Pubblicazione: (2024)
di: Lai, Junyu, et al.
Pubblicazione: (2024)
MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models
di: Xu, Jingwei, et al.
Pubblicazione: (2024)
di: Xu, Jingwei, et al.
Pubblicazione: (2024)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
di: Huang, Yunpeng, et al.
Pubblicazione: (2023)
di: Huang, Yunpeng, et al.
Pubblicazione: (2023)
Rethinking Supervision Granularity: Segment-Level Learning for LLM-Based Theorem Proving
di: Xu, Shuo, et al.
Pubblicazione: (2026)
di: Xu, Shuo, et al.
Pubblicazione: (2026)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
di: Zhang, Yizhuo, et al.
Pubblicazione: (2025)
di: Zhang, Yizhuo, et al.
Pubblicazione: (2025)
Synthia: Scalable Grounded Persona Generation from Social Media Data
di: Rahimzadeh, Vahid, et al.
Pubblicazione: (2025)
di: Rahimzadeh, Vahid, et al.
Pubblicazione: (2025)
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
di: Kang, Migyeong, et al.
Pubblicazione: (2026)
di: Kang, Migyeong, et al.
Pubblicazione: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
di: Tu, Songjun, et al.
Pubblicazione: (2026)
di: Tu, Songjun, et al.
Pubblicazione: (2026)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
di: Xu, Qiyuan, et al.
Pubblicazione: (2026)
di: Xu, Qiyuan, et al.
Pubblicazione: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
Rango: Adaptive Retrieval-Augmented Proving for Automated Software Verification
di: Thompson, Kyle, et al.
Pubblicazione: (2024)
di: Thompson, Kyle, et al.
Pubblicazione: (2024)
Using LLM-Based Approaches to Enhance and Automate Topic Labeling
di: Khandelwal, Trishia
Pubblicazione: (2025)
di: Khandelwal, Trishia
Pubblicazione: (2025)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
di: DeRenzi, Brian, et al.
Pubblicazione: (2025)
di: DeRenzi, Brian, et al.
Pubblicazione: (2025)
Curating Grounded Synthetic Data with Global Perspectives for Equitable AI
di: Törnquist, Elin, et al.
Pubblicazione: (2024)
di: Törnquist, Elin, et al.
Pubblicazione: (2024)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
di: Souza, Débora, et al.
Pubblicazione: (2026)
di: Souza, Débora, et al.
Pubblicazione: (2026)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
di: Zhang, Yizhuo, et al.
Pubblicazione: (2024)
di: Zhang, Yizhuo, et al.
Pubblicazione: (2024)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
di: Liu, Zhongxin, et al.
Pubblicazione: (2025)
di: Liu, Zhongxin, et al.
Pubblicazione: (2025)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
di: Hu, Yuxuan, et al.
Pubblicazione: (2026)
di: Hu, Yuxuan, et al.
Pubblicazione: (2026)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
di: Wang, Changsheng, et al.
Pubblicazione: (2025)
di: Wang, Changsheng, et al.
Pubblicazione: (2025)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
di: Klöser, Lars, et al.
Pubblicazione: (2024)
di: Klöser, Lars, et al.
Pubblicazione: (2024)
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
di: Lajčinová, Bibiána, et al.
Pubblicazione: (2024)
di: Lajčinová, Bibiána, et al.
Pubblicazione: (2024)
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
di: Chen, Mu-Chi, et al.
Pubblicazione: (2026)
di: Chen, Mu-Chi, et al.
Pubblicazione: (2026)
HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents
di: Jiang, Yilin, et al.
Pubblicazione: (2026)
di: Jiang, Yilin, et al.
Pubblicazione: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
di: Liao, Jianxing, et al.
Pubblicazione: (2025)
di: Liao, Jianxing, et al.
Pubblicazione: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
di: Wang, Liang, et al.
Pubblicazione: (2026)
di: Wang, Liang, et al.
Pubblicazione: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
di: Stacey, Joe, et al.
Pubblicazione: (2024)
di: Stacey, Joe, et al.
Pubblicazione: (2024)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
SOCIA-$\nabla$: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
"The Data Says Otherwise"-Towards Automated Fact-checking and Communication of Data Claims
di: Fu, Yu, et al.
Pubblicazione: (2024)
di: Fu, Yu, et al.
Pubblicazione: (2024)
SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models
di: Rubinstein, Beny, et al.
Pubblicazione: (2026)
di: Rubinstein, Beny, et al.
Pubblicazione: (2026)
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
di: Fang, Bowen, et al.
Pubblicazione: (2026)
di: Fang, Bowen, et al.
Pubblicazione: (2026)
Leveraging Retrieval Augmented Generative LLMs For Automated Metadata Description Generation to Enhance Data Catalogs
di: Singh, Mayank, et al.
Pubblicazione: (2025)
di: Singh, Mayank, et al.
Pubblicazione: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
di: Drchal, Jan, et al.
Pubblicazione: (2023)
di: Drchal, Jan, et al.
Pubblicazione: (2023)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
di: Hinterleitner, Lukas, et al.
Pubblicazione: (2026)
di: Hinterleitner, Lukas, et al.
Pubblicazione: (2026)
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
di: Dong, Wenxin, et al.
Pubblicazione: (2026)
di: Dong, Wenxin, et al.
Pubblicazione: (2026)
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
di: Hahn, Udo
Pubblicazione: (2024)
di: Hahn, Udo
Pubblicazione: (2024)
Applying Cognitive Design Patterns to General LLM Agents
di: Wray, Robert E., et al.
Pubblicazione: (2025)
di: Wray, Robert E., et al.
Pubblicazione: (2025)
Documenti analoghi
-
Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines
di: Lai, Junyu, et al.
Pubblicazione: (2024) -
MeteoRA: Multiple-tasks Embedded LoRA for Large Language Models
di: Xu, Jingwei, et al.
Pubblicazione: (2024) -
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
di: Huang, Yunpeng, et al.
Pubblicazione: (2023) -
Rethinking Supervision Granularity: Segment-Level Learning for LLM-Based Theorem Proving
di: Xu, Shuo, et al.
Pubblicazione: (2026) -
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
di: Zhang, Yizhuo, et al.
Pubblicazione: (2025)