Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Lacombe, Valentin, Quesnel, Valentin, Sileo, Damien |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
by: Lacombe, Valentin, et al.
Published: (2025)
by: Lacombe, Valentin, et al.
Published: (2025)
Saturation-Driven Dataset Generation for LLM Mathematical Reasoning in the TPTP Ecosystem
by: Quesnel, Valentin, et al.
Published: (2025)
by: Quesnel, Valentin, et al.
Published: (2025)
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
by: Sileo, Damien
Published: (2025)
by: Sileo, Damien
Published: (2025)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
by: Sileo, Damien
Published: (2024)
by: Sileo, Damien
Published: (2024)
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
by: Sileo, Damien
Published: (2024)
by: Sileo, Damien
Published: (2024)
MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts
by: Lanzeray, Etienne, et al.
Published: (2026)
by: Lanzeray, Etienne, et al.
Published: (2026)
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
by: Calbucura, Nicolas, et al.
Published: (2025)
by: Calbucura, Nicolas, et al.
Published: (2025)
MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment
by: Xiao, Mengxi, et al.
Published: (2025)
by: Xiao, Mengxi, et al.
Published: (2025)
Tau-Eval: A Unified Evaluation Framework for Useful and Private Text Anonymization
by: Loiseau, Gabriel, et al.
Published: (2025)
by: Loiseau, Gabriel, et al.
Published: (2025)
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
by: Noël, Valentin
Published: (2026)
by: Noël, Valentin
Published: (2026)
Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization
by: Loiseau, Gabriel, et al.
Published: (2026)
by: Loiseau, Gabriel, et al.
Published: (2026)
Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
by: Loiseau, Gabriel, et al.
Published: (2026)
by: Loiseau, Gabriel, et al.
Published: (2026)
TAROT: Task-Oriented Authorship Obfuscation Using Policy Optimization Methods
by: Loiseau, Gabriel, et al.
Published: (2024)
by: Loiseau, Gabriel, et al.
Published: (2024)
Training-Free Spectral Fingerprints of Voice Processing in Transformers
by: Noël, Valentin
Published: (2025)
by: Noël, Valentin
Published: (2025)
On Predicting the Post-training Potential of Pre-trained LLMs
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
Order-Based Pre-training Strategies for Procedural Text Understanding
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
by: Meadows, Jordan, et al.
Published: (2023)
by: Meadows, Jordan, et al.
Published: (2023)
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
by: Feng, Xiang, et al.
Published: (2025)
by: Feng, Xiang, et al.
Published: (2025)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
by: Ceron, Tanise, et al.
Published: (2025)
by: Ceron, Tanise, et al.
Published: (2025)
Iterative Self-Training for Code Generation via Reinforced Re-Ranking
by: Sorokin, Nikita, et al.
Published: (2025)
by: Sorokin, Nikita, et al.
Published: (2025)
Investigating the 'Autoencoder Behavior' in Speech Self-Supervised Models: a focus on HuBERT's Pretraining
by: Vielzeuf, Valentin
Published: (2024)
by: Vielzeuf, Valentin
Published: (2024)
On Data Synthesis and Post-training for Visual Abstract Reasoning
by: Zhu, Ke, et al.
Published: (2025)
by: Zhu, Ke, et al.
Published: (2025)
Can Large Language Models Generalize Procedures Across Representations?
by: Lin, Fangru, et al.
Published: (2026)
by: Lin, Fangru, et al.
Published: (2026)
AIR: Post-training Data Selection for Reasoning via Attention Head Influence
by: Liu, Jinrui, et al.
Published: (2025)
by: Liu, Jinrui, et al.
Published: (2025)
Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese Newspaper
by: Ishihara, Shotaro, et al.
Published: (2024)
by: Ishihara, Shotaro, et al.
Published: (2024)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers
by: Barriere, Valentin, et al.
Published: (2024)
by: Barriere, Valentin, et al.
Published: (2024)
A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
by: Noël, Valentin
Published: (2025)
by: Noël, Valentin
Published: (2025)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
by: Finlayson, Matthew, et al.
Published: (2025)
by: Finlayson, Matthew, et al.
Published: (2025)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
by: Iyer, Vivek, et al.
Published: (2025)
by: Iyer, Vivek, et al.
Published: (2025)
PhoGPT: Generative Pre-training for Vietnamese
by: Nguyen, Dat Quoc, et al.
Published: (2023)
by: Nguyen, Dat Quoc, et al.
Published: (2023)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
by: Zhang, Charlie, et al.
Published: (2025)
by: Zhang, Charlie, et al.
Published: (2025)
Fantastic Biases (What are They) and Where to Find Them
by: Barriere, Valentin
Published: (2024)
by: Barriere, Valentin
Published: (2024)
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026)
by: Jiang, Liangze, et al.
Published: (2026)
Sustainable self-supervised learning for speech representations
by: Lugo, Luis, et al.
Published: (2024)
by: Lugo, Luis, et al.
Published: (2024)
Spectral Archaeology: The Causal Topology of Model Evolution
by: Noël, Valentin
Published: (2026)
by: Noël, Valentin
Published: (2026)
S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models
by: Lei, Fangyu, et al.
Published: (2023)
by: Lei, Fangyu, et al.
Published: (2023)
On the Impact of Calibration Data in Post-training Quantization and Pruning
by: Williams, Miles, et al.
Published: (2023)
by: Williams, Miles, et al.
Published: (2023)
Similar Items
-
Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
by: Lacombe, Valentin, et al.
Published: (2025) -
Saturation-Driven Dataset Generation for LLM Mathematical Reasoning in the TPTP Ecosystem
by: Quesnel, Valentin, et al.
Published: (2025) -
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
by: Sileo, Damien
Published: (2025) -
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
by: Sileo, Damien
Published: (2024) -
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
by: Sileo, Damien
Published: (2024)