Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
Fuente:
arXiv
Saved in:
| Main Author: | Patel, Rohit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
by: Grigaliūnas, Domas, et al.
Published: (2024)
by: Grigaliūnas, Domas, et al.
Published: (2024)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
Feature Selection Based on Reinforcement Learning and Hazard State Classification for Magnetic Adhesion Wall-Climbing Robots
by: Ma, Zhen, et al.
Published: (2025)
by: Ma, Zhen, et al.
Published: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
by: Yang, Yibo
Published: (2025)
by: Yang, Yibo
Published: (2025)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
by: Mathew, Aby Mammen
Published: (2026)
by: Mathew, Aby Mammen
Published: (2026)
How much do LLMs learn from negative examples?
by: Hamdan, Shadi, et al.
Published: (2025)
by: Hamdan, Shadi, et al.
Published: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
by: Lee, Wooin, et al.
Published: (2026)
by: Lee, Wooin, et al.
Published: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
by: Mutlu, Abdulvahap, et al.
Published: (2026)
by: Mutlu, Abdulvahap, et al.
Published: (2026)
Extracting Sentence Embeddings from Pretrained Transformer Models
by: Stankevičius, Lukas, et al.
Published: (2024)
by: Stankevičius, Lukas, et al.
Published: (2024)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
by: Vileikytė, Brigita, et al.
Published: (2024)
by: Vileikytė, Brigita, et al.
Published: (2024)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
by: Chen, Yihong, et al.
Published: (2022)
by: Chen, Yihong, et al.
Published: (2022)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
by: Keeman, Michael
Published: (2026)
by: Keeman, Michael
Published: (2026)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
BreakFun: Jailbreaking LLMs via Schema Exploitation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
GIM: Evaluating models via tasks that integrate multiple cognitive domains
by: Patel, Rohit, et al.
Published: (2026)
by: Patel, Rohit, et al.
Published: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
by: Ma, Yueen, et al.
Published: (2024)
by: Ma, Yueen, et al.
Published: (2024)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
by: Hari, Vishnu, et al.
Published: (2025)
by: Hari, Vishnu, et al.
Published: (2025)
Model Surgery: Modulating LLM's Behavior Via Simple Parameter Editing
by: Wang, Huanqian, et al.
Published: (2024)
by: Wang, Huanqian, et al.
Published: (2024)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
by: Kim, Heejun, et al.
Published: (2026)
by: Kim, Heejun, et al.
Published: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
by: Tikhonov, Alexey, et al.
Published: (2026)
by: Tikhonov, Alexey, et al.
Published: (2026)
Mitigating Position-Shift Failures in Text-Based Modular Arithmetic via Position Curriculum and Template Diversity
by: Yudin, Nikolay
Published: (2026)
by: Yudin, Nikolay
Published: (2026)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
by: Sela, Omer
Published: (2026)
by: Sela, Omer
Published: (2026)
HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
by: Hill, Brennen
Published: (2025)
by: Hill, Brennen
Published: (2025)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
by: Cui, Hejie, et al.
Published: (2024)
by: Cui, Hejie, et al.
Published: (2024)
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
by: Mitchell, Rupert, et al.
Published: (2025)
by: Mitchell, Rupert, et al.
Published: (2025)
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
by: Lin, Shuhang, et al.
Published: (2026)
by: Lin, Shuhang, et al.
Published: (2026)
Improved ICNN-LSTM Model Classification Based on Attitude Sensor Data for Hazardous State Assessment of Magnetic Adhesion Climbing Wall Robots
by: Ma, Zhen, et al.
Published: (2024)
by: Ma, Zhen, et al.
Published: (2024)
Similar Items
-
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026) -
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026) -
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025) -
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
by: Imanov, Olaf Yunus Laitinen
Published: (2026) -
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
by: Alpay, Faruk, et al.
Published: (2026)