MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Mitra, Purbesh, Ulukus, Sennur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
by: Mitra, Purbesh, et al.
Published: (2025)
by: Mitra, Purbesh, et al.
Published: (2025)
Scale-Robust Timely Asynchronous Decentralized Learning
by: Mitra, Purbesh, et al.
Published: (2024)
by: Mitra, Purbesh, et al.
Published: (2024)
Distributed Mixture-of-Agents for Edge Inference with Large Language Models
by: Mitra, Purbesh, et al.
Published: (2024)
by: Mitra, Purbesh, et al.
Published: (2024)
Multi-Modal Semantic Communication
by: Mortaheb, Matin, et al.
Published: (2025)
by: Mortaheb, Matin, et al.
Published: (2025)
Fine-tuning Smaller Language Models for Question Answering over Financial Documents
by: Phogat, Karmvir Singh, et al.
Published: (2024)
by: Phogat, Karmvir Singh, et al.
Published: (2024)
Joint Age-State Belief is All You Need: Minimizing AoII via Pull-Based Remote Estimation
by: Cosandal, Ismail, et al.
Published: (2024)
by: Cosandal, Ismail, et al.
Published: (2024)
Optimum Monitoring and Job Assignment with Multiple Markov Machines
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
Age of Estimates: When to Submit Jobs to a Markov Machine to Maximize Revenue
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
JCAS-MARL: Joint Communication and Sensing UAV Networks via Resource-Constrained Multi-Agent Reinforcement Learning
by: Guven, Islam, et al.
Published: (2026)
by: Guven, Islam, et al.
Published: (2026)
Utilizing the Perceived Age to Maximize Freshness in Query-Based Update Systems
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
Beyond Martingale Estimators: Structured Estimators for Maximizing Information Freshness in Query-Based Update Systems
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
When to Preempt in a Status Update System?
by: Banerjee, Subhankar, et al.
Published: (2024)
by: Banerjee, Subhankar, et al.
Published: (2024)
Equidistant-Sample or Wait-and-Sample to Minimize Age Under Sampling Constraint?
by: Banerjee, Subhankar, et al.
Published: (2025)
by: Banerjee, Subhankar, et al.
Published: (2025)
Tracking and Assigning Jobs to a Markov Machine
by: Banerjee, Subhankar, et al.
Published: (2025)
by: Banerjee, Subhankar, et al.
Published: (2025)
Preemptive Scheduling for Age of Job Minimization in Task-Specific Machine Networks
by: Banerjee, Subhankar, et al.
Published: (2026)
by: Banerjee, Subhankar, et al.
Published: (2026)
FALCON: Autonomous Cyber Threat Intelligence Mining with LLMs for IDS Rule Generation
by: Mitra, Shaswata, et al.
Published: (2025)
by: Mitra, Shaswata, et al.
Published: (2025)
Structure-Enhanced Deep Reinforcement Learning for Optimal Transmission Scheduling
by: Chen, Jiazheng, et al.
Published: (2022)
by: Chen, Jiazheng, et al.
Published: (2022)
Capacity-Constrained Continual Learning
by: Wen, Zheng, et al.
Published: (2025)
by: Wen, Zheng, et al.
Published: (2025)
Semantic-aware Transmission Scheduling: a Monotonicity-driven Deep Reinforcement Learning Approach
by: Chen, Jiazheng, et al.
Published: (2023)
by: Chen, Jiazheng, et al.
Published: (2023)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
by: Kharrat, Salma, et al.
Published: (2024)
by: Kharrat, Salma, et al.
Published: (2024)
Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement
by: Xie, Guanwen, et al.
Published: (2024)
by: Xie, Guanwen, et al.
Published: (2024)
Semi-Markov Decision Process Framework for Age of Incorrect Information Minimization
by: Cosandal, Ismail, et al.
Published: (2025)
by: Cosandal, Ismail, et al.
Published: (2025)
Source Coding for a Wiener Process
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
Structured Estimators: A New Perspective on Information Freshness
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
by: Liyanaarachchi, Sahan, et al.
Published: (2025)
AoII-Optimum Sampling of CTMC Information Sources Under Sampling Rate Constraints
by: Cosandal, Ismail, et al.
Published: (2024)
by: Cosandal, Ismail, et al.
Published: (2024)
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
by: Niu, Xueyan, et al.
Published: (2026)
by: Niu, Xueyan, et al.
Published: (2026)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
Privacy-Preserving Semantic Communications via Multi-Task Learning and Adversarial Perturbations
by: Sagduyu, Yalin E., et al.
Published: (2025)
by: Sagduyu, Yalin E., et al.
Published: (2025)
Learning What Matters: Adaptive Information-Theoretic Objectives for Robot Exploration
by: Yu, Youwei, et al.
Published: (2026)
by: Yu, Youwei, et al.
Published: (2026)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
by: Sitdhipol, Supawich, et al.
Published: (2025)
by: Sitdhipol, Supawich, et al.
Published: (2025)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Index-Based Scheduling for a Resource-Constrained Quantum Switch
by: Banerjee, Subhankar, et al.
Published: (2026)
by: Banerjee, Subhankar, et al.
Published: (2026)
Scheduling Policies in a Multi-Source Status Update System with Dedicated and Shared Servers
by: Liyanaarachchi, Sahan, et al.
Published: (2024)
by: Liyanaarachchi, Sahan, et al.
Published: (2024)
Preemption Revisited: Multi-Threshold Preemption Policies for AoI Minimization
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
Minimizing the Age of Two Heterogeneous Sources With Packet Drops Via Cyclic Schedulers
by: Liyanaarachchi, Sahan, et al.
Published: (2024)
by: Liyanaarachchi, Sahan, et al.
Published: (2024)
Age-Based Scheduling for a Memory-Constrained Quantum Switch
by: Mitrolaris, Stavros, et al.
Published: (2026)
by: Mitrolaris, Stavros, et al.
Published: (2026)
6G at $\frac{1}{6}g$: The Future of Cislunar Communications
by: Liyanaarachchi, Sahan, et al.
Published: (2024)
by: Liyanaarachchi, Sahan, et al.
Published: (2024)
FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast
by: Bogdanov, Igor, et al.
Published: (2026)
by: Bogdanov, Igor, et al.
Published: (2026)
Topology-Aware Exploration of Energy-Based Models Equilibrium: Toric QC-LDPC Codes and Hyperbolic MET QC-LDPC Codes
by: Usatyuk, Vasiliy, et al.
Published: (2024)
by: Usatyuk, Vasiliy, et al.
Published: (2024)
Similar Items
-
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
by: Mitra, Purbesh, et al.
Published: (2025) -
Scale-Robust Timely Asynchronous Decentralized Learning
by: Mitra, Purbesh, et al.
Published: (2024) -
Distributed Mixture-of-Agents for Edge Inference with Large Language Models
by: Mitra, Purbesh, et al.
Published: (2024) -
Multi-Modal Semantic Communication
by: Mortaheb, Matin, et al.
Published: (2025) -
Fine-tuning Smaller Language Models for Question Answering over Financial Documents
by: Phogat, Karmvir Singh, et al.
Published: (2024)