Saved in:
| Main Authors: | Yang, Zhiying, Liu, Fang, Zhang, Wei, Lou, Xin, Low, Malcolm Yoke Hean, Gan, Boon Ping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.06351 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing
by: Zhang, Hanwen, et al.
Published: (2026)
by: Zhang, Hanwen, et al.
Published: (2026)
Learning-enabled Flexible Job-shop Scheduling for Scalable Smart Manufacturing
by: Moon, Sihoon, et al.
Published: (2024)
by: Moon, Sihoon, et al.
Published: (2024)
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
Smart Audit System Empowered by LLM
by: Yao, Xu, et al.
Published: (2024)
by: Yao, Xu, et al.
Published: (2024)
Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents
by: Ni, Hongqiu, et al.
Published: (2025)
by: Ni, Hongqiu, et al.
Published: (2025)
HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
by: Shen, Maohao, et al.
Published: (2025)
by: Shen, Maohao, et al.
Published: (2025)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
by: Park, Jungsoo, et al.
Published: (2026)
by: Park, Jungsoo, et al.
Published: (2026)
LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
by: Pihulski, Dzmitry, et al.
Published: (2025)
by: Pihulski, Dzmitry, et al.
Published: (2025)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
by: Yang, Senqiao, et al.
Published: (2025)
by: Yang, Senqiao, et al.
Published: (2025)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
by: Seo, Gyuhyeon, et al.
Published: (2025)
by: Seo, Gyuhyeon, et al.
Published: (2025)
MockLLM: A Multi-Agent Behavior Collaboration Framework for Online Job Seeking and Recruiting
by: Sun, Hongda, et al.
Published: (2024)
by: Sun, Hongda, et al.
Published: (2024)
Path-LLM: A Shortest-Path-based LLM Learning for Unified Graph Representation
by: Shang, Wenbo, et al.
Published: (2024)
by: Shang, Wenbo, et al.
Published: (2024)
PJB: A Reasoning-Aware Benchmark for Person-Job Retrieval
by: Wang, Guangzhi, et al.
Published: (2026)
by: Wang, Guangzhi, et al.
Published: (2026)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning
by: Ma, Weiyu, et al.
Published: (2026)
by: Ma, Weiyu, et al.
Published: (2026)
Optimizing the Training Schedule of Multilingual NMT using Reinforcement Learning
by: Allemann, Alexis, et al.
Published: (2024)
by: Allemann, Alexis, et al.
Published: (2024)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
by: Zhu, Zhiying, et al.
Published: (2024)
by: Zhu, Zhiying, et al.
Published: (2024)
HyperGraphPro: Progress-Aware Reinforcement Learning for Structure-Guided Hypergraph RAG
by: Park, Jinyoung, et al.
Published: (2026)
by: Park, Jinyoung, et al.
Published: (2026)
D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training
by: Xu, Yuanjian, et al.
Published: (2026)
by: Xu, Yuanjian, et al.
Published: (2026)
Smart-Hiring: An Explainable end-to-end Pipeline for CV Information Extraction and Job Matching
by: Khelkhal, Kenza, et al.
Published: (2025)
by: Khelkhal, Kenza, et al.
Published: (2025)
FlowCompile: An Optimizing Compiler for Structured LLM Workflows
by: Li, Junyan, et al.
Published: (2026)
by: Li, Junyan, et al.
Published: (2026)
Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
by: Sun, Yiliu, et al.
Published: (2025)
by: Sun, Yiliu, et al.
Published: (2025)
AUTOCIRCUIT-RL: Reinforcement Learning-Driven LLM for Automated Circuit Topology Generation
by: Vijayaraghavan, Prashanth, et al.
Published: (2025)
by: Vijayaraghavan, Prashanth, et al.
Published: (2025)
Learning Job Title Representation from Job Description Aggregation Network
by: Laosaengpha, Napat, et al.
Published: (2024)
by: Laosaengpha, Napat, et al.
Published: (2024)
Fairness-Aware Job Scheduling for Multi-Job Federated Learning
by: Shi, Yuxin, et al.
Published: (2024)
by: Shi, Yuxin, et al.
Published: (2024)
From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
A Survey on Data Selection for LLM Instruction Tuning
by: Zhang, Bolin, et al.
Published: (2024)
by: Zhang, Bolin, et al.
Published: (2024)
Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness
by: An, Hao, et al.
Published: (2026)
by: An, Hao, et al.
Published: (2026)
SynGraph: A Dynamic Graph-LLM Synthesis Framework for Sparse Streaming User Sentiment Modeling
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Improving Interagency Collaboration, Innovation and Learning in Criminal Justice Systems Supporting Offender Rehabilitation
by: Sarah Hean
by: Sarah Hean
ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQL
by: Dai, Yaxun, et al.
Published: (2025)
by: Dai, Yaxun, et al.
Published: (2025)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards
by: Zhang, Jiajie, et al.
Published: (2026)
by: Zhang, Jiajie, et al.
Published: (2026)
Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models
by: Liu, Runxuan, et al.
Published: (2026)
by: Liu, Runxuan, et al.
Published: (2026)
ArbGraph: Conflict-Aware Evidence Arbitration for Reliable Long-Form Retrieval-Augmented Generation
by: Niu, Qingying, et al.
Published: (2026)
by: Niu, Qingying, et al.
Published: (2026)
STRuCT-LLM: Unifying Tabular and Graph Reasoning with Reinforcement Learning for Semantic Parsing
by: Stoisser, Josefa Lia, et al.
Published: (2025)
by: Stoisser, Josefa Lia, et al.
Published: (2025)
Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
Learning Latency-Aware Orchestration for Parallel Multi-Agent Systems
by: Shi, Xi, et al.
Published: (2026)
by: Shi, Xi, et al.
Published: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Similar Items
-
An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing
by: Zhang, Hanwen, et al.
Published: (2026) -
Learning-enabled Flexible Job-shop Scheduling for Scalable Smart Manufacturing
by: Moon, Sihoon, et al.
Published: (2024) -
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
by: Wan, Zhongwei, et al.
Published: (2025) -
Smart Audit System Empowered by LLM
by: Yao, Xu, et al.
Published: (2024) -
Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents
by: Ni, Hongqiu, et al.
Published: (2025)