Saved in:
| Main Authors: | Zhou, Yuhang, Zhang, Mingrui, Li, Ke, Wang, Mingyi, Liu, Qiao, Wang, Qifei, Liu, Jiayi, Liu, Fei, Li, Serena, Li, Weiwei, Gao, Mingze, Kumar, Abhishek, Fan, Xiangjun, Zhao, Zhuokai, Zhang, Lizhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.20176 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Driven Reasoning for Constraint-Aware Feature Selection in Industrial Systems
by: Zhou, Yuhang, et al.
Published: (2026)
by: Zhou, Yuhang, et al.
Published: (2026)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
by: Zhou, Yuhang, et al.
Published: (2026)
by: Zhou, Yuhang, et al.
Published: (2026)
Synthetic Sandbox for Training Machine Learning Engineering Agents
by: Zhou, Yuhang, et al.
Published: (2026)
by: Zhou, Yuhang, et al.
Published: (2026)
EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimization
by: Han, Kevin, et al.
Published: (2026)
by: Han, Kevin, et al.
Published: (2026)
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
by: Zeng, Hanqing, et al.
Published: (2025)
by: Zeng, Hanqing, et al.
Published: (2025)
GEM: Empowering LLM for both Embedding Generation and Language Understanding
by: Zhang, Caojin, et al.
Published: (2025)
by: Zhang, Caojin, et al.
Published: (2025)
RecoWorld: Building Simulated Environments for Agentic Recommender Systems
by: Liu, Fei, et al.
Published: (2025)
by: Liu, Fei, et al.
Published: (2025)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Thought Communication in Multiagent Collaboration
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
Exploring System 1 and 2 communication for latent reasoning in LLMs
by: Coda-Forno, Julian, et al.
Published: (2025)
by: Coda-Forno, Julian, et al.
Published: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
Agentic Recommender System with Hierarchical Belief-State Memory
by: Shen, Xiang, et al.
Published: (2026)
by: Shen, Xiang, et al.
Published: (2026)
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
Token-Level LLM Collaboration via FusionRoute
by: Xiong, Nuoya, et al.
Published: (2026)
by: Xiong, Nuoya, et al.
Published: (2026)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Facet-Aware Multi-Head Mixture-of-Experts Model for Sequential Recommendation
by: Liu, Mingrui, et al.
Published: (2024)
by: Liu, Mingrui, et al.
Published: (2024)
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
by: Fostiropoulos, Iordanis, et al.
Published: (2026)
by: Fostiropoulos, Iordanis, et al.
Published: (2026)
Grid Evolution for Doubly Fractional Channel Estimation in OTFS Systems
by: Li, Xiangjun, et al.
Published: (2024)
by: Li, Xiangjun, et al.
Published: (2024)
Facet-Aware Multi-Head Mixture-of-Experts Model with Text-Enhanced Pre-training for Sequential Recommendation
by: Liu, Mingrui, et al.
Published: (2026)
by: Liu, Mingrui, et al.
Published: (2026)
A Joint Prediction Method of Multi-Agent to Reduce Collision Rate
by: Wang, Mingyi, et al.
Published: (2024)
by: Wang, Mingyi, et al.
Published: (2024)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
by: Feng, Jiarui, et al.
Published: (2026)
by: Feng, Jiarui, et al.
Published: (2026)
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
by: Kong, Fanqi, et al.
Published: (2026)
by: Kong, Fanqi, et al.
Published: (2026)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Understanding Inverse Reinforcement Learning under Overparameterization: Non-Asymptotic Analysis and Global Optimality
by: Zhang, Ruijia, et al.
Published: (2025)
by: Zhang, Ruijia, et al.
Published: (2025)
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
by: Wang, Mingze, et al.
Published: (2025)
by: Wang, Mingze, et al.
Published: (2025)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
by: Ropero, Fernando, et al.
Published: (2026)
by: Ropero, Fernando, et al.
Published: (2026)
UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language Models
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Destroy and Repair Using Hyper Graphs for Routing
by: Li, Ke, et al.
Published: (2025)
by: Li, Ke, et al.
Published: (2025)
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
by: Li, Sijie, et al.
Published: (2025)
by: Li, Sijie, et al.
Published: (2025)
Chain-of-Query: Unleashing the Power of LLMs in SQL-Aided Table Understanding via Multi-Agent Collaboration
by: Sui, Songyuan, et al.
Published: (2025)
by: Sui, Songyuan, et al.
Published: (2025)
Matched Filtering-Based Channel Estimation for AFDM Systems in Doubly Selective Channels
by: Li, Xiangjun, et al.
Published: (2025)
by: Li, Xiangjun, et al.
Published: (2025)
Agentic Reinforcement Learning with Implicit Step Rewards
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
An Online Review‐Driven Multiattribute Decision‐Making Model Based on Rough‐Cloud‐Integrated Three‐Way Decisions
by: Fan Jia, et al.
Published: (2026)
by: Fan Jia, et al.
Published: (2026)
Understanding Nonlinear Collaboration between Human and AI Agents: A Co-design Framework for Creative Design
by: Zhou, Jiayi, et al.
Published: (2024)
by: Zhou, Jiayi, et al.
Published: (2024)
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
by: Liu, Yuhong, et al.
Published: (2025)
by: Liu, Yuhong, et al.
Published: (2025)
Functionality Locality, Mixture & Control = Logic = Memory
by: Peng, Xiangjun
Published: (2024)
by: Peng, Xiangjun
Published: (2024)
Preliminary analysis on the interdecadal variation characteristics of typhoon over the Northwestern Pacific in the past sixty years.
by: Yu, Fan, et al.
Published: (2012)
by: Yu, Fan, et al.
Published: (2012)
Preliminary analysis on the interdecadal variation characteristics of typhoon over the Northwestern Pacific in the past sixty years.
by: Yu, Fan, et al.
Published: (2012)
by: Yu, Fan, et al.
Published: (2012)
Similar Items
-
LLM-Driven Reasoning for Constraint-Aware Feature Selection in Industrial Systems
by: Zhou, Yuhang, et al.
Published: (2026) -
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
by: Zhou, Yuhang, et al.
Published: (2026) -
Synthetic Sandbox for Training Machine Learning Engineering Agents
by: Zhou, Yuhang, et al.
Published: (2026) -
EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimization
by: Han, Kevin, et al.
Published: (2026) -
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
by: Zeng, Hanqing, et al.
Published: (2025)