SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Chenglong, Li, Canjia, Zhu, Xingzhao, Huo, Yifu, Wang, Huiyu, Lin, Weixiong, Yang, Yun, He, Qiaozhi, Zhou, Tianhua, Chang, Xiaojia, Zhu, Jingbo, Xiao, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
by: Huo, Yifu, et al.
Published: (2026)
by: Huo, Yifu, et al.
Published: (2026)
MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning
by: Wang, Chenglong, et al.
Published: (2026)
by: Wang, Chenglong, et al.
Published: (2026)
APR: Penalizing Structural Redundancy in Large Reasoning Models via Anchor-based Process Rewards
by: Chang, Kaiyan, et al.
Published: (2026)
by: Chang, Kaiyan, et al.
Published: (2026)
LRHP: Learning Representations for Human Preferences via Preference Pairs
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
GRAM: A Generative Foundation Reward Model for Reward Generalization
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordination
by: Zhu, Ziming, et al.
Published: (2025)
by: Zhu, Ziming, et al.
Published: (2025)
When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
by: Mao, Ruixiang, et al.
Published: (2026)
by: Mao, Ruixiang, et al.
Published: (2026)
Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
by: Huo, Yifu, et al.
Published: (2025)
by: Huo, Yifu, et al.
Published: (2025)
Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
by: Mu, Yongyu, et al.
Published: (2025)
by: Mu, Yongyu, et al.
Published: (2025)
DaPT: A Dual-Path Framework for Multilingual Multi-hop Question Answering
by: Wang, Yilin, et al.
Published: (2026)
by: Wang, Yilin, et al.
Published: (2026)
Self-Consolidation for Self-Evolving Agents
by: Yu, Hongzhuo, et al.
Published: (2026)
by: Yu, Hongzhuo, et al.
Published: (2026)
Efficient Prompting Methods for Large Language Models: A Survey
by: Chang, Kaiyan, et al.
Published: (2024)
by: Chang, Kaiyan, et al.
Published: (2024)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
by: Wu, Rong, et al.
Published: (2025)
by: Wu, Rong, et al.
Published: (2025)
Prior Constraints-based Reward Model Training for Aligning Large Language Models
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Robust Wideband Channel Estimation for mmWave Massive MIMO Systems With Beam Squint
by: Ge, Li, et al.
Published: (2022)
by: Ge, Li, et al.
Published: (2022)
On Safety Risks in Experience-Driven Self-Evolving Agents
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
Cross-layer Attention Sharing for Pre-trained Large Language Models
by: Mu, Yongyu, et al.
Published: (2024)
by: Mu, Yongyu, et al.
Published: (2024)
Hybrid Alignment Training for Large Language Models
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
Foundations of Large Language Models
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization
by: He, Shan, et al.
Published: (2026)
by: He, Shan, et al.
Published: (2026)
Mem$^2$Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
by: Li, Jindong, et al.
Published: (2026)
by: Li, Jindong, et al.
Published: (2026)
EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks
by: Feng, Tongtong, et al.
Published: (2025)
by: Feng, Tongtong, et al.
Published: (2025)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Randomized Neural Networks for Partial Differential Equation on Static and Evolving Surfaces
by: Sun, Jingbo, et al.
Published: (2026)
by: Sun, Jingbo, et al.
Published: (2026)
WISE-Flow: Workflow-Induced Structured Experience for Self-Evolving Conversational Service Agents
by: Zhou, Yuqing, et al.
Published: (2026)
by: Zhou, Yuqing, et al.
Published: (2026)
MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents
by: Xie, Zhengwei, et al.
Published: (2026)
by: Xie, Zhengwei, et al.
Published: (2026)
RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment
by: Luo, Yingfeng, et al.
Published: (2026)
by: Luo, Yingfeng, et al.
Published: (2026)
Learning Evaluation Models from Large Language Models for Sequence Generation
by: Wang, Chenglong, et al.
Published: (2023)
by: Wang, Chenglong, et al.
Published: (2023)
Multi-Agent Evolve: LLM Self-Improve through Co-evolution
by: Chen, Yixing, et al.
Published: (2025)
by: Chen, Yixing, et al.
Published: (2025)
Multi-Satellite Multi-Stream Beamspace Massive MIMO Transmission
by: Wang, Yafei, et al.
Published: (2025)
by: Wang, Yafei, et al.
Published: (2025)
Agents of Change: Self-Evolving LLM Agents for Strategic Planning
by: Belle, Nikolas, et al.
Published: (2025)
by: Belle, Nikolas, et al.
Published: (2025)
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
by: Fan, Zhiyuan, et al.
Published: (2026)
by: Fan, Zhiyuan, et al.
Published: (2026)
Alita-G: Self-Evolving Generative Agent for Agent Generation
by: Qiu, Jiahao, et al.
Published: (2025)
by: Qiu, Jiahao, et al.
Published: (2025)
Smoothed Score Queries and the Complexity of Sampling
by: Liu, Jingbo
Published: (2026)
by: Liu, Jingbo
Published: (2026)
Proteoform Identification Using Multiplexed Top‐Down Mass Spectra
by: Zhige Wang, et al.
Published: (2025)
by: Zhige Wang, et al.
Published: (2025)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Similar Items
-
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
by: Huo, Yifu, et al.
Published: (2026) -
MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement Learning
by: Wang, Chenglong, et al.
Published: (2026) -
APR: Penalizing Structural Redundancy in Large Reasoning Models via Anchor-based Process Rewards
by: Chang, Kaiyan, et al.
Published: (2026) -
LRHP: Learning Representations for Human Preferences via Preference Pairs
by: Wang, Chenglong, et al.
Published: (2024) -
GRAM: A Generative Foundation Reward Model for Reward Generalization
by: Wang, Chenglong, et al.
Published: (2025)