Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Songwei, Chen, Zihan, Shi, Chengshuai, Wang, Peng, Li, Jundong, Shen, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Prompt Optimization Through the Lens of Best Arm Identification
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
FastGAS: Fast Graph-based Annotation Selection for In-Context Learning
by: Chen, Zihan, et al.
Published: (2024)
by: Chen, Zihan, et al.
Published: (2024)
FedHERO: A Federated Learning Approach for Node Classification Task on Heterophilic Graphs
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
by: Li, Donghao, et al.
Published: (2026)
by: Li, Donghao, et al.
Published: (2026)
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
by: Wu, Di, et al.
Published: (2026)
by: Wu, Di, et al.
Published: (2026)
Probing to Refine: Reinforcement Distillation of LLMs via Explanatory Inversion
by: Tan, Zhen, et al.
Published: (2026)
by: Tan, Zhen, et al.
Published: (2026)
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Graph Prompting for Graph Learning Models: Recent Advances and Future Directions
by: Fu, Xingbo, et al.
Published: (2025)
by: Fu, Xingbo, et al.
Published: (2025)
ST-FiT: Inductive Spatial-Temporal Forecasting with Limited Training Data
by: Lei, Zhenyu, et al.
Published: (2024)
by: Lei, Zhenyu, et al.
Published: (2024)
From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
Memory Retention Is Not Enough to Master Memory Tasks in Reinforcement Learning
by: Shchendrigin, Oleg, et al.
Published: (2026)
by: Shchendrigin, Oleg, et al.
Published: (2026)
Potent but Stealthy: Rethink Profile Pollution against Sequential Recommendation via Bi-level Constrained Reinforcement Paradigm
by: Su, Jiajie, et al.
Published: (2025)
by: Su, Jiajie, et al.
Published: (2025)
GraphTOP: Graph Topology-Oriented Prompting for Graph Neural Networks
by: Fu, Xingbo, et al.
Published: (2025)
by: Fu, Xingbo, et al.
Published: (2025)
One Sample is Enough to Make Conformal Prediction Robust
by: Zargarbashi, Soroush H., et al.
Published: (2025)
by: Zargarbashi, Soroush H., et al.
Published: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
One-Way Policy Optimization for Self-Evolving LLMs
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Parameter Efficiency Is Not Memory Efficiency: Rethinking Fine-Tuning for On-Device LLM Adaptation
by: Tenison, Irene, et al.
Published: (2026)
by: Tenison, Irene, et al.
Published: (2026)
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Graph Neural Networks Are More Than Filters: Revisiting and Benchmarking from A Spectral Perspective
by: Dong, Yushun, et al.
Published: (2024)
by: Dong, Yushun, et al.
Published: (2024)
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models
by: Balef, Amir Rezaei, et al.
Published: (2026)
by: Balef, Amir Rezaei, et al.
Published: (2026)
Sequential Stochastic Combinatorial Optimization Using Hierarchal Reinforcement Learning
by: Feng, Xinsong, et al.
Published: (2025)
by: Feng, Xinsong, et al.
Published: (2025)
Federated Graph Learning with Structure Proxy Alignment
by: Fu, Xingbo, et al.
Published: (2024)
by: Fu, Xingbo, et al.
Published: (2024)
Self-Evolving Curriculum for LLM Reasoning
by: Chen, Xiaoyin, et al.
Published: (2025)
by: Chen, Xiaoyin, et al.
Published: (2025)
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
When Drafts Evolve: Speculative Decoding Meets Online Learning
by: Qian, Yu-Yang, et al.
Published: (2026)
by: Qian, Yu-Yang, et al.
Published: (2026)
FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation
by: Ding, Zhihao, et al.
Published: (2026)
by: Ding, Zhihao, et al.
Published: (2026)
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
by: Lin, Huawei, et al.
Published: (2026)
by: Lin, Huawei, et al.
Published: (2026)
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
by: Song, Chuxu, et al.
Published: (2026)
by: Song, Chuxu, et al.
Published: (2026)
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
by: Wang, Yuyao, et al.
Published: (2026)
by: Wang, Yuyao, et al.
Published: (2026)
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
by: Zhang, Yanzhi, et al.
Published: (2025)
by: Zhang, Yanzhi, et al.
Published: (2025)
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory
by: Li, Sijia, et al.
Published: (2025)
by: Li, Sijia, et al.
Published: (2025)
Edge Prompt Tuning for Graph Neural Networks
by: Fu, Xingbo, et al.
Published: (2025)
by: Fu, Xingbo, et al.
Published: (2025)
Is Distance Matrix Enough for Geometric Deep Learning?
by: Li, Zian, et al.
Published: (2023)
by: Li, Zian, et al.
Published: (2023)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Similar Items
-
Efficient Prompt Optimization Through the Lens of Best Arm Identification
by: Shi, Chengshuai, et al.
Published: (2024) -
FastGAS: Fast Graph-based Annotation Selection for In-Context Learning
by: Chen, Zihan, et al.
Published: (2024) -
FedHERO: A Federated Learning Approach for Node Classification Task on Heterophilic Graphs
by: Chen, Zihan, et al.
Published: (2025) -
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
by: Li, Donghao, et al.
Published: (2026) -
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)