Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yifei, Guo, Weidong, Zhang, Lingling, Xu, Rongman, Huang, Muye, Liu, Hui, Xu, Lijiao, Xu, Yu, Liu, Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
di: Shi, Zhichao, et al.
Pubblicazione: (2025)
di: Shi, Zhichao, et al.
Pubblicazione: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
di: Tan, Haoran, et al.
Pubblicazione: (2025)
di: Tan, Haoran, et al.
Pubblicazione: (2025)
StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall
di: Wu, Yerong, et al.
Pubblicazione: (2026)
di: Wu, Yerong, et al.
Pubblicazione: (2026)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
di: Kim, Doyoung, et al.
Pubblicazione: (2026)
di: Kim, Doyoung, et al.
Pubblicazione: (2026)
On the Structural Memory of LLM Agents
di: Zeng, Ruihong, et al.
Pubblicazione: (2024)
di: Zeng, Ruihong, et al.
Pubblicazione: (2024)
Learning from Semi-Factuals: A Debiased and Semantic-Aware Framework for Generalized Relation Discovery
di: Wang, Jiaxin, et al.
Pubblicazione: (2024)
di: Wang, Jiaxin, et al.
Pubblicazione: (2024)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
di: Xu, Yunqi, et al.
Pubblicazione: (2024)
di: Xu, Yunqi, et al.
Pubblicazione: (2024)
Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual Errors
di: Chandler, Alex, et al.
Pubblicazione: (2024)
di: Chandler, Alex, et al.
Pubblicazione: (2024)
Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
di: Wang, Haoyu, et al.
Pubblicazione: (2026)
Beyond Static Summarization: Proactive Memory Extraction for LLM Agents
di: Yang, Chengyuan, et al.
Pubblicazione: (2026)
di: Yang, Chengyuan, et al.
Pubblicazione: (2026)
Evaluating the Factuality of Large Language Models using Large-Scale Knowledge Graphs
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
di: Ning, Yucheng, et al.
Pubblicazione: (2025)
di: Ning, Yucheng, et al.
Pubblicazione: (2025)
DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automation
di: You, Ziming, et al.
Pubblicazione: (2025)
di: You, Ziming, et al.
Pubblicazione: (2025)
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
di: Hu, Yuanzhe, et al.
Pubblicazione: (2025)
di: Hu, Yuanzhe, et al.
Pubblicazione: (2025)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
di: Deng, Shihan, et al.
Pubblicazione: (2024)
di: Deng, Shihan, et al.
Pubblicazione: (2024)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
di: Wang, Xintao, et al.
Pubblicazione: (2025)
di: Wang, Xintao, et al.
Pubblicazione: (2025)
Adaptive Memory Admission Control for LLM Agents
di: Zhang, Guilin, et al.
Pubblicazione: (2026)
di: Zhang, Guilin, et al.
Pubblicazione: (2026)
Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
di: Xu, Fangzhi, et al.
Pubblicazione: (2023)
di: Xu, Fangzhi, et al.
Pubblicazione: (2023)
JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
di: Sun, Wangtao, et al.
Pubblicazione: (2024)
di: Sun, Wangtao, et al.
Pubblicazione: (2024)
ExpeL: LLM Agents Are Experiential Learners
di: Zhao, Andrew, et al.
Pubblicazione: (2023)
di: Zhao, Andrew, et al.
Pubblicazione: (2023)
Reasoning Factual Knowledge in Structured Data with Large Language Models
di: Huang, Sirui, et al.
Pubblicazione: (2024)
di: Huang, Sirui, et al.
Pubblicazione: (2024)
Permutation-Consensus Listwise Judging for Robust Factuality Evaluation
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
MIRIX: Multi-Agent Memory System for LLM-Based Agents
di: Wang, Yu, et al.
Pubblicazione: (2025)
di: Wang, Yu, et al.
Pubblicazione: (2025)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
di: Xu, Derong, et al.
Pubblicazione: (2024)
di: Xu, Derong, et al.
Pubblicazione: (2024)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
di: Yang, Xintong, et al.
Pubblicazione: (2026)
di: Yang, Xintong, et al.
Pubblicazione: (2026)
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
G-MemLLM: Gated Latent Memory Augmentation for Long-Context Reasoning in Large Language Models
di: Xu, Xun
Pubblicazione: (2026)
di: Xu, Xun
Pubblicazione: (2026)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
di: Chen, Yuhao, et al.
Pubblicazione: (2026)
di: Chen, Yuhao, et al.
Pubblicazione: (2026)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
di: Hu, Xuming, et al.
Pubblicazione: (2024)
di: Hu, Xuming, et al.
Pubblicazione: (2024)
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
di: Huang, Minghui
Pubblicazione: (2025)
di: Huang, Minghui
Pubblicazione: (2025)
How Does Response Length Affect Long-Form Factuality
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
di: Ramprasad, Sanjana, et al.
Pubblicazione: (2024)
di: Ramprasad, Sanjana, et al.
Pubblicazione: (2024)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
di: Jiao, Rui, et al.
Pubblicazione: (2025)
di: Jiao, Rui, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering
di: Xu, Yao, et al.
Pubblicazione: (2024)
di: Xu, Yao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM
di: Yang, Jiuding, et al.
Pubblicazione: (2024) -
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
di: Shi, Zhichao, et al.
Pubblicazione: (2025) -
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
di: Tan, Haoran, et al.
Pubblicazione: (2025) -
StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall
di: Wu, Yerong, et al.
Pubblicazione: (2026) -
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
di: Kim, Doyoung, et al.
Pubblicazione: (2026)