MEME: Multi-entity & Evolving Memory Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Jung, Seokwon, Rubinstein, Alexander, Uselis, Arnas, Yun, Sangdoo, Oh, Seong Joon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Intermediate Layer Classifiers for OOD generalization
por: Uselis, Arnas, et al.
Publicado: (2025)
por: Uselis, Arnas, et al.
Publicado: (2025)
MASEval: Extending Multi-Agent Evaluation from Models to Systems
por: Emde, Cornelius, et al.
Publicado: (2026)
por: Emde, Cornelius, et al.
Publicado: (2026)
How can embedding models bind concepts?
por: Uselis, Arnas, et al.
Publicado: (2026)
por: Uselis, Arnas, et al.
Publicado: (2026)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
por: Uselis, Arnas, et al.
Publicado: (2026)
por: Uselis, Arnas, et al.
Publicado: (2026)
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
por: Koishigarina, Darina, et al.
Publicado: (2025)
por: Koishigarina, Darina, et al.
Publicado: (2025)
Does Data Scaling Lead to Visual Compositional Generalization?
por: Uselis, Arnas, et al.
Publicado: (2025)
por: Uselis, Arnas, et al.
Publicado: (2025)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
por: Puerto, Haritz, et al.
Publicado: (2024)
por: Puerto, Haritz, et al.
Publicado: (2024)
Calibrating Large Language Models Using Their Generations Only
por: Ulmer, Dennis, et al.
Publicado: (2024)
por: Ulmer, Dennis, et al.
Publicado: (2024)
Dr.LLM: Dynamic Layer Routing in LLMs
por: Heakl, Ahmed, et al.
Publicado: (2025)
por: Heakl, Ahmed, et al.
Publicado: (2025)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
por: Gubri, Martin, et al.
Publicado: (2024)
por: Gubri, Martin, et al.
Publicado: (2024)
Compressed Context Memory For Online Language Model Interaction
por: Kim, Jang-Hyun, et al.
Publicado: (2023)
por: Kim, Jang-Hyun, et al.
Publicado: (2023)
Half-Truths Break Similarity-Based Retrieval
por: Kargi, Bora, et al.
Publicado: (2026)
por: Kargi, Bora, et al.
Publicado: (2026)
On the rankability of visual embeddings
por: Sonthalia, Ankit, et al.
Publicado: (2025)
por: Sonthalia, Ankit, et al.
Publicado: (2025)
Task-Synchronized Recurrent Neural Networks
por: Lukoševičius, Mantas, et al.
Publicado: (2022)
por: Lukoševičius, Mantas, et al.
Publicado: (2022)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
por: Do, Heejin, et al.
Publicado: (2025)
por: Do, Heejin, et al.
Publicado: (2025)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
por: Morelli, Fabian, et al.
Publicado: (2026)
por: Morelli, Fabian, et al.
Publicado: (2026)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
por: Jeong, Yujin, et al.
Publicado: (2025)
por: Jeong, Yujin, et al.
Publicado: (2025)
Studying Large Language Model Behaviors Under Context-Memory Conflicts With Real Documents
por: Kortukov, Evgenii, et al.
Publicado: (2024)
por: Kortukov, Evgenii, et al.
Publicado: (2024)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
por: Kim, Jang-Hyun, et al.
Publicado: (2026)
por: Kim, Jang-Hyun, et al.
Publicado: (2026)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
por: Yoo, Haneul, et al.
Publicado: (2024)
por: Yoo, Haneul, et al.
Publicado: (2024)
DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
por: Rubinstein, Alexander, et al.
Publicado: (2025)
por: Rubinstein, Alexander, et al.
Publicado: (2025)
Scalable Ensemble Diversification for OOD Generalization and Detection
por: Rubinstein, Alexander, et al.
Publicado: (2024)
por: Rubinstein, Alexander, et al.
Publicado: (2024)
Are We Done with Object-Centric Learning?
por: Rubinstein, Alexander, et al.
Publicado: (2025)
por: Rubinstein, Alexander, et al.
Publicado: (2025)
Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
por: Goel, Anmol, et al.
Publicado: (2026)
por: Goel, Anmol, et al.
Publicado: (2026)
Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
por: Goo, Sungwoo, et al.
Publicado: (2026)
por: Goo, Sungwoo, et al.
Publicado: (2026)
Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
por: Shim, Jung-Woo, et al.
Publicado: (2025)
por: Shim, Jung-Woo, et al.
Publicado: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
por: Jeong, Yujin, et al.
Publicado: (2026)
por: Jeong, Yujin, et al.
Publicado: (2026)
LLM generation novelty through the lens of semantic similarity
por: Davydov, Philipp, et al.
Publicado: (2025)
por: Davydov, Philipp, et al.
Publicado: (2025)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
por: Scimeca, Luca, et al.
Publicado: (2023)
por: Scimeca, Luca, et al.
Publicado: (2023)
An Evolved Universal Transformer Memory
por: Cetin, Edoardo, et al.
Publicado: (2024)
por: Cetin, Edoardo, et al.
Publicado: (2024)
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
por: Green, Tommaso, et al.
Publicado: (2025)
por: Green, Tommaso, et al.
Publicado: (2025)
C-SEO Bench: Does Conversational SEO Work?
por: Puerto, Haritz, et al.
Publicado: (2025)
por: Puerto, Haritz, et al.
Publicado: (2025)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
por: Kim, Seungone, et al.
Publicado: (2023)
por: Kim, Seungone, et al.
Publicado: (2023)
Do Deep Neural Network Solutions Form a Star Domain?
por: Sonthalia, Ankit, et al.
Publicado: (2024)
por: Sonthalia, Ankit, et al.
Publicado: (2024)
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
por: Zhang, Haozhen, et al.
Publicado: (2026)
por: Zhang, Haozhen, et al.
Publicado: (2026)
CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
por: Shim, Jung-Woo, et al.
Publicado: (2025)
por: Shim, Jung-Woo, et al.
Publicado: (2025)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
por: Kirchhof, Michael, et al.
Publicado: (2025)
por: Kirchhof, Michael, et al.
Publicado: (2025)
Exploring prompts to elicit memorization in masked language model-based named entity recognition
por: Xia, Yuxi, et al.
Publicado: (2024)
por: Xia, Yuxi, et al.
Publicado: (2024)
SNOBERT: A Benchmark for clinical notes entity linking in the SNOMED CT clinical terminology
por: Kulyabin, Mikhail, et al.
Publicado: (2024)
por: Kulyabin, Mikhail, et al.
Publicado: (2024)
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
por: Lin, Huawei, et al.
Publicado: (2026)
por: Lin, Huawei, et al.
Publicado: (2026)
Ejemplares similares
-
Intermediate Layer Classifiers for OOD generalization
por: Uselis, Arnas, et al.
Publicado: (2025) -
MASEval: Extending Multi-Agent Evaluation from Models to Systems
por: Emde, Cornelius, et al.
Publicado: (2026) -
How can embedding models bind concepts?
por: Uselis, Arnas, et al.
Publicado: (2026) -
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
por: Uselis, Arnas, et al.
Publicado: (2026) -
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
por: Koishigarina, Darina, et al.
Publicado: (2025)