Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
Fuente:
arXiv
Guardado en:
| Autores principales: | Cohen-Inger, Nurit, Elisha, Yehonatan, Shapira, Bracha, Rokach, Lior, Cohen, Seffi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
por: Cohen, Seffi, et al.
Publicado: (2025)
por: Cohen, Seffi, et al.
Publicado: (2025)
FairTTTS: A Tree Test Time Simulation Method for Fairness-Aware Classification
por: Cohen-Inger, Nurit, et al.
Publicado: (2025)
por: Cohen-Inger, Nurit, et al.
Publicado: (2025)
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
por: Cohen-Inger, Nurit, et al.
Publicado: (2025)
por: Cohen-Inger, Nurit, et al.
Publicado: (2025)
BagStacking: An Integrated Ensemble Learning Approach for Freezing of Gait Detection in Parkinson's Disease
por: Cohen, Seffi, et al.
Publicado: (2024)
por: Cohen, Seffi, et al.
Publicado: (2024)
Boosting Anomaly Detection Using Unsupervised Diverse Test-Time Augmentation
por: Cohen, Seffi, et al.
Publicado: (2021)
por: Cohen, Seffi, et al.
Publicado: (2021)
Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations
por: Elisha, Yehonatan, et al.
Publicado: (2025)
por: Elisha, Yehonatan, et al.
Publicado: (2025)
X-Cross: Dynamic Integration of Language Models for Cross-Domain Sequential Recommendation
por: Hadad, Guy, et al.
Publicado: (2025)
por: Hadad, Guy, et al.
Publicado: (2025)
SHAPoint: Task-Agnostic, Efficient, and Interpretable Point-Based Risk Scoring via Shapley Values
por: Meirman, Tomer D., et al.
Publicado: (2025)
por: Meirman, Tomer D., et al.
Publicado: (2025)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
por: Ovadia, Oded, et al.
Publicado: (2023)
por: Ovadia, Oded, et al.
Publicado: (2023)
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
por: Abdi, Immanuel, et al.
Publicado: (2026)
por: Abdi, Immanuel, et al.
Publicado: (2026)
Dark LLMs: The Growing Threat of Unaligned AI Models
por: Fire, Michael, et al.
Publicado: (2025)
por: Fire, Michael, et al.
Publicado: (2025)
What can Large Language Models Capture about Code Functional Equivalence?
por: Maveli, Nickil, et al.
Publicado: (2024)
por: Maveli, Nickil, et al.
Publicado: (2024)
Wings: Learning Multimodal LLMs without Text-only Forgetting
por: Zhang, Yi-Kai, et al.
Publicado: (2024)
por: Zhang, Yi-Kai, et al.
Publicado: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
por: Wang, Guangtao, et al.
Publicado: (2025)
por: Wang, Guangtao, et al.
Publicado: (2025)
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
por: Mittal, Avni
Publicado: (2026)
por: Mittal, Avni
Publicado: (2026)
Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
por: Maveli, Nickil, et al.
Publicado: (2026)
por: Maveli, Nickil, et al.
Publicado: (2026)
PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
por: Margalit, Yanki, et al.
Publicado: (2026)
por: Margalit, Yanki, et al.
Publicado: (2026)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
por: Park, Jungsoo, et al.
Publicado: (2025)
por: Park, Jungsoo, et al.
Publicado: (2025)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
por: Liu, Chi, et al.
Publicado: (2026)
por: Liu, Chi, et al.
Publicado: (2026)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
por: McGovern, Hope, et al.
Publicado: (2026)
por: McGovern, Hope, et al.
Publicado: (2026)
Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
por: Elisha, Yehonatan, et al.
Publicado: (2026)
por: Elisha, Yehonatan, et al.
Publicado: (2026)
Watch Your Steps: Observable and Modular Chains of Thought
por: Cohen, Cassandra A., et al.
Publicado: (2024)
por: Cohen, Cassandra A., et al.
Publicado: (2024)
Attention Is All You Need for KV Cache in Diffusion LLMs
por: Nguyen-Tri, Quan, et al.
Publicado: (2025)
por: Nguyen-Tri, Quan, et al.
Publicado: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
por: Mayne, Harry, et al.
Publicado: (2025)
por: Mayne, Harry, et al.
Publicado: (2025)
How Likely Do LLMs with CoT Mimic Human Reasoning?
por: Bao, Guangsheng, et al.
Publicado: (2024)
por: Bao, Guangsheng, et al.
Publicado: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
por: Kapoor, Sanyam, et al.
Publicado: (2024)
por: Kapoor, Sanyam, et al.
Publicado: (2024)
How Much Can We Forget about Data Contamination?
por: Bordt, Sebastian, et al.
Publicado: (2024)
por: Bordt, Sebastian, et al.
Publicado: (2024)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
por: Iskander, Shadi, et al.
Publicado: (2024)
por: Iskander, Shadi, et al.
Publicado: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
What Large Language Models Know and What People Think They Know
por: Steyvers, Mark, et al.
Publicado: (2024)
por: Steyvers, Mark, et al.
Publicado: (2024)
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
por: Miceli-Barone, Antonio Valerio, et al.
Publicado: (2026)
por: Miceli-Barone, Antonio Valerio, et al.
Publicado: (2026)
AgentBench: Evaluating LLMs as Agents
por: Liu, Xiao, et al.
Publicado: (2023)
por: Liu, Xiao, et al.
Publicado: (2023)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
por: Davoodi, Arash Gholami, et al.
Publicado: (2024)
por: Davoodi, Arash Gholami, et al.
Publicado: (2024)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
por: Xu, Yijie, et al.
Publicado: (2025)
por: Xu, Yijie, et al.
Publicado: (2025)
NanoKnow: How to Know What Your Language Model Knows
por: Gu, Lingwei, et al.
Publicado: (2026)
por: Gu, Lingwei, et al.
Publicado: (2026)
Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens
por: Koh, Seunghee, et al.
Publicado: (2026)
por: Koh, Seunghee, et al.
Publicado: (2026)
PersonaGym: Evaluating Persona Agents and LLMs
por: Samuel, Vinay, et al.
Publicado: (2024)
por: Samuel, Vinay, et al.
Publicado: (2024)
Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment
por: Kumar, Sayantan, et al.
Publicado: (2026)
por: Kumar, Sayantan, et al.
Publicado: (2026)
Learning is Forgetting: LLM Training As Lossy Compression
por: Conklin, Henry C., et al.
Publicado: (2026)
por: Conklin, Henry C., et al.
Publicado: (2026)
More Compute Is What You Need
por: Guo, Zhen
Publicado: (2024)
por: Guo, Zhen
Publicado: (2024)
Ejemplares similares
-
DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance
por: Cohen, Seffi, et al.
Publicado: (2025) -
FairTTTS: A Tree Test Time Simulation Method for Fairness-Aware Classification
por: Cohen-Inger, Nurit, et al.
Publicado: (2025) -
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
por: Cohen-Inger, Nurit, et al.
Publicado: (2025) -
BagStacking: An Integrated Ensemble Learning Approach for Freezing of Gait Detection in Parkinson's Disease
por: Cohen, Seffi, et al.
Publicado: (2024) -
Boosting Anomaly Detection Using Unsupervised Diverse Test-Time Augmentation
por: Cohen, Seffi, et al.
Publicado: (2021)