LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
Fuente:
arXiv
Guardado en:
| Autores principales: | Ivanov, Igor, Africa, David Demitri |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Steering Awareness: Detecting Activation Steering from Within
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026)
por: Imran, Sohaib, et al.
Publicado: (2026)
Learning Dynamics of Meta-Learning in Small Model Pretraining
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
por: Weiss, Yuval, et al.
Publicado: (2025)
por: Weiss, Yuval, et al.
Publicado: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Does Self-Evaluation Enable Wireheading in Language Models?
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
por: Martinez, Richard Diehl, et al.
Publicado: (2025)
por: Martinez, Richard Diehl, et al.
Publicado: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
por: Tan, Daniel, et al.
Publicado: (2025)
por: Tan, Daniel, et al.
Publicado: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025)
por: Africa, David Demitri
Publicado: (2025)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
por: Goel, Shashwat, et al.
Publicado: (2026)
por: Goel, Shashwat, et al.
Publicado: (2026)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
por: Montalan, Jann Railey, et al.
Publicado: (2025)
por: Montalan, Jann Railey, et al.
Publicado: (2025)
WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
por: Ning, Kangyun, et al.
Publicado: (2024)
por: Ning, Kangyun, et al.
Publicado: (2024)
GameArena: Evaluating LLM Reasoning through Live Computer Games
por: Hu, Lanxiang, et al.
Publicado: (2024)
por: Hu, Lanxiang, et al.
Publicado: (2024)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
por: Lu, Yiyang, et al.
Publicado: (2026)
por: Lu, Yiyang, et al.
Publicado: (2026)
Probing and Steering Evaluation Awareness of Language Models
por: Nguyen, Jord, et al.
Publicado: (2025)
por: Nguyen, Jord, et al.
Publicado: (2025)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
por: Guo, Pei-Fu, et al.
Publicado: (2025)
por: Guo, Pei-Fu, et al.
Publicado: (2025)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
por: Patel, Liana, et al.
Publicado: (2025)
por: Patel, Liana, et al.
Publicado: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
por: Zhang, Hongbin, et al.
Publicado: (2025)
por: Zhang, Hongbin, et al.
Publicado: (2025)
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
por: Hua, Tim Tian, et al.
Publicado: (2025)
por: Hua, Tim Tian, et al.
Publicado: (2025)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
por: Yuan, Youliang, et al.
Publicado: (2025)
por: Yuan, Youliang, et al.
Publicado: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
por: Mekky, Ali, et al.
Publicado: (2025)
por: Mekky, Ali, et al.
Publicado: (2025)
Syntactic Evolution in Language Usage
por: Kumar, Surbhit
Publicado: (2025)
por: Kumar, Surbhit
Publicado: (2025)
Decomposing and Measuring Evaluation Awareness
por: Li, Changling, et al.
Publicado: (2026)
por: Li, Changling, et al.
Publicado: (2026)
FineSurE: Fine-grained Summarization Evaluation using LLMs
por: Song, Hwanjun, et al.
Publicado: (2024)
por: Song, Hwanjun, et al.
Publicado: (2024)
CALRK-Bench: Evaluating Context-Aware Legal Reasoning in Korean Law
por: Jung, JiHyeok, et al.
Publicado: (2026)
por: Jung, JiHyeok, et al.
Publicado: (2026)
PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
por: Pang, Tsz Fung, et al.
Publicado: (2025)
por: Pang, Tsz Fung, et al.
Publicado: (2025)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
por: Cettolo, Mauro, et al.
Publicado: (2025)
por: Cettolo, Mauro, et al.
Publicado: (2025)
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
por: Imai, Saki, et al.
Publicado: (2025)
por: Imai, Saki, et al.
Publicado: (2025)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
por: Luo, Zhimeng, et al.
Publicado: (2025)
por: Luo, Zhimeng, et al.
Publicado: (2025)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
por: Kumar, Shanu, et al.
Publicado: (2024)
por: Kumar, Shanu, et al.
Publicado: (2024)
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
por: Martynov, Nikita, et al.
Publicado: (2025)
por: Martynov, Nikita, et al.
Publicado: (2025)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
por: Dawson, Fiifi, et al.
Publicado: (2024)
por: Dawson, Fiifi, et al.
Publicado: (2024)
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
por: Ding, Liang
Publicado: (2026)
por: Ding, Liang
Publicado: (2026)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
por: Chhabra, Mukul, et al.
Publicado: (2026)
por: Chhabra, Mukul, et al.
Publicado: (2026)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
por: Xu, Xi, et al.
Publicado: (2024)
por: Xu, Xi, et al.
Publicado: (2024)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
por: Dekoninck, Jasper, et al.
Publicado: (2024)
por: Dekoninck, Jasper, et al.
Publicado: (2024)
Prioritized Replay for RL Post-training
por: Fatemi, Mehdi
Publicado: (2026)
por: Fatemi, Mehdi
Publicado: (2026)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
por: Schlangen, David
Publicado: (2024)
por: Schlangen, David
Publicado: (2024)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
por: Butt, Natasha, et al.
Publicado: (2024)
por: Butt, Natasha, et al.
Publicado: (2024)
Ejemplares similares
-
Steering Awareness: Detecting Activation Steering from Within
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025) -
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026) -
Learning Dynamics of Meta-Learning in Small Model Pretraining
por: Africa, David Demitri, et al.
Publicado: (2025) -
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
por: Weiss, Yuval, et al.
Publicado: (2025) -
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
por: Africa, David Demitri, et al.
Publicado: (2025)