Decomposing and Measuring Evaluation Awareness
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Changling, Zhang, Terry Jingchen, Zhang, Jie, Jin, Zhijing, Abdelnabi, Sahar, Andriushchenko, Maksym |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Does Refusal Training in LLMs Generalize to the Past Tense?
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
di: Zhao, Hao, et al.
Pubblicazione: (2024)
di: Zhao, Hao, et al.
Pubblicazione: (2024)
Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
di: Schmotz, David, et al.
Pubblicazione: (2025)
di: Schmotz, David, et al.
Pubblicazione: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
di: Goel, Shashwat, et al.
Pubblicazione: (2026)
di: Goel, Shashwat, et al.
Pubblicazione: (2026)
Causality for Natural Language Processing
di: Jin, Zhijing
Pubblicazione: (2025)
di: Jin, Zhijing
Pubblicazione: (2025)
QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
di: Qin, Jeremy, et al.
Pubblicazione: (2026)
di: Qin, Jeremy, et al.
Pubblicazione: (2026)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
di: Liu, Xinge, et al.
Pubblicazione: (2026)
di: Liu, Xinge, et al.
Pubblicazione: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
di: Schmotz, David, et al.
Pubblicazione: (2026)
di: Schmotz, David, et al.
Pubblicazione: (2026)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
di: Qi, Xuan, et al.
Pubblicazione: (2025)
di: Qi, Xuan, et al.
Pubblicazione: (2025)
Evaluating Cooperation in LLM Social Groups through Elected Leadership
di: Faulkner, Ryan, et al.
Pubblicazione: (2026)
di: Faulkner, Ryan, et al.
Pubblicazione: (2026)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
di: Rando, Javier, et al.
Pubblicazione: (2024)
di: Rando, Javier, et al.
Pubblicazione: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
di: Simko, Samuel, et al.
Pubblicazione: (2025)
di: Simko, Samuel, et al.
Pubblicazione: (2025)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
di: Lv, Keyu, et al.
Pubblicazione: (2026)
di: Lv, Keyu, et al.
Pubblicazione: (2026)
Models That Know How Evaluations Are Designed Score Safer
di: Deckenbach, Katharina, et al.
Pubblicazione: (2026)
di: Deckenbach, Katharina, et al.
Pubblicazione: (2026)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Voices of Her: Analyzing Gender Differences in the AI Publication World
di: Ding, Yiwen, et al.
Pubblicazione: (2023)
di: Ding, Yiwen, et al.
Pubblicazione: (2023)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
di: Kundurthy, Srivatsa, et al.
Pubblicazione: (2026)
di: Kundurthy, Srivatsa, et al.
Pubblicazione: (2026)
Towards More Effective Table-to-Text Generation: Assessing In-Context Learning and Self-Evaluation with Open-Source Models
di: Iravani, Sahar, et al.
Pubblicazione: (2024)
di: Iravani, Sahar, et al.
Pubblicazione: (2024)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Decomposing Attention To Find Context-Sensitive Neurons
di: Gibson, Alex
Pubblicazione: (2025)
di: Gibson, Alex
Pubblicazione: (2025)
Deep Language Geometry: Constructing a Metric Space from LLM Weights
di: Shamrai, Maksym, et al.
Pubblicazione: (2025)
di: Shamrai, Maksym, et al.
Pubblicazione: (2025)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
di: Huang, Xinting, et al.
Pubblicazione: (2025)
di: Huang, Xinting, et al.
Pubblicazione: (2025)
3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection
di: Ding, Hongxin, et al.
Pubblicazione: (2024)
di: Ding, Hongxin, et al.
Pubblicazione: (2024)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
di: Tsujimura, Hikaru, et al.
Pubblicazione: (2025)
di: Tsujimura, Hikaru, et al.
Pubblicazione: (2025)
Lean Meets Theoretical Computer Science: Scalable Synthesis of Theorem Proving Challenges in Formal-Informal Pairs
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
di: Qin, Tian, et al.
Pubblicazione: (2025)
di: Qin, Tian, et al.
Pubblicazione: (2025)
PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head
di: Lin, Chieh-Yen, et al.
Pubblicazione: (2026)
di: Lin, Chieh-Yen, et al.
Pubblicazione: (2026)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
di: Hamilton, Sil, et al.
Pubblicazione: (2025)
di: Hamilton, Sil, et al.
Pubblicazione: (2025)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
di: Liao, Huanxuan, et al.
Pubblicazione: (2025)
di: Liao, Huanxuan, et al.
Pubblicazione: (2025)
Can Large Language Models Infer Causation from Correlation?
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory
di: Zhang, Haozhen, et al.
Pubblicazione: (2026)
di: Zhang, Haozhen, et al.
Pubblicazione: (2026)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
di: Feng, Jie, et al.
Pubblicazione: (2024)
di: Feng, Jie, et al.
Pubblicazione: (2024)
A Theory of Response Sampling in LLMs: Part Descriptive and Part Prescriptive
di: Sivaprasad, Sarath, et al.
Pubblicazione: (2024)
di: Sivaprasad, Sarath, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Does Refusal Training in LLMs Generalize to the Past Tense?
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024) -
Is In-Context Learning Sufficient for Instruction Following in LLMs?
di: Zhao, Hao, et al.
Pubblicazione: (2024) -
Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
di: Schmotz, David, et al.
Pubblicazione: (2025) -
Capability-Based Scaling Trends for LLM-Based Red-Teaming
di: Panfilov, Alexander, et al.
Pubblicazione: (2025) -
FutureSim: Replaying World Events to Evaluate Adaptive Agents
di: Goel, Shashwat, et al.
Pubblicazione: (2026)