SpecEval: Evaluating Model Adherence to Behavior Specifications
Fuente:
arXiv
Salvato in:
| Autori principali: | Ahmed, Ahmed, Klyman, Kevin, Zeng, Yi, Koyejo, Sanmi, Liang, Percy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
Reliable and Efficient Amortized Model-based Evaluation
di: Truong, Sang, et al.
Pubblicazione: (2025)
di: Truong, Sang, et al.
Pubblicazione: (2025)
SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications
di: Ma, Lezhi, et al.
Pubblicazione: (2024)
di: Ma, Lezhi, et al.
Pubblicazione: (2024)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
di: Vo, Truong, et al.
Pubblicazione: (2025)
di: Vo, Truong, et al.
Pubblicazione: (2025)
Scalable Ensembling For Mitigating Reward Overoptimisation
di: Ahmed, Ahmed M., et al.
Pubblicazione: (2024)
di: Ahmed, Ahmed M., et al.
Pubblicazione: (2024)
Acceptable Use Policies for Foundation Models
di: Klyman, Kevin
Pubblicazione: (2024)
di: Klyman, Kevin
Pubblicazione: (2024)
Discovering Implicit Large Language Model Alignment Objectives
di: Chen, Edward, et al.
Pubblicazione: (2026)
di: Chen, Edward, et al.
Pubblicazione: (2026)
Reasoning Models Don't Just Think Longer, They Move Differently
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
Independence Tests for Language Models
di: Zhu, Sally, et al.
Pubblicazione: (2025)
di: Zhu, Sally, et al.
Pubblicazione: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
di: Rafid, Ahmed, et al.
Pubblicazione: (2026)
di: Rafid, Ahmed, et al.
Pubblicazione: (2026)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
di: Schaeffer, Rylan, et al.
Pubblicazione: (2026)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2026)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
di: Liu, Ken Ziyu, et al.
Pubblicazione: (2025)
di: Liu, Ken Ziyu, et al.
Pubblicazione: (2025)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
di: McDuff, Daniel, et al.
Pubblicazione: (2025)
di: McDuff, Daniel, et al.
Pubblicazione: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
di: Dubois, Yann, et al.
Pubblicazione: (2024)
di: Dubois, Yann, et al.
Pubblicazione: (2024)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
di: Agarwal, Anmol, et al.
Pubblicazione: (2026)
di: Agarwal, Anmol, et al.
Pubblicazione: (2026)
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
di: Yang, Haote, et al.
Pubblicazione: (2025)
di: Yang, Haote, et al.
Pubblicazione: (2025)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
di: Patel, Fagun, et al.
Pubblicazione: (2025)
di: Patel, Fagun, et al.
Pubblicazione: (2025)
Scaling Laws for Downstream Task Performance of Large Language Models
di: Isik, Berivan, et al.
Pubblicazione: (2024)
di: Isik, Berivan, et al.
Pubblicazione: (2024)
AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
di: Liang, Yuanzhi, et al.
Pubblicazione: (2024)
di: Liang, Yuanzhi, et al.
Pubblicazione: (2024)
General Preference Reinforcement Learning
di: Umer, Muhammad, et al.
Pubblicazione: (2026)
di: Umer, Muhammad, et al.
Pubblicazione: (2026)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
di: Dejl, Adam, et al.
Pubblicazione: (2026)
di: Dejl, Adam, et al.
Pubblicazione: (2026)
AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
di: Shea, Ryan, et al.
Pubblicazione: (2025)
di: Shea, Ryan, et al.
Pubblicazione: (2025)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
di: Truong, Sang T., et al.
Pubblicazione: (2024)
di: Truong, Sang T., et al.
Pubblicazione: (2024)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
di: Zhou, Zhanke, et al.
Pubblicazione: (2025)
di: Zhou, Zhanke, et al.
Pubblicazione: (2025)
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
GATech at AbjadGenEval Shared Task: Multilingual Embeddings for Arabic Machine-Generated Text Classification
di: Khamis, Ahmed Khaled
Pubblicazione: (2026)
di: Khamis, Ahmed Khaled
Pubblicazione: (2026)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
di: Liang, Chumeng, et al.
Pubblicazione: (2025)
di: Liang, Chumeng, et al.
Pubblicazione: (2025)
Is Pre-training Truly Better Than Meta-Learning?
di: Miranda, Brando, et al.
Pubblicazione: (2023)
di: Miranda, Brando, et al.
Pubblicazione: (2023)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
di: Sun, Chongyan, et al.
Pubblicazione: (2024)
di: Sun, Chongyan, et al.
Pubblicazione: (2024)
SocialEval: Evaluating Social Intelligence of Large Language Models
di: Zhou, Jinfeng, et al.
Pubblicazione: (2025)
di: Zhou, Jinfeng, et al.
Pubblicazione: (2025)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
di: Chan, Willy, et al.
Pubblicazione: (2025)
di: Chan, Willy, et al.
Pubblicazione: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026) -
Reliable and Efficient Amortized Model-based Evaluation
di: Truong, Sang, et al.
Pubblicazione: (2025) -
SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications
di: Ma, Lezhi, et al.
Pubblicazione: (2024) -
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025) -
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
di: Vo, Truong, et al.
Pubblicazione: (2025)