LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
Fuente:
arXiv
Saved in:
| Main Authors: | Saad-Falcon, Jon, Vivek, Rajan, Berrios, William, Naik, Nandita Shankar, Franklin, Matija, Vidgen, Bertie, Singh, Amanpreet, Kiela, Douwe, Mehri, Shikib |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reflective Context Learning: Studying the Optimization Primitives of Context Space
by: Vassilyev, Nikita, et al.
Published: (2026)
by: Vassilyev, Nikita, et al.
Published: (2026)
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
by: D'Oosterlinck, Karel, et al.
Published: (2024)
by: D'Oosterlinck, Karel, et al.
Published: (2024)
Anchor Points: Benchmarking Models with Much Fewer Examples
by: Vivek, Rajan, et al.
Published: (2023)
by: Vivek, Rajan, et al.
Published: (2023)
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
by: Tran, Dat, et al.
Published: (2026)
by: Tran, Dat, et al.
Published: (2026)
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
by: Lui, Nicholas, et al.
Published: (2023)
by: Lui, Nicholas, et al.
Published: (2023)
Generative Representational Instruction Tuning
by: Muennighoff, Niklas, et al.
Published: (2024)
by: Muennighoff, Niklas, et al.
Published: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025)
by: Willats, Richard, et al.
Published: (2025)
I am a Strange Dataset: Metalinguistic Tests for Language Models
by: Thrush, Tristan, et al.
Published: (2024)
by: Thrush, Tristan, et al.
Published: (2024)
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
by: Agrawal, Sheshansh, et al.
Published: (2026)
by: Agrawal, Sheshansh, et al.
Published: (2026)
Nearest Neighbor Normalization Improves Multimodal Retrieval
by: Chowdhury, Neil, et al.
Published: (2024)
by: Chowdhury, Neil, et al.
Published: (2024)
Document Optimization for Black-Box Retrieval via Reinforcement Learning
by: Uzan, Omri, et al.
Published: (2026)
by: Uzan, Omri, et al.
Published: (2026)
Goal Alignment in LLM-Based User Simulators for Conversational AI
by: Mehri, Shuhaib, et al.
Published: (2025)
by: Mehri, Shuhaib, et al.
Published: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
by: Vidgen, Bertie, et al.
Published: (2023)
by: Vidgen, Bertie, et al.
Published: (2023)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
by: Kirk, Hannah Rose, et al.
Published: (2026)
by: Kirk, Hannah Rose, et al.
Published: (2026)
KTO: Model Alignment as Prospect Theoretic Optimization
by: Ethayarajh, Kawin, et al.
Published: (2024)
by: Ethayarajh, Kawin, et al.
Published: (2024)
Approximate Analytical Solution using Power Series Method for the Propagation of Blast Waves in a Rotational Axisymmetric non-ideal Gas
by: Nandita, et al.
Published: (2023)
by: Nandita, et al.
Published: (2023)
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
Lynx: An Open Source Hallucination Evaluation Model
by: Ravi, Selvan Sunitha, et al.
Published: (2024)
by: Ravi, Selvan Sunitha, et al.
Published: (2024)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
by: Styles, Olly, et al.
Published: (2024)
by: Styles, Olly, et al.
Published: (2024)
Enhancing Software Vulnerability Detection Using Code Property Graphs and Convolutional Neural Networks
by: Saimbhi, Amanpreet Singh
Published: (2025)
by: Saimbhi, Amanpreet Singh
Published: (2025)
Model-Free RL Agents Demonstrate System 1-Like Intentionality
by: Ashton, Hal, et al.
Published: (2025)
by: Ashton, Hal, et al.
Published: (2025)
Fine-grained Testing for Autonomous Driving Software: a Study on Autoware with LLM-driven Unit Testing
by: Wang, Wenhan, et al.
Published: (2025)
by: Wang, Wenhan, et al.
Published: (2025)
The AI Consumer Index (ACE)
by: Benchek, Julien, et al.
Published: (2025)
by: Benchek, Julien, et al.
Published: (2025)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
by: Ferguson, Nick, et al.
Published: (2026)
by: Ferguson, Nick, et al.
Published: (2026)
Extending the noise of splitting to its completion and stability of Brownian maxima
by: Vidmar, Matija, et al.
Published: (2024)
by: Vidmar, Matija, et al.
Published: (2024)
Upcycling waste plastics into 3D printing filaments for a circular and sustainable future
by: Man Vir Singh, et al.
Published: (2026)
by: Man Vir Singh, et al.
Published: (2026)
Unit and distinct distances in typical norms
by: Alon, Noga, et al.
Published: (2023)
by: Alon, Noga, et al.
Published: (2023)
Transcendental Optimization in Geometric Programming via Power Series Approximations
by: singh, Amanpreet
Published: (2025)
by: singh, Amanpreet
Published: (2025)
Corporate Governance Reforms in India: Evolution and Observance
by: Amanpreet Kaur
Published: (2026)
by: Amanpreet Kaur
Published: (2026)
Cross Culturalism in Bharati Mukherjee's Jasmine
by: Amanpreet Kaur
Published: (2018)
by: Amanpreet Kaur
Published: (2018)
Intelligent AI Delegation
by: Tomašev, Nenad, et al.
Published: (2026)
by: Tomašev, Nenad, et al.
Published: (2026)
TRACE: Capability-Targeted Agentic Training
by: Kang, Hangoo, et al.
Published: (2026)
by: Kang, Hangoo, et al.
Published: (2026)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
Increasing the seed production efficiency of autumn potato with plant growth regulators
by: Amanpreet Singh, et al.
Published: (2024)
by: Amanpreet Singh, et al.
Published: (2024)
DouweGeurtjens/unfolding-alignments: v1.0
by: DouweGeurtjens
Published: (2025)
by: DouweGeurtjens
Published: (2025)
Perfect Worlds
by: Fokkema, Douwe
Published: (2011)
by: Fokkema, Douwe
Published: (2011)
Similar Items
-
Reflective Context Learning: Studying the Optimization Primitives of Context Space
by: Vassilyev, Nikita, et al.
Published: (2026) -
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
by: D'Oosterlinck, Karel, et al.
Published: (2024) -
Anchor Points: Benchmarking Models with Much Fewer Examples
by: Vivek, Rajan, et al.
Published: (2023) -
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
by: Tran, Dat, et al.
Published: (2026) -
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
by: Lui, Nicholas, et al.
Published: (2023)