LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hashemi, Helia, Eisner, Jason, Rosset, Corby, Van Durme, Benjamin, Kedzie, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dodo: Dynamic Contextual Compression for Decoder-only LMs
von: Qin, Guanghui, et al.
Veröffentlicht: (2023)
von: Qin, Guanghui, et al.
Veröffentlicht: (2023)
BLT: Can Large Language Models Handle Basic Legal Text?
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2023)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2023)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
von: Sauter, Andreas, et al.
Veröffentlicht: (2026)
von: Sauter, Andreas, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
von: Hartnett, Peter, et al.
Veröffentlicht: (2026)
von: Hartnett, Peter, et al.
Veröffentlicht: (2026)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
von: Mullens, Drake, et al.
Veröffentlicht: (2026)
von: Mullens, Drake, et al.
Veröffentlicht: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
von: Liao, Jianxing, et al.
Veröffentlicht: (2025)
von: Liao, Jianxing, et al.
Veröffentlicht: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
von: Kugler, Kai
Veröffentlicht: (2025)
von: Kugler, Kai
Veröffentlicht: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
von: Voss, Lukas
Veröffentlicht: (2026)
von: Voss, Lukas
Veröffentlicht: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
von: Ye, Hua, et al.
Veröffentlicht: (2025)
von: Ye, Hua, et al.
Veröffentlicht: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
von: Patarlapalli, Sai Babu, et al.
Veröffentlicht: (2026)
von: Patarlapalli, Sai Babu, et al.
Veröffentlicht: (2026)
Can LLMs Identify Tax Abuse?
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
On the Influence of Discourse Relations in Persuasive Texts
von: Turk, Nawar, et al.
Veröffentlicht: (2025)
von: Turk, Nawar, et al.
Veröffentlicht: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
Latent Cache Flow: Model-to-Model Communication Without Text
von: Rossi, Maximillian, et al.
Veröffentlicht: (2026)
von: Rossi, Maximillian, et al.
Veröffentlicht: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
Automated Circuit Interpretation via Probe Prompting
von: Birardi, Giuseppe
Veröffentlicht: (2025)
von: Birardi, Giuseppe
Veröffentlicht: (2025)
When Reasoning Fails: Evaluating 'Thinking' LLMs for Stock Prediction
von: Sodha, Rakeshkumar H
Veröffentlicht: (2025)
von: Sodha, Rakeshkumar H
Veröffentlicht: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
von: Zhang, Xue
Veröffentlicht: (2025)
von: Zhang, Xue
Veröffentlicht: (2025)
Towards Conditioning Clinical Text Generation for User Control
von: Koraş, Osman Alperen, et al.
Veröffentlicht: (2025)
von: Koraş, Osman Alperen, et al.
Veröffentlicht: (2025)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
Cost-Aware Model Selection for Text Classification: Multi-Objective Trade-offs Between Fine-Tuned Encoders and LLM Prompting in Production
von: Gonzalez, Alberto Andres Valdes
Veröffentlicht: (2026)
von: Gonzalez, Alberto Andres Valdes
Veröffentlicht: (2026)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
von: Zhang, Luyan, et al.
Veröffentlicht: (2025)
von: Zhang, Luyan, et al.
Veröffentlicht: (2025)
A Graph-based RAG for Energy Efficiency Question Answering
von: Campi, Riccardo, et al.
Veröffentlicht: (2025)
von: Campi, Riccardo, et al.
Veröffentlicht: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
von: Du, Bangde, et al.
Veröffentlicht: (2025)
von: Du, Bangde, et al.
Veröffentlicht: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
Generalization of Graph Neural Network Models for Distribution Grid Fault Detection
von: Karabulut, Burak, et al.
Veröffentlicht: (2025)
von: Karabulut, Burak, et al.
Veröffentlicht: (2025)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning
von: Luijkx, Jelle, et al.
Veröffentlicht: (2025)
von: Luijkx, Jelle, et al.
Veröffentlicht: (2025)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
KAConvText: Novel Approach to Burmese Sentence Classification using Kolmogorov-Arnold Convolution
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
von: Idahl, Maximilian, et al.
Veröffentlicht: (2026)
von: Idahl, Maximilian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dodo: Dynamic Contextual Compression for Decoder-only LMs
von: Qin, Guanghui, et al.
Veröffentlicht: (2023) -
BLT: Can Large Language Models Handle Basic Legal Text?
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2023) -
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025) -
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
von: Sauter, Andreas, et al.
Veröffentlicht: (2026) -
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
von: Hartnett, Peter, et al.
Veröffentlicht: (2026)