Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
Fuente:
arXiv
Salvato in:
| Autori principali: | Raju, Ravi, Jain, Swayambhoo, Li, Bo, Li, Jonathan, Thakker, Urmish |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
di: Hong, Fenglu, et al.
Pubblicazione: (2025)
di: Hong, Fenglu, et al.
Pubblicazione: (2025)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
di: Jain, Swayambhoo, et al.
Pubblicazione: (2024)
di: Jain, Swayambhoo, et al.
Pubblicazione: (2024)
The Limits of Long-Context Reasoning in Automated Bug Fixing
di: Raju, Ravi, et al.
Pubblicazione: (2026)
di: Raju, Ravi, et al.
Pubblicazione: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
SambaLingo: Teaching Large Language Models New Languages
di: Csaki, Zoltan, et al.
Pubblicazione: (2024)
di: Csaki, Zoltan, et al.
Pubblicazione: (2024)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
di: Zhao, Xueliang, et al.
Pubblicazione: (2024)
di: Zhao, Xueliang, et al.
Pubblicazione: (2024)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
di: Li, Dawei, et al.
Pubblicazione: (2025)
di: Li, Dawei, et al.
Pubblicazione: (2025)
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
di: Fathullah, Yassir, et al.
Pubblicazione: (2025)
di: Fathullah, Yassir, et al.
Pubblicazione: (2025)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
di: Zhang, Qizheng, et al.
Pubblicazione: (2025)
di: Zhang, Qizheng, et al.
Pubblicazione: (2025)
Implicit Federated In-context Learning For Task-Specific LLM Fine-Tuning
di: Li, Dongcheng, et al.
Pubblicazione: (2025)
di: Li, Dongcheng, et al.
Pubblicazione: (2025)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
Evaluating Differentially Private Generation of Domain-Specific Text
di: Sun, Yidan, et al.
Pubblicazione: (2025)
di: Sun, Yidan, et al.
Pubblicazione: (2025)
Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
di: Chatterjee, Abhijit, et al.
Pubblicazione: (2025)
di: Chatterjee, Abhijit, et al.
Pubblicazione: (2025)
Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain
di: Okochi, Yuma, et al.
Pubblicazione: (2026)
di: Okochi, Yuma, et al.
Pubblicazione: (2026)
Causal-Driven Feature Evaluation for Cross-Domain Image Classification
di: Cheng, Chen, et al.
Pubblicazione: (2026)
di: Cheng, Chen, et al.
Pubblicazione: (2026)
Evaluating Fairness and Mitigating Bias in Machine Learning: A Novel Technique using Tensor Data and Bayesian Regression
di: Paxton, Kuniko, et al.
Pubblicazione: (2025)
di: Paxton, Kuniko, et al.
Pubblicazione: (2025)
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
di: Ostapenko, Oleksiy, et al.
Pubblicazione: (2025)
di: Ostapenko, Oleksiy, et al.
Pubblicazione: (2025)
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
di: Jain, Purvam, et al.
Pubblicazione: (2026)
di: Jain, Purvam, et al.
Pubblicazione: (2026)
Guiding Exploration in Reinforcement Learning Through LLM-Augmented Observations
di: Jain, Vaibhav, et al.
Pubblicazione: (2025)
di: Jain, Vaibhav, et al.
Pubblicazione: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
di: Mohammadi, Mahmoud, et al.
Pubblicazione: (2025)
di: Mohammadi, Mahmoud, et al.
Pubblicazione: (2025)
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
di: Xia, Xiaojie, et al.
Pubblicazione: (2026)
di: Xia, Xiaojie, et al.
Pubblicazione: (2026)
Online Domain-aware LLM Decoding for Continual Domain Evolution
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2026)
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2026)
LLM-Assisted Logic Rule Learning: Scaling Human Expertise for Time Series Anomaly Detection
di: Zhang, Haoting, et al.
Pubblicazione: (2026)
di: Zhang, Haoting, et al.
Pubblicazione: (2026)
Gains: Fine-grained Federated Domain Adaptation in Open Set
di: Zhong, Zhengyi, et al.
Pubblicazione: (2025)
di: Zhong, Zhengyi, et al.
Pubblicazione: (2025)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
di: Wang, Yutong, et al.
Pubblicazione: (2025)
di: Wang, Yutong, et al.
Pubblicazione: (2025)
Graph Edit Distance with General Costs Using Neural Set Divergence
di: Jain, Eeshaan, et al.
Pubblicazione: (2024)
di: Jain, Eeshaan, et al.
Pubblicazione: (2024)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
di: Sun, Lihao, et al.
Pubblicazione: (2026)
di: Sun, Lihao, et al.
Pubblicazione: (2026)
Predicting Survival of Hemodialysis Patients using Federated Learning
di: Raju, Abhiram, et al.
Pubblicazione: (2024)
di: Raju, Abhiram, et al.
Pubblicazione: (2024)
AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment
di: Fu, Yonggan, et al.
Pubblicazione: (2024)
di: Fu, Yonggan, et al.
Pubblicazione: (2024)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
di: Chen, Rubing, et al.
Pubblicazione: (2025)
di: Chen, Rubing, et al.
Pubblicazione: (2025)
Towards Efficient Pareto Set Approximation via Mixture of Experts Based Model Fusion
di: Tang, Anke, et al.
Pubblicazione: (2024)
di: Tang, Anke, et al.
Pubblicazione: (2024)
From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation
di: Wang, Yongqi, et al.
Pubblicazione: (2026)
di: Wang, Yongqi, et al.
Pubblicazione: (2026)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
di: Li, Xiaochuan, et al.
Pubblicazione: (2025)
di: Li, Xiaochuan, et al.
Pubblicazione: (2025)
Refinement Provenance Inference: Detecting LLM-Refined Training Prompts from Model Behavior
di: Yin, Bo, et al.
Pubblicazione: (2026)
di: Yin, Bo, et al.
Pubblicazione: (2026)
Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
di: Rawal, Kaivalya, et al.
Pubblicazione: (2026)
di: Rawal, Kaivalya, et al.
Pubblicazione: (2026)
FD-LLM: Large Language Model for Fault Diagnosis of Machines
di: Qaid, Hamzah A. A. M., et al.
Pubblicazione: (2024)
di: Qaid, Hamzah A. A. M., et al.
Pubblicazione: (2024)
Data Trajectory Alignment for LLM Domain Adaptation: A Two-Phase Synthesis Framework for Telecommunications Mathematics
di: Zhou, Zhicheng, et al.
Pubblicazione: (2025)
di: Zhou, Zhicheng, et al.
Pubblicazione: (2025)
MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task Learning
di: Ma, Yufei, et al.
Pubblicazione: (2024)
di: Ma, Yufei, et al.
Pubblicazione: (2024)
Thinking in Different Spaces: Domain-Specific Latent Geometry Survives Cross-Architecture Translation
di: Armstrong, Marcus, et al.
Pubblicazione: (2026)
di: Armstrong, Marcus, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
di: Hong, Fenglu, et al.
Pubblicazione: (2025) -
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
di: Jain, Swayambhoo, et al.
Pubblicazione: (2024) -
The Limits of Long-Context Reasoning in Automated Bug Fixing
di: Raju, Ravi, et al.
Pubblicazione: (2026) -
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025) -
SambaLingo: Teaching Large Language Models New Languages
di: Csaki, Zoltan, et al.
Pubblicazione: (2024)