Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boughorbel, Sabri, Dalvi, Fahim, Durrani, Nadir, Hawasly, Majd
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911171765862400
author Boughorbel, Sabri
Dalvi, Fahim
Durrani, Nadir
Hawasly, Majd
author_facet Boughorbel, Sabri
Dalvi, Fahim
Durrani, Nadir
Hawasly, Majd
contents As fine-tuning becomes the dominant paradigm for improving large language models (LLMs), understanding what changes during this process is increasingly important. Traditional benchmarking often fails to explain why one model outperforms another. In this work, we use model diffing, a mechanistic interpretability approach, to analyze the specific capability differences between Gemma-2-9b-it and a SimPO-enhanced variant. Using crosscoders, we identify and categorize latent representations that differentiate the two models. We find that SimPO acquired latent concepts predominantly enhance safety mechanisms (+32.8%), multilingual capabilities (+43.8%), and instruction-following (+151.7%), while its additional training also reduces emphasis on model self-reference (-44.1%) and hallucination management (-68.5%). Our analysis shows that model diffing can yield fine-grained insights beyond leaderboard metrics, attributing performance gaps to concrete mechanistic capabilities. This approach offers a transparent and targeted framework for comparing LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18792
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
Boughorbel, Sabri
Dalvi, Fahim
Durrani, Nadir
Hawasly, Majd
Computation and Language
As fine-tuning becomes the dominant paradigm for improving large language models (LLMs), understanding what changes during this process is increasingly important. Traditional benchmarking often fails to explain why one model outperforms another. In this work, we use model diffing, a mechanistic interpretability approach, to analyze the specific capability differences between Gemma-2-9b-it and a SimPO-enhanced variant. Using crosscoders, we identify and categorize latent representations that differentiate the two models. We find that SimPO acquired latent concepts predominantly enhance safety mechanisms (+32.8%), multilingual capabilities (+43.8%), and instruction-following (+151.7%), while its additional training also reduces emphasis on model self-reference (-44.1%) and hallucination management (-68.5%). Our analysis shows that model diffing can yield fine-grained insights beyond leaderboard metrics, attributing performance gaps to concrete mechanistic capabilities. This approach offers a transparent and targeted framework for comparing LLMs.
title Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
topic Computation and Language
url https://arxiv.org/abs/2509.18792