Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kerkouri, Mohamed Amine, Tliba, Marouane, Chetouani, Aladine, Aburaed, Nour, Bruno, Alessandro
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918034763939840
author Kerkouri, Mohamed Amine
Tliba, Marouane
Chetouani, Aladine
Aburaed, Nour
Bruno, Alessandro
author_facet Kerkouri, Mohamed Amine
Tliba, Marouane
Chetouani, Aladine
Aburaed, Nour
Bruno, Alessandro
contents This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments to a single scalar, obscuring semantic failures, user intent, and the rationale behind quality decisions. We contend that modern quality assessment models must integrate three interdependent capabilities: (1) context-awareness, to adapt evaluations to task-specific goals and viewing conditions; (2) reasoning, to produce interpretable, evidence-grounded justifications for quality judgments; and (3) multimodality, to align perceptual and semantic cues using vision-language models. We critique the limitations of current MOS-centric benchmarks and propose a roadmap for reform: richer datasets with contextual metadata and expert rationales, and new evaluation metrics that assess semantic alignment, reasoning fidelity, and contextual sensitivity. By reframing quality assessment as a contextual, explainable, and multimodal modeling task, we aim to catalyze a shift toward more robust, human-aligned, and trustworthy evaluation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19696
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
Kerkouri, Mohamed Amine
Tliba, Marouane
Chetouani, Aladine
Aburaed, Nour
Bruno, Alessandro
Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments to a single scalar, obscuring semantic failures, user intent, and the rationale behind quality decisions. We contend that modern quality assessment models must integrate three interdependent capabilities: (1) context-awareness, to adapt evaluations to task-specific goals and viewing conditions; (2) reasoning, to produce interpretable, evidence-grounded justifications for quality judgments; and (3) multimodality, to align perceptual and semantic cues using vision-language models. We critique the limitations of current MOS-centric benchmarks and propose a roadmap for reform: richer datasets with contextual metadata and expert rationales, and new evaluation metrics that assess semantic alignment, reasoning fidelity, and contextual sensitivity. By reframing quality assessment as a contextual, explainable, and multimodal modeling task, we aim to catalyze a shift toward more robust, human-aligned, and trustworthy evaluation systems.
title Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
topic Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
url https://arxiv.org/abs/2505.19696