Measuring What Matters Beyond Text: Evaluating Multimodal Summaries by Quality, Alignment, and Diversity
Fuente:
arXiv
Saved in:
| Main Authors: | Ali, Abid, Molla-Aliod, Diego, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention
by: Ali, Abid, et al.
Published: (2026)
by: Ali, Abid, et al.
Published: (2026)
Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
by: Alawwad, Hessa, et al.
Published: (2025)
by: Alawwad, Hessa, et al.
Published: (2025)
Synthetic Dialogue Dataset Generation using LLM Agents
by: Abdullin, Yelaman, et al.
Published: (2024)
by: Abdullin, Yelaman, et al.
Published: (2024)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
by: Nadeem, Afrozah, et al.
Published: (2026)
by: Nadeem, Afrozah, et al.
Published: (2026)
Fairness Evaluation and Inference Level Mitigation in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering
by: Alawwad, Hessa A., et al.
Published: (2025)
by: Alawwad, Hessa A., et al.
Published: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
by: Shetty, Anudeex, et al.
Published: (2025)
by: Shetty, Anudeex, et al.
Published: (2025)
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
by: Mustafa, Akram, et al.
Published: (2025)
by: Mustafa, Akram, et al.
Published: (2025)
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
by: Desikan, Prasanna, et al.
Published: (2026)
by: Desikan, Prasanna, et al.
Published: (2026)
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
by: Almutairi, Ali, et al.
Published: (2025)
by: Almutairi, Ali, et al.
Published: (2025)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
by: Zhong, Jiayou, et al.
Published: (2025)
by: Zhong, Jiayou, et al.
Published: (2025)
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
by: Islam, Muhammad, et al.
Published: (2025)
by: Islam, Muhammad, et al.
Published: (2025)
Fair Summarization: Bridging Quality and Diversity in Extractive Summaries
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
by: Nezhad, Sina Bagheri, et al.
Published: (2024)
LLM Ensemble for RAG: Role of Context Length in Zero-Shot Question Answering for BioASQ Challenge
by: Galat, Dima, et al.
Published: (2025)
by: Galat, Dima, et al.
Published: (2025)
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
by: Zhang, Zhihao, et al.
Published: (2026)
by: Zhang, Zhihao, et al.
Published: (2026)
Steering Towards Fairness: Mitigating Political Bias in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
by: Zheng, Shenyan, et al.
Published: (2026)
by: Zheng, Shenyan, et al.
Published: (2026)
What Matters For Safety Alignment?
by: Li, Xing, et al.
Published: (2026)
by: Li, Xing, et al.
Published: (2026)
Can Reasoning LLMs Enhance Clinical Document Classification?
by: Mustafa, Akram, et al.
Published: (2025)
by: Mustafa, Akram, et al.
Published: (2025)
Measuring What Matters: The AI Pluralism Index
by: Mushkani, Rashid
Published: (2025)
by: Mushkani, Rashid
Published: (2025)
Measuring What Matters: Intrinsic Distance Preservation as a Robust Metric for Embedding Quality
by: Hart, Steven N., et al.
Published: (2024)
by: Hart, Steven N., et al.
Published: (2024)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
Beyond One-Size-Fits-All Summarization: Customizing Summaries for Diverse Users
by: Duran, Mehmet Samet, et al.
Published: (2025)
by: Duran, Mehmet Samet, et al.
Published: (2025)
Leveraging Taxonomy and LLMs for Improved Multimodal Hierarchical Classification
by: Chen, Shijing, et al.
Published: (2025)
by: Chen, Shijing, et al.
Published: (2025)
MRGAgents: A Multi-Agent Framework for Improved Medical Report Generation with Med-LVLMs
by: Wang, Pengyu, et al.
Published: (2025)
by: Wang, Pengyu, et al.
Published: (2025)
MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
by: Wang, Pengyu, et al.
Published: (2025)
by: Wang, Pengyu, et al.
Published: (2025)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
by: Bhandari, Pranav, et al.
Published: (2026)
by: Bhandari, Pranav, et al.
Published: (2026)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
by: Gero, Zelalem, et al.
Published: (2024)
by: Gero, Zelalem, et al.
Published: (2024)
Competing LLM Agents in a Non-Cooperative Game of Opinion Polarisation
by: Qasmi, Amin, et al.
Published: (2025)
by: Qasmi, Amin, et al.
Published: (2025)
Pay Attention to What Matters
by: Silva, Pedro Luiz, et al.
Published: (2024)
by: Silva, Pedro Luiz, et al.
Published: (2024)
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires
by: Bhandari, Pranav, et al.
Published: (2025)
by: Bhandari, Pranav, et al.
Published: (2025)
What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge
by: Shin, Yosub, et al.
Published: (2026)
by: Shin, Yosub, et al.
Published: (2026)
Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment
by: Wang, Haoyuan, et al.
Published: (2026)
by: Wang, Haoyuan, et al.
Published: (2026)
Geometry-Aware Semantic Reasoning for Training Free Video Anomaly Detection
by: Zia, Ali, et al.
Published: (2026)
by: Zia, Ali, et al.
Published: (2026)
Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
by: Galatolo, Alessio, et al.
Published: (2025)
by: Galatolo, Alessio, et al.
Published: (2025)
Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up
by: Yuan, Jiahao, et al.
Published: (2024)
by: Yuan, Jiahao, et al.
Published: (2024)
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
Similar Items
-
Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention
by: Ali, Abid, et al.
Published: (2026) -
Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
by: Alawwad, Hessa, et al.
Published: (2025) -
Synthetic Dialogue Dataset Generation using LLM Agents
by: Abdullin, Yelaman, et al.
Published: (2024) -
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
by: Nadeem, Afrozah, et al.
Published: (2026) -
Fairness Evaluation and Inference Level Mitigation in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)