MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Epstein, Elliot L., Yao, Kaisheng, Li, Jing, Bai, Xinyi, Palangi, Hamid |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Set-Sequence Model for Time Series
por: Epstein, Elliot L., et al.
Publicado: (2025)
por: Epstein, Elliot L., et al.
Publicado: (2025)
LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
por: Epstein, Elliot L., et al.
Publicado: (2025)
por: Epstein, Elliot L., et al.
Publicado: (2025)
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
por: Zhang, Gongbo, et al.
Publicado: (2026)
por: Zhang, Gongbo, et al.
Publicado: (2026)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
por: Manikantan, Kawshik, et al.
Publicado: (2024)
por: Manikantan, Kawshik, et al.
Publicado: (2024)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
por: Adapala, Sai Teja Reddy
Publicado: (2025)
por: Adapala, Sai Teja Reddy
Publicado: (2025)
Attention Factors for Statistical Arbitrage
por: Epstein, Elliot L., et al.
Publicado: (2025)
por: Epstein, Elliot L., et al.
Publicado: (2025)
Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking
por: Epstein, Elliot L., et al.
Publicado: (2026)
por: Epstein, Elliot L., et al.
Publicado: (2026)
Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
por: Zhang, Liangliang, et al.
Publicado: (2025)
por: Zhang, Liangliang, et al.
Publicado: (2025)
Eureka: Evaluating and Understanding Large Foundation Models
por: Balachandran, Vidhisha, et al.
Publicado: (2024)
por: Balachandran, Vidhisha, et al.
Publicado: (2024)
LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
por: Niklaus, Joel, et al.
Publicado: (2023)
por: Niklaus, Joel, et al.
Publicado: (2023)
EasyMath: A 0-shot Math Benchmark for SLMs
por: Karki, Drishya, et al.
Publicado: (2025)
por: Karki, Drishya, et al.
Publicado: (2025)
BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, Rerankers and LLMs
por: Aarab, Ilias
Publicado: (2026)
por: Aarab, Ilias
Publicado: (2026)
Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities
por: Dai, Qirun, et al.
Publicado: (2025)
por: Dai, Qirun, et al.
Publicado: (2025)
ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series
por: Niu, Shuai, et al.
Publicado: (2025)
por: Niu, Shuai, et al.
Publicado: (2025)
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
por: González-Pizarro, Felipe, et al.
Publicado: (2024)
por: González-Pizarro, Felipe, et al.
Publicado: (2024)
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
por: Yin, Yuwei, et al.
Publicado: (2026)
por: Yin, Yuwei, et al.
Publicado: (2026)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
por: Koo, Ryan, et al.
Publicado: (2023)
por: Koo, Ryan, et al.
Publicado: (2023)
Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
por: Guo, Xu, et al.
Publicado: (2024)
por: Guo, Xu, et al.
Publicado: (2024)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
por: Zhang, Chuyifei, et al.
Publicado: (2026)
por: Zhang, Chuyifei, et al.
Publicado: (2026)
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
por: Bandarkar, Lucas, et al.
Publicado: (2023)
por: Bandarkar, Lucas, et al.
Publicado: (2023)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
por: Kurtic, Eldar, et al.
Publicado: (2024)
por: Kurtic, Eldar, et al.
Publicado: (2024)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
por: Vieira, Inês, et al.
Publicado: (2026)
por: Vieira, Inês, et al.
Publicado: (2026)
A Llama walks into the 'Bar': Efficient Supervised Fine-Tuning for Legal Reasoning in the Multi-state Bar Exam
por: Fernandes, Rean, et al.
Publicado: (2025)
por: Fernandes, Rean, et al.
Publicado: (2025)
GMoE: Empowering LLMs Fine-Tuning via MoE Graph Collaboration
por: Bai, Ting, et al.
Publicado: (2024)
por: Bai, Ting, et al.
Publicado: (2024)
What Should Embeddings Embed? Autoregressive Models Represent Latent Generating Distributions
por: Zhang, Liyi, et al.
Publicado: (2024)
por: Zhang, Liyi, et al.
Publicado: (2024)
Flash Multi-Head Feed-Forward Network
por: Zhang, Minshen, et al.
Publicado: (2025)
por: Zhang, Minshen, et al.
Publicado: (2025)
SocraSynth: Multi-LLM Reasoning with Conditional Statistics
por: Chang, Edward Y.
Publicado: (2024)
por: Chang, Edward Y.
Publicado: (2024)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
por: Niklaus, Joel, et al.
Publicado: (2025)
por: Niklaus, Joel, et al.
Publicado: (2025)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
por: Fan, Yu, et al.
Publicado: (2025)
por: Fan, Yu, et al.
Publicado: (2025)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
por: Ghosh, Shubhra, et al.
Publicado: (2025)
por: Ghosh, Shubhra, et al.
Publicado: (2025)
Dealing with Annotator Disagreement in Hate Speech Classification
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
por: Hong, Chunsan, et al.
Publicado: (2025)
por: Hong, Chunsan, et al.
Publicado: (2025)
The Efficiency Attenuation Phenomenon: A Computational Challenge to the Language of Thought Hypothesis
por: Zhang, Di
Publicado: (2026)
por: Zhang, Di
Publicado: (2026)
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification
por: Sirbu, Iustin, et al.
Publicado: (2025)
por: Sirbu, Iustin, et al.
Publicado: (2025)
Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing
por: Rohanian, Omid, et al.
Publicado: (2023)
por: Rohanian, Omid, et al.
Publicado: (2023)
Modality Interactive Mixture-of-Experts for Fake News Detection
por: Liu, Yifan, et al.
Publicado: (2025)
por: Liu, Yifan, et al.
Publicado: (2025)
Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression
por: Baxi, Rahul
Publicado: (2025)
por: Baxi, Rahul
Publicado: (2025)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
por: Abramov, Roman, et al.
Publicado: (2025)
por: Abramov, Roman, et al.
Publicado: (2025)
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
por: Stern, Ronja, et al.
Publicado: (2023)
por: Stern, Ronja, et al.
Publicado: (2023)
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset
por: S, Santosh T. Y. S., et al.
Publicado: (2024)
por: S, Santosh T. Y. S., et al.
Publicado: (2024)
Ejemplares similares
-
A Set-Sequence Model for Time Series
por: Epstein, Elliot L., et al.
Publicado: (2025) -
LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
por: Epstein, Elliot L., et al.
Publicado: (2025) -
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
por: Zhang, Gongbo, et al.
Publicado: (2026) -
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
por: Manikantan, Kawshik, et al.
Publicado: (2024) -
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
por: Adapala, Sai Teja Reddy
Publicado: (2025)