MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts
Fuente:
arXiv
Guardado en:
| Autores principales: | Lanzeray, Etienne, Meilliez, Stephane, Ruelle, Malo, Sileo, Damien |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
por: Sileo, Damien
Publicado: (2025)
por: Sileo, Damien
Publicado: (2025)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
por: Sileo, Damien
Publicado: (2024)
por: Sileo, Damien
Publicado: (2024)
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
por: Sileo, Damien
Publicado: (2024)
por: Sileo, Damien
Publicado: (2024)
Saturation-Driven Dataset Generation for LLM Mathematical Reasoning in the TPTP Ecosystem
por: Quesnel, Valentin, et al.
Publicado: (2025)
por: Quesnel, Valentin, et al.
Publicado: (2025)
Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
por: Lacombe, Valentin, et al.
Publicado: (2025)
por: Lacombe, Valentin, et al.
Publicado: (2025)
Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training
por: Lacombe, Valentin, et al.
Publicado: (2026)
por: Lacombe, Valentin, et al.
Publicado: (2026)
Tau-Eval: A Unified Evaluation Framework for Useful and Private Text Anonymization
por: Loiseau, Gabriel, et al.
Publicado: (2025)
por: Loiseau, Gabriel, et al.
Publicado: (2025)
Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization
por: Loiseau, Gabriel, et al.
Publicado: (2026)
por: Loiseau, Gabriel, et al.
Publicado: (2026)
Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
por: Loiseau, Gabriel, et al.
Publicado: (2026)
por: Loiseau, Gabriel, et al.
Publicado: (2026)
TAROT: Task-Oriented Authorship Obfuscation Using Policy Optimization Methods
por: Loiseau, Gabriel, et al.
Publicado: (2024)
por: Loiseau, Gabriel, et al.
Publicado: (2024)
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
por: Sun, Kai, et al.
Publicado: (2024)
por: Sun, Kai, et al.
Publicado: (2024)
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
por: Wang, Ke, et al.
Publicado: (2024)
por: Wang, Ke, et al.
Publicado: (2024)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
por: Chernyshev, Konstantin, et al.
Publicado: (2024)
por: Chernyshev, Konstantin, et al.
Publicado: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
por: Huang, Kaixuan, et al.
Publicado: (2025)
por: Huang, Kaixuan, et al.
Publicado: (2025)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
por: Saha, Soumadeep, et al.
Publicado: (2025)
por: Saha, Soumadeep, et al.
Publicado: (2025)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
por: Teixeira, Tiago, et al.
Publicado: (2026)
por: Teixeira, Tiago, et al.
Publicado: (2026)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
por: Gao, Bofei, et al.
Publicado: (2024)
por: Gao, Bofei, et al.
Publicado: (2024)
Reward-free Alignment for Conflicting Objectives
por: Chen, Peter, et al.
Publicado: (2026)
por: Chen, Peter, et al.
Publicado: (2026)
InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating
por: Wang, Fuyu, et al.
Publicado: (2025)
por: Wang, Fuyu, et al.
Publicado: (2025)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
por: Wang, Yiming, et al.
Publicado: (2025)
por: Wang, Yiming, et al.
Publicado: (2025)
Evaluating Multilingual Long-Context Models for Retrieval and Reasoning
por: Agrawal, Ameeta, et al.
Publicado: (2024)
por: Agrawal, Ameeta, et al.
Publicado: (2024)
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
por: Meadows, Jordan, et al.
Publicado: (2023)
por: Meadows, Jordan, et al.
Publicado: (2023)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
por: Nachshoni, Eviatar, et al.
Publicado: (2025)
por: Nachshoni, Eviatar, et al.
Publicado: (2025)
Recipient Profiling: Predicting Characteristics from Messages
por: Borquez, Martin, et al.
Publicado: (2024)
por: Borquez, Martin, et al.
Publicado: (2024)
A Compute-Matched Re-Evaluation of TroVE on MATH
por: Sesterhenn, Tobias, et al.
Publicado: (2025)
por: Sesterhenn, Tobias, et al.
Publicado: (2025)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
por: Petty, Jackson, et al.
Publicado: (2025)
por: Petty, Jackson, et al.
Publicado: (2025)
DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
por: Xu, Zhe, et al.
Publicado: (2024)
por: Xu, Zhe, et al.
Publicado: (2024)
Open Domain Question Answering with Conflicting Contexts
por: Liu, Siyi, et al.
Publicado: (2024)
por: Liu, Siyi, et al.
Publicado: (2024)
Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation
por: Chandra, Arjun, et al.
Publicado: (2026)
por: Chandra, Arjun, et al.
Publicado: (2026)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
por: Zhao, Weixiang, et al.
Publicado: (2026)
por: Zhao, Weixiang, et al.
Publicado: (2026)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
por: Bertsch, Amanda, et al.
Publicado: (2025)
por: Bertsch, Amanda, et al.
Publicado: (2025)
On the Emergence of Induction Heads for In-Context Learning
por: Musat, Tiberiu, et al.
Publicado: (2025)
por: Musat, Tiberiu, et al.
Publicado: (2025)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
por: Li, Moxin, et al.
Publicado: (2025)
por: Li, Moxin, et al.
Publicado: (2025)
ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
por: Mehandru, Nikita, et al.
Publicado: (2025)
por: Mehandru, Nikita, et al.
Publicado: (2025)
Lost in the Middle, and In-Between: Enhancing Language Models' Ability to Reason Over Long Contexts in Multi-Hop QA
por: Baker, George Arthur, et al.
Publicado: (2024)
por: Baker, George Arthur, et al.
Publicado: (2024)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
por: Su, Zhaochen, et al.
Publicado: (2024)
por: Su, Zhaochen, et al.
Publicado: (2024)
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
por: Ling, Zhan, et al.
Publicado: (2025)
por: Ling, Zhan, et al.
Publicado: (2025)
ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts
por: Noh, Dongwon, et al.
Publicado: (2025)
por: Noh, Dongwon, et al.
Publicado: (2025)
Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflicts
por: Tighidet, Zineddine, et al.
Publicado: (2025)
por: Tighidet, Zineddine, et al.
Publicado: (2025)
FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation
por: Zhang, Qinggang, et al.
Publicado: (2025)
por: Zhang, Qinggang, et al.
Publicado: (2025)
Ejemplares similares
-
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
por: Sileo, Damien
Publicado: (2025) -
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
por: Sileo, Damien
Publicado: (2024) -
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
por: Sileo, Damien
Publicado: (2024) -
Saturation-Driven Dataset Generation for LLM Mathematical Reasoning in the TPTP Ecosystem
por: Quesnel, Valentin, et al.
Publicado: (2025) -
Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning
por: Lacombe, Valentin, et al.
Publicado: (2025)