Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations
Fuente:
arXiv
Guardado en:
| Autores principales: | Tomar, Aditya, Sahoo, Nihar Ranjan, Mittal, Ashish, Murthy, Rudra, Bhattacharyya, Pushpak |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
por: Tomar, Aditya, et al.
Publicado: (2025)
por: Tomar, Aditya, et al.
Publicado: (2025)
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
por: Tomar, Aditya, et al.
Publicado: (2025)
por: Tomar, Aditya, et al.
Publicado: (2025)
Evaluating Dialect Robustness of Language Models via Conversation Understanding
por: Srirag, Dipankar, et al.
Publicado: (2024)
por: Srirag, Dipankar, et al.
Publicado: (2024)
IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian Context
por: Sahoo, Nihar Ranjan, et al.
Publicado: (2024)
por: Sahoo, Nihar Ranjan, et al.
Publicado: (2024)
Inverse Scaling: When Bigger Isn't Better
por: McKenzie, Ian R., et al.
Publicado: (2023)
por: McKenzie, Ian R., et al.
Publicado: (2023)
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities
por: Sravanthi, Settaluri Lakshmi, et al.
Publicado: (2024)
por: Sravanthi, Settaluri Lakshmi, et al.
Publicado: (2024)
Word Boundary Information Isn't Useful for Encoder Language Models
por: Gow-Smith, Edward, et al.
Publicado: (2024)
por: Gow-Smith, Edward, et al.
Publicado: (2024)
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
por: Ghosh, Poulami, et al.
Publicado: (2024)
por: Ghosh, Poulami, et al.
Publicado: (2024)
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
por: Long, Zhuohan, et al.
Publicado: (2026)
por: Long, Zhuohan, et al.
Publicado: (2026)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
por: Barkett, Emilio, et al.
Publicado: (2025)
por: Barkett, Emilio, et al.
Publicado: (2025)
Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
por: Kim, Sewon, et al.
Publicado: (2025)
por: Kim, Sewon, et al.
Publicado: (2025)
When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
por: Das, Sarmistha, et al.
Publicado: (2026)
por: Das, Sarmistha, et al.
Publicado: (2026)
Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents
por: Wang, Hanlin, et al.
Publicado: (2026)
por: Wang, Hanlin, et al.
Publicado: (2026)
Expect the unexpected: Harnessing Sentence Completion for Sarcasm Detection
por: Joshi, Aditya, et al.
Publicado: (2017)
por: Joshi, Aditya, et al.
Publicado: (2017)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
por: Loem, Mengsay, et al.
Publicado: (2023)
por: Loem, Mengsay, et al.
Publicado: (2023)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
por: Khan, Mohammed Safi Ur Rahman, et al.
Publicado: (2026)
por: Khan, Mohammed Safi Ur Rahman, et al.
Publicado: (2026)
The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
por: Menta, Venkata Pushpak Teja
Publicado: (2026)
por: Menta, Venkata Pushpak Teja
Publicado: (2026)
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
por: Barale, Claire, et al.
Publicado: (2025)
por: Barale, Claire, et al.
Publicado: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
por: Xu, Xiaoyu, et al.
Publicado: (2025)
por: Xu, Xiaoyu, et al.
Publicado: (2025)
P-ReMIS: Pragmatic Reasoning in Mental Health and a Social Implication
por: Oram, Sneha, et al.
Publicado: (2025)
por: Oram, Sneha, et al.
Publicado: (2025)
Main Predicate and Their Arguments as Explanation Signals For Intent Classification
por: Pimparkhede, Sameer, et al.
Publicado: (2025)
por: Pimparkhede, Sameer, et al.
Publicado: (2025)
Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair
por: Yousofi, Waisullah, et al.
Publicado: (2024)
por: Yousofi, Waisullah, et al.
Publicado: (2024)
Facts-and-Feelings: Capturing both Objectivity and Subjectivity in Table-to-Text Generation
por: Dey, Tathagata, et al.
Publicado: (2024)
por: Dey, Tathagata, et al.
Publicado: (2024)
We Care: Multimodal Depression Detection and Knowledge Infused Mental Health Therapeutic Response Generation
por: Moon, Palash, et al.
Publicado: (2024)
por: Moon, Palash, et al.
Publicado: (2024)
Recall Isn't Enough: Bounding Commitments in Personalized Language Systems
por: Tang, Rui, et al.
Publicado: (2026)
por: Tang, Rui, et al.
Publicado: (2026)
Ta-G-T: Subjectivity Capture in Table to Text Generation via RDF Graphs
por: Upasham, Ronak, et al.
Publicado: (2025)
por: Upasham, Ronak, et al.
Publicado: (2025)
Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
por: Wynn, Andrea, et al.
Publicado: (2025)
por: Wynn, Andrea, et al.
Publicado: (2025)
When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation
por: Huang, Nannan, et al.
Publicado: (2026)
por: Huang, Nannan, et al.
Publicado: (2026)
Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
por: Tsipidi, Eleftheria, et al.
Publicado: (2024)
por: Tsipidi, Eleftheria, et al.
Publicado: (2024)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
por: Kashid, Harshvivek, et al.
Publicado: (2024)
por: Kashid, Harshvivek, et al.
Publicado: (2024)
Open-DeBias: Toward Mitigating Open-Set Bias in Language Models
por: Rani, Arti, et al.
Publicado: (2025)
por: Rani, Arti, et al.
Publicado: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
por: Zhang, Yue, et al.
Publicado: (2026)
por: Zhang, Yue, et al.
Publicado: (2026)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
por: Devanathan, Rishikesh, et al.
Publicado: (2025)
por: Devanathan, Rishikesh, et al.
Publicado: (2025)
Enhancing Food-Domain Question Answering with a Multimodal Knowledge Graph: Hybrid QA Generation and Diversity Analysis
por: B, Srihari K, et al.
Publicado: (2025)
por: B, Srihari K, et al.
Publicado: (2025)
Striking a Balance between Classical and Deep Learning Approaches in Natural Language Processing Pedagogy
por: Joshi, Aditya, et al.
Publicado: (2024)
por: Joshi, Aditya, et al.
Publicado: (2024)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
por: Galeone, Cosimo, et al.
Publicado: (2026)
por: Galeone, Cosimo, et al.
Publicado: (2026)
StereoDetect: Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological Underpinnings
por: Shejole, Kaustubh Shivshankar, et al.
Publicado: (2025)
por: Shejole, Kaustubh Shivshankar, et al.
Publicado: (2025)
Giving the Old a Fresh Spin: Quality Estimation-Assisted Constrained Decoding for Automatic Post-Editing
por: Deoghare, Sourabh, et al.
Publicado: (2025)
por: Deoghare, Sourabh, et al.
Publicado: (2025)
Pretraining Language Models Using Translationese
por: Doshi, Meet, et al.
Publicado: (2024)
por: Doshi, Meet, et al.
Publicado: (2024)
Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics
por: Kumar, Sanjeev, et al.
Publicado: (2026)
por: Kumar, Sanjeev, et al.
Publicado: (2026)
Ejemplares similares
-
BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
por: Tomar, Aditya, et al.
Publicado: (2025) -
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
por: Tomar, Aditya, et al.
Publicado: (2025) -
Evaluating Dialect Robustness of Language Models via Conversation Understanding
por: Srirag, Dipankar, et al.
Publicado: (2024) -
IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian Context
por: Sahoo, Nihar Ranjan, et al.
Publicado: (2024) -
Inverse Scaling: When Bigger Isn't Better
por: McKenzie, Ian R., et al.
Publicado: (2023)