No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Permadi, Vynska Amalia, Tan, Xingwei, Moosavi, Nafise Sadat, Aletras, Nikos |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiHoax: A Dataset of Multi-hop False-Premise Questions
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025)
by: Xue, Huiyin, et al.
Published: (2025)
How to Leverage Digit Embeddings to Represent Numbers?
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026)
by: Kennedy, Ian W., et al.
Published: (2026)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
by: Liu, Yiqi, et al.
Published: (2023)
by: Liu, Yiqi, et al.
Published: (2023)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
by: Mi, Maggie, et al.
Published: (2024)
by: Mi, Maggie, et al.
Published: (2024)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
by: Mi, Maggie, et al.
Published: (2025)
by: Mi, Maggie, et al.
Published: (2025)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
by: Pandya, Mugdha, et al.
Published: (2024)
by: Pandya, Mugdha, et al.
Published: (2024)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
by: Pastorino, Valeria, et al.
Published: (2024)
by: Pastorino, Valeria, et al.
Published: (2024)
Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
by: Alajrami, Ahmed, et al.
Published: (2025)
by: Alajrami, Ahmed, et al.
Published: (2025)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
Multi-hop Question Answering
by: Mavi, Vaibhav, et al.
Published: (2022)
by: Mavi, Vaibhav, et al.
Published: (2022)
MultiCube-RAG for Multi-hop Question Answering
by: Shi, Jimeng, et al.
Published: (2026)
by: Shi, Jimeng, et al.
Published: (2026)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026)
by: James, Joseph, et al.
Published: (2026)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
by: Saffari, Hamidreza, et al.
Published: (2024)
by: Saffari, Hamidreza, et al.
Published: (2024)
Where does output diversity collapse in post-training?
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
Exploring Gender Disparities in Automatic Speech Recognition Technology
by: ElGhazaly, Hend, et al.
Published: (2025)
by: ElGhazaly, Hend, et al.
Published: (2025)
Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks
by: Barati, Iman, et al.
Published: (2025)
by: Barati, Iman, et al.
Published: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
by: Lewis-Lim, Samuel, et al.
Published: (2025)
by: Lewis-Lim, Samuel, et al.
Published: (2025)
PokeMQA: Programmable knowledge editing for Multi-hop Question Answering
by: Gu, Hengrui, et al.
Published: (2023)
by: Gu, Hengrui, et al.
Published: (2023)
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
by: Lee, Jihyung, et al.
Published: (2025)
by: Lee, Jihyung, et al.
Published: (2025)
Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering
by: Chu, Zheng, et al.
Published: (2025)
by: Chu, Zheng, et al.
Published: (2025)
AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering
by: Wang, Yuxin, et al.
Published: (2026)
by: Wang, Yuxin, et al.
Published: (2026)
TRACE: An Experiential Framework for Coherent Multi-hop Knowledge Graph Question Answering
by: Wang, Yingxu, et al.
Published: (2026)
by: Wang, Yingxu, et al.
Published: (2026)
Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision
by: Tan, Xingwei, et al.
Published: (2025)
by: Tan, Xingwei, et al.
Published: (2025)
Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
by: Lewis-Lim, Samuel, et al.
Published: (2025)
by: Lewis-Lim, Samuel, et al.
Published: (2025)
Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering
by: Shi, Zhengliang, et al.
Published: (2024)
by: Shi, Zhengliang, et al.
Published: (2024)
Multi-hop Question Answering under Temporal Knowledge Editing
by: Cheng, Keyuan, et al.
Published: (2024)
by: Cheng, Keyuan, et al.
Published: (2024)
Afri-MCQA: Multimodal Cultural Question Answering for African Languages
by: Tonja, Atnafu Lambebo, et al.
Published: (2026)
by: Tonja, Atnafu Lambebo, et al.
Published: (2026)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering
by: Sun, Duolin, et al.
Published: (2025)
by: Sun, Duolin, et al.
Published: (2025)
HOLMES: Hyper-Relational Knowledge Graphs for Multi-hop Question Answering using LLMs
by: Panda, Pranoy, et al.
Published: (2024)
by: Panda, Pranoy, et al.
Published: (2024)
MQA-KEAL: Multi-hop Question Answering under Knowledge Editing for Arabic Language
by: Ali, Muhammad Asif, et al.
Published: (2024)
by: Ali, Muhammad Asif, et al.
Published: (2024)
RISE: Reasoning Enhancement via Iterative Self-Exploration in Multi-hop Question Answering
by: He, Bolei, et al.
Published: (2025)
by: He, Bolei, et al.
Published: (2025)
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
by: Lim, Qi Zhi, et al.
Published: (2025)
by: Lim, Qi Zhi, et al.
Published: (2025)
CIRAG: Construction-Integration Retrieval and Adaptive Generation for Multi-hop Question Answering
by: Wei, Zili, et al.
Published: (2026)
by: Wei, Zili, et al.
Published: (2026)
Similar Items
-
MultiHoax: A Dataset of Multi-hop False-Premise Questions
by: Shafiei, Mohammadamin, et al.
Published: (2025) -
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025) -
How to Leverage Digit Embeddings to Represent Numbers?
by: Sivakumar, Jasivan Alex, et al.
Published: (2024) -
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026) -
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
by: Liu, Yiqi, et al.
Published: (2023)