Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
Fuente:
arXiv
Saved in:
| Main Authors: | Wibowo, Haryo Akbarianto, Elsetohy, Alaa, Cui, Qinrong, Aji, Alham Fikri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
IteRABRe: Iterative Recovery-Aided Block Reduction
by: Wibowo, Haryo Akbarianto, et al.
Published: (2025)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2025)
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
Unveiling the Influence of Amplifying Language-Specific Neurons
by: Rahmanisa, Inaya, et al.
Published: (2025)
by: Rahmanisa, Inaya, et al.
Published: (2025)
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming
by: Hadhoud, Sama, et al.
Published: (2026)
by: Hadhoud, Sama, et al.
Published: (2026)
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
by: Aji, Alham Fikri, et al.
Published: (2025)
by: Aji, Alham Fikri, et al.
Published: (2025)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
by: Li, Mingda, et al.
Published: (2024)
by: Li, Mingda, et al.
Published: (2024)
Duluth at SemEval-2025 Task 7: TF-IDF with Optimized Vector Dimensions for Multilingual Fact-Checked Claim Retrieval
by: Syed, Shujauddin, et al.
Published: (2025)
by: Syed, Shujauddin, et al.
Published: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations
by: Bhattacharya, Shamik, et al.
Published: (2026)
by: Bhattacharya, Shamik, et al.
Published: (2026)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
by: Alhamzeh, Alaa, et al.
Published: (2025)
by: Alhamzeh, Alaa, et al.
Published: (2025)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
by: Gómez-Rodríguez, Carlos, et al.
Published: (2024)
by: Gómez-Rodríguez, Carlos, et al.
Published: (2024)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024)
by: Liu, Linyu, et al.
Published: (2024)
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
by: Marcuzzo, Matteo, et al.
Published: (2025)
by: Marcuzzo, Matteo, et al.
Published: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
by: Yao, Huaiyuan, et al.
Published: (2026)
by: Yao, Huaiyuan, et al.
Published: (2026)
Meaning-infused grammar: Gradient Acceptability Shapes the Geometric Representations of Constructions in LLMs
by: Rakshit, Supantho, et al.
Published: (2025)
by: Rakshit, Supantho, et al.
Published: (2025)
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
by: Harris, Joshua, et al.
Published: (2025)
by: Harris, Joshua, et al.
Published: (2025)
One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
by: Wang, Pengbo, et al.
Published: (2025)
by: Wang, Pengbo, et al.
Published: (2025)
BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
by: Johnson, Warren
Published: (2026)
by: Johnson, Warren
Published: (2026)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024)
by: Land, Sander, et al.
Published: (2024)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
by: Ikoma, Hayato, et al.
Published: (2025)
by: Ikoma, Hayato, et al.
Published: (2025)
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
by: Li, Yongchang, et al.
Published: (2024)
by: Li, Yongchang, et al.
Published: (2024)
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
by: Alkan, Atilla Kaan, et al.
Published: (2025)
by: Alkan, Atilla Kaan, et al.
Published: (2025)
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2024)
by: Goldin, Gili, et al.
Published: (2024)
Can LLMs Compute with Reasons?
by: Sandilya, Harshit, et al.
Published: (2024)
by: Sandilya, Harshit, et al.
Published: (2024)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
by: Yoo, Seunghyun
Published: (2025)
by: Yoo, Seunghyun
Published: (2025)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Multilinguality as Sense Adaptation
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
LLMs for Legal Subsumption in German Employment Contracts
by: Wardas, Oliver, et al.
Published: (2025)
by: Wardas, Oliver, et al.
Published: (2025)
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
by: Balter, Samuel G., et al.
Published: (2026)
by: Balter, Samuel G., et al.
Published: (2026)
Similar Items
-
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024) -
IteRABRe: Iterative Recovery-Aided Block Reduction
by: Wibowo, Haryo Akbarianto, et al.
Published: (2025) -
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023) -
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
by: Andrylie, Lyzander Marciano, et al.
Published: (2025) -
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)