Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming
Fuente:
arXiv
Saved in:
| Main Authors: | Hadhoud, Sama, Elsetohy, Alaa, Hudi, Frederikus, Cruz, Jan Christian Blaise, Halim, Steven, Aji, Alham Fikri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
by: Wibowo, Haryo Akbarianto, et al.
Published: (2026)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2026)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025)
by: Hudi, Frederikus, et al.
Published: (2025)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
Sense Representations Are Inducible Interfaces
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation
by: Cruz, Jan Christian Blaise, et al.
Published: (2025)
by: Cruz, Jan Christian Blaise, et al.
Published: (2025)
Multilinguality as Sense Adaptation
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
Khattat: Enhancing Readability and Concept Representation of Semantic Typography
by: Hussein, Ahmed, et al.
Published: (2024)
by: Hussein, Ahmed, et al.
Published: (2024)
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
by: Aji, Alham Fikri, et al.
Published: (2025)
by: Aji, Alham Fikri, et al.
Published: (2025)
Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning
by: Attia, Ahmed, et al.
Published: (2026)
by: Attia, Ahmed, et al.
Published: (2026)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
by: Chevi, Rendi, et al.
Published: (2024)
by: Chevi, Rendi, et al.
Published: (2024)
Beyond Probabilities: Unveiling the Misalignment in Evaluating Large Language Models
by: Lyu, Chenyang, et al.
Published: (2024)
by: Lyu, Chenyang, et al.
Published: (2024)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
by: Imam, Mohamed Fazli, et al.
Published: (2025)
by: Imam, Mohamed Fazli, et al.
Published: (2025)
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
by: Mansurov, Jonibek, et al.
Published: (2024)
by: Mansurov, Jonibek, et al.
Published: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Sense
by: Cahyawijaya, Samuel, et al.
Published: (2024)
by: Cahyawijaya, Samuel, et al.
Published: (2024)
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
Language-Specific Latent Process Hinders Cross-Lingual Performance
by: Lim, Zheng Wei, et al.
Published: (2025)
by: Lim, Zheng Wei, et al.
Published: (2025)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
Efficient and Interpretable Grammatical Error Correction with Mixture of Experts
by: Qorib, Muhammad Reza, et al.
Published: (2024)
by: Qorib, Muhammad Reza, et al.
Published: (2024)
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning
by: Elshabrawy, Ahmed, et al.
Published: (2024)
by: Elshabrawy, Ahmed, et al.
Published: (2024)
How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
by: Chevi, Rendi, et al.
Published: (2025)
by: Chevi, Rendi, et al.
Published: (2025)
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
by: Tasawong, Panuthep, et al.
Published: (2025)
by: Tasawong, Panuthep, et al.
Published: (2025)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
by: Nur'aini, Khumaisa, et al.
Published: (2026)
by: Nur'aini, Khumaisa, et al.
Published: (2026)
Predicting the Order of Upcoming Tokens Improves Language Modeling
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
StructLens: A Structural Lens for Language Models via Maximum Spanning Trees
by: Sakajo, Haruki, et al.
Published: (2026)
by: Sakajo, Haruki, et al.
Published: (2026)
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
by: Wu, Minghao, et al.
Published: (2023)
by: Wu, Minghao, et al.
Published: (2023)
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
by: Tasawong, Panuthep, et al.
Published: (2026)
by: Tasawong, Panuthep, et al.
Published: (2026)
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024)
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024)
MapCoder: Multi-Agent Code Generation for Competitive Problem Solving
by: Islam, Md. Ashraful, et al.
Published: (2024)
by: Islam, Md. Ashraful, et al.
Published: (2024)
PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2026)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2026)
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
by: Villa-Cueva, Emilio, et al.
Published: (2025)
by: Villa-Cueva, Emilio, et al.
Published: (2025)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
QLESS: A Quantized Approach for Data Valuation and Selection in Large Language Model Fine-Tuning
by: Ananta, Moses, et al.
Published: (2025)
by: Ananta, Moses, et al.
Published: (2025)
Similar Items
-
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026) -
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
by: Wibowo, Haryo Akbarianto, et al.
Published: (2026) -
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025) -
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
by: Cruz, Jan Christian Blaise, et al.
Published: (2026) -
Sense Representations Are Inducible Interfaces
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)