AlgoSimBench: Identifying Algorithmically Similar Problems for Competitive Programming
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jierui, Mooney, Raymond |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distilling Algorithmic Reasoning from LLMs via Explaining Solution Programs
by: Li, Jierui, et al.
Published: (2024)
by: Li, Jierui, et al.
Published: (2024)
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
by: Cohen, Vanya, et al.
Published: (2025)
by: Cohen, Vanya, et al.
Published: (2025)
Problem-Solving Guide: Predicting the Algorithm Tags and Difficulty for Competitive Programming Problems
by: Kim, Juntae, et al.
Published: (2023)
by: Kim, Juntae, et al.
Published: (2023)
ProBench: Benchmarking Large Language Models in Competitive Programming
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
Natural Language Can Help Bridge the Sim2Real Gap
by: Yu, Albert, et al.
Published: (2024)
by: Yu, Albert, et al.
Published: (2024)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
by: Li, Jierui, et al.
Published: (2023)
by: Li, Jierui, et al.
Published: (2023)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization
by: Sun, Yan, et al.
Published: (2026)
by: Sun, Yan, et al.
Published: (2026)
FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction
by: Johnson, Natasha, et al.
Published: (2025)
by: Johnson, Natasha, et al.
Published: (2025)
SimSMoE: Solving Representational Collapse via Similarity Measure
by: Do, Giang, et al.
Published: (2024)
by: Do, Giang, et al.
Published: (2024)
mPLM-Sim: Better Cross-Lingual Similarity and Transfer in Multilingual Pretrained Language Models
by: Lin, Peiqin, et al.
Published: (2023)
by: Lin, Peiqin, et al.
Published: (2023)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
FMC: Formalization of Natural Language Mathematical Competition Problems
by: Xie, Jiaxuan, et al.
Published: (2025)
by: Xie, Jiaxuan, et al.
Published: (2025)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming
by: Hadhoud, Sama, et al.
Published: (2026)
by: Hadhoud, Sama, et al.
Published: (2026)
Learning Task Decomposition to Assist Humans in Competitive Programming
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
by: Zheng, Zihan, et al.
Published: (2025)
by: Zheng, Zihan, et al.
Published: (2025)
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
by: Tsoukalas, George, et al.
Published: (2024)
by: Tsoukalas, George, et al.
Published: (2024)
NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
by: Song, Dinghong, et al.
Published: (2025)
by: Song, Dinghong, et al.
Published: (2025)
CaT-BENCH: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans
by: Lal, Yash Kumar, et al.
Published: (2024)
by: Lal, Yash Kumar, et al.
Published: (2024)
Multimodal Contextualized Semantic Parsing from Speech
by: Voas, Jordan, et al.
Published: (2024)
by: Voas, Jordan, et al.
Published: (2024)
Sampling the Swadesh List to Identify Similar Languages with Tree Spaces
by: Ordway, Garett, et al.
Published: (2024)
by: Ordway, Garett, et al.
Published: (2024)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming
by: Wang, Peisong, et al.
Published: (2026)
by: Wang, Peisong, et al.
Published: (2026)
AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?
by: Press, Ori, et al.
Published: (2025)
by: Press, Ori, et al.
Published: (2025)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
by: An, Shengnan, et al.
Published: (2025)
by: An, Shengnan, et al.
Published: (2025)
SimGRAG: Leveraging Similar Subgraphs for Knowledge Graphs Driven Retrieval-Augmented Generation
by: Cai, Yuzheng, et al.
Published: (2024)
by: Cai, Yuzheng, et al.
Published: (2024)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
by: Cao, Yuhan, et al.
Published: (2025)
by: Cao, Yuhan, et al.
Published: (2025)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
by: Li, Jierui, et al.
Published: (2024)
by: Li, Jierui, et al.
Published: (2024)
Competition-Level Problems are Effective LLM Evaluators
by: Huang, Yiming, et al.
Published: (2023)
by: Huang, Yiming, et al.
Published: (2023)
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
by: Xu, Shiyi, et al.
Published: (2025)
by: Xu, Shiyi, et al.
Published: (2025)
BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Codehacks: A Dataset of Adversarial Tests for Competitive Programming Problems Obtained from Codeforces
by: Hort, Max, et al.
Published: (2025)
by: Hort, Max, et al.
Published: (2025)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
LLMLagBench: Identifying Temporal Training Boundaries in Large Language Models
by: Pęzik, Piotr, et al.
Published: (2025)
by: Pęzik, Piotr, et al.
Published: (2025)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
Competitive Programming with Large Reasoning Models
by: OpenAI, et al.
Published: (2025)
by: OpenAI, et al.
Published: (2025)
Similar Items
-
Distilling Algorithmic Reasoning from LLMs via Explaining Solution Programs
by: Li, Jierui, et al.
Published: (2024) -
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
by: Cohen, Vanya, et al.
Published: (2025) -
Problem-Solving Guide: Predicting the Algorithm Tags and Difficulty for Competitive Programming Problems
by: Kim, Juntae, et al.
Published: (2023) -
ProBench: Benchmarking Large Language Models in Competitive Programming
by: Yang, Lei, et al.
Published: (2025) -
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)