GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Qiming, Chen, Zichen, Corcoran, Will, Sra, Misha, Singh, Ambuj K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving
von: Luyten, Max Ruiz, et al.
Veröffentlicht: (2026)
von: Luyten, Max Ruiz, et al.
Veröffentlicht: (2026)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
von: Fang, Qitong, et al.
Veröffentlicht: (2026)
von: Fang, Qitong, et al.
Veröffentlicht: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
von: Steele, Brady, et al.
Veröffentlicht: (2026)
von: Steele, Brady, et al.
Veröffentlicht: (2026)
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
von: Koc, Vincent
Veröffentlicht: (2025)
von: Koc, Vincent
Veröffentlicht: (2025)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
von: Ghandi, Taraneh, et al.
Veröffentlicht: (2026)
von: Ghandi, Taraneh, et al.
Veröffentlicht: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
von: Raman, Vishal, et al.
Veröffentlicht: (2025)
von: Raman, Vishal, et al.
Veröffentlicht: (2025)
On the Limits of Learned Importance Scoring for KV Cache Compression
von: Steele, Brady
Veröffentlicht: (2026)
von: Steele, Brady
Veröffentlicht: (2026)
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
von: Cui, Jian, et al.
Veröffentlicht: (2026)
von: Cui, Jian, et al.
Veröffentlicht: (2026)
Can AI Assist in Olympiad Coding
von: Ren, Samuel
Veröffentlicht: (2025)
von: Ren, Samuel
Veröffentlicht: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
von: Zolduoarrati, Elijah, et al.
Veröffentlicht: (2025)
von: Zolduoarrati, Elijah, et al.
Veröffentlicht: (2025)
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
von: Palacios, Diego Cabezas
Veröffentlicht: (2026)
von: Palacios, Diego Cabezas
Veröffentlicht: (2026)
From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
von: Gill, Gurbinder, et al.
Veröffentlicht: (2025)
von: Gill, Gurbinder, et al.
Veröffentlicht: (2025)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
von: He, Langzhou, et al.
Veröffentlicht: (2026)
von: He, Langzhou, et al.
Veröffentlicht: (2026)
Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
von: Xu, Jiexi
Veröffentlicht: (2025)
von: Xu, Jiexi
Veröffentlicht: (2025)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
Dynamics of COVID-19 Misinformation: An Analysis of Conspiracy Theories, Fake Remedies, and False Reports
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2025)
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2025)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
von: Li, Xu, et al.
Veröffentlicht: (2026)
von: Li, Xu, et al.
Veröffentlicht: (2026)
ChatGPT4PCG 2 Competition: Prompt Engineering for Science Birds Level Generation
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2024)
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2024)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
von: Tanjim, Md Mehrab, et al.
Veröffentlicht: (2026)
von: Tanjim, Md Mehrab, et al.
Veröffentlicht: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
von: Yousaf, Iqra
Veröffentlicht: (2024)
von: Yousaf, Iqra
Veröffentlicht: (2024)
CoupleEvo: Evolving Heuristics for Coupled Optimization Problems Using Large Language Models
von: Bömer, Thomas, et al.
Veröffentlicht: (2026)
von: Bömer, Thomas, et al.
Veröffentlicht: (2026)
Open-TI: Open Traffic Intelligence with Augmented Language Model
von: Da, Longchao, et al.
Veröffentlicht: (2023)
von: Da, Longchao, et al.
Veröffentlicht: (2023)
Systematic Classification of Studies Investigating Social Media Conversations about Long COVID Using a Novel Zero-Shot Transformer Framework
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2025)
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2025)
Critical Insights into Leading Conversational AI Models
von: Kohli, Urja, et al.
Veröffentlicht: (2025)
von: Kohli, Urja, et al.
Veröffentlicht: (2025)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
von: Hildebrand, Samuel, et al.
Veröffentlicht: (2025)
von: Hildebrand, Samuel, et al.
Veröffentlicht: (2025)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
von: Ansell, Rebecca, et al.
Veröffentlicht: (2026)
von: Ansell, Rebecca, et al.
Veröffentlicht: (2026)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
von: Jang, Yoonjin, et al.
Veröffentlicht: (2026)
von: Jang, Yoonjin, et al.
Veröffentlicht: (2026)
Deployment-Time Reliability of Learned Robot Policies
von: Agia, Christopher
Veröffentlicht: (2026)
von: Agia, Christopher
Veröffentlicht: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
von: Iscan, Mehmet
Veröffentlicht: (2026)
von: Iscan, Mehmet
Veröffentlicht: (2026)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
von: Agrawal, Lakshya A, et al.
Veröffentlicht: (2025)
von: Agrawal, Lakshya A, et al.
Veröffentlicht: (2025)
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
von: Thakur, Nirmalya
Veröffentlicht: (2024)
von: Thakur, Nirmalya
Veröffentlicht: (2024)
A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and Other Sources about the 2024 Outbreak of Measles
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2024)
von: Thakur, Nirmalya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
von: Zhou, Yue, et al.
Veröffentlicht: (2025) -
The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving
von: Luyten, Max Ruiz, et al.
Veröffentlicht: (2026) -
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
von: Fang, Qitong, et al.
Veröffentlicht: (2026) -
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
von: Steele, Brady, et al.
Veröffentlicht: (2026) -
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
von: Koc, Vincent
Veröffentlicht: (2025)