LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Yanan, Salem, Ahmed, Nushi, Besmira, Russinovich, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking is (Mostly) Simpler Than You Think
by: Russinovich, Mark, et al.
Published: (2025)
by: Russinovich, Mark, et al.
Published: (2025)
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
by: Russinovich, Mark, et al.
Published: (2026)
by: Russinovich, Mark, et al.
Published: (2026)
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
by: Russinovich, Mark, et al.
Published: (2025)
by: Russinovich, Mark, et al.
Published: (2025)
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)
by: Naik, Ranjita, et al.
Published: (2023)
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
by: Pfister, Rolf, et al.
Published: (2025)
by: Pfister, Rolf, et al.
Published: (2025)
Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
by: Vilas, Martina G., et al.
Published: (2025)
by: Vilas, Martina G., et al.
Published: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
by: Bi, Zhen, et al.
Published: (2026)
by: Bi, Zhen, et al.
Published: (2026)
Reliability by design: quantifying and eliminating fabrication risk in LLMs. From generative to consultative AI: a comparative analysis in the legal domain and lessons for high-stakes knowledge bases
by: Dantart, Alex
Published: (2026)
by: Dantart, Alex
Published: (2026)
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
by: Bolet, Gregory, et al.
Published: (2025)
by: Bolet, Gregory, et al.
Published: (2025)
Detecting Data Contamination in LLMs via In-Context Learning
by: Zawalski, Michał, et al.
Published: (2025)
by: Zawalski, Michał, et al.
Published: (2025)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
by: Pinnock, Alyssa, et al.
Published: (2025)
by: Pinnock, Alyssa, et al.
Published: (2025)
MoEITS: A Green AI approach for simplifying MoE-LLMs
by: Balderas, Luis, et al.
Published: (2026)
by: Balderas, Luis, et al.
Published: (2026)
Prompting for Performance: Exploring LLMs for Configuring Software
by: Spieker, Helge, et al.
Published: (2025)
by: Spieker, Helge, et al.
Published: (2025)
LLMs for Analog Circuit Design Continuum (ACDC)
by: Esfandiari, Yasaman, et al.
Published: (2025)
by: Esfandiari, Yasaman, et al.
Published: (2025)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
by: Hourri, Younes, et al.
Published: (2025)
by: Hourri, Younes, et al.
Published: (2025)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
by: Mozaffari, Mohammad, et al.
Published: (2025)
by: Mozaffari, Mohammad, et al.
Published: (2025)
This Is Taking Too Long -- Investigating Time as a Proxy for Energy Consumption of LLMs
by: Krupp, Lars, et al.
Published: (2026)
by: Krupp, Lars, et al.
Published: (2026)
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
XTC, A Research Platform for Optimizing AI Workload Operators
by: Hugo, Pompougnac, et al.
Published: (2025)
by: Hugo, Pompougnac, et al.
Published: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
by: Liu, Jiashuo, et al.
Published: (2025)
by: Liu, Jiashuo, et al.
Published: (2025)
Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
by: Zhao, Youpeng, et al.
Published: (2025)
by: Zhao, Youpeng, et al.
Published: (2025)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
by: Rani, Pooja, et al.
Published: (2025)
by: Rani, Pooja, et al.
Published: (2025)
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
by: Atinafu, Yonas, et al.
Published: (2026)
by: Atinafu, Yonas, et al.
Published: (2026)
DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning Workloads
by: Zhao, Qidong, et al.
Published: (2024)
by: Zhao, Qidong, et al.
Published: (2024)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
by: Arantes, Gabriel M., et al.
Published: (2025)
by: Arantes, Gabriel M., et al.
Published: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
by: Ivanov, Andrei, et al.
Published: (2025)
by: Ivanov, Andrei, et al.
Published: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
by: Prieto, Pablo, et al.
Published: (2025)
by: Prieto, Pablo, et al.
Published: (2025)
PixLift: Accelerating Web Browsing via AI Upscaling
by: Atinafu, Yonas, et al.
Published: (2025)
by: Atinafu, Yonas, et al.
Published: (2025)
Tiny-QMoE
by: Cashman, Jack, et al.
Published: (2025)
by: Cashman, Jack, et al.
Published: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
by: Chai, Yuji, et al.
Published: (2025)
by: Chai, Yuji, et al.
Published: (2025)
WANDER: An Explainable Decision-Support Framework for HPC
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
Similar Items
-
Jailbreaking is (Mostly) Simpler Than You Think
by: Russinovich, Mark, et al.
Published: (2025) -
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
by: Russinovich, Mark, et al.
Published: (2024) -
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
by: Russinovich, Mark, et al.
Published: (2026) -
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
by: Russinovich, Mark, et al.
Published: (2025) -
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)