LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Mondorf, Philipp, Bell, Samuel J., Dodge, Jesse, Hupkes, Dieuwke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
by: Roberts, Nicholas, et al.
Published: (2025)
by: Roberts, Nicholas, et al.
Published: (2025)
Quantifying Variance in Evaluation Benchmarks
by: Madaan, Lovish, et al.
Published: (2024)
by: Madaan, Lovish, et al.
Published: (2024)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
by: Hupkes, Dieuwke, et al.
Published: (2025)
by: Hupkes, Dieuwke, et al.
Published: (2025)
Brittlebench: Quantifying LLM robustness via prompt sensitivity
by: Romanou, Angelika, et al.
Published: (2026)
by: Romanou, Angelika, et al.
Published: (2026)
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale
by: Thomas, Noel
Published: (2026)
by: Thomas, Noel
Published: (2026)
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
by: Ohmer, Xenia, et al.
Published: (2024)
by: Ohmer, Xenia, et al.
Published: (2024)
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
by: Eichin, Florian, et al.
Published: (2025)
by: Eichin, Florian, et al.
Published: (2025)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Reason to Rote: Rethinking Memorization in Reasoning
by: Du, Yupei, et al.
Published: (2025)
by: Du, Yupei, et al.
Published: (2025)
Holistically Evaluating the Environmental Impact of Creating Language Models
by: Morrison, Jacob, et al.
Published: (2025)
by: Morrison, Jacob, et al.
Published: (2025)
LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
Difficulties with Evaluating a Deception Detector for AIs
by: Smith, Lewis, et al.
Published: (2025)
by: Smith, Lewis, et al.
Published: (2025)
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
by: Fairoze, Jaiden, et al.
Published: (2025)
by: Fairoze, Jaiden, et al.
Published: (2025)
Fake News Detection After LLM Laundering: Measurement and Explanation
by: Das, Rupak Kumar, et al.
Published: (2025)
by: Das, Rupak Kumar, et al.
Published: (2025)
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment
by: Li, Cheryl, et al.
Published: (2025)
by: Li, Cheryl, et al.
Published: (2025)
Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
by: Heineman, David, et al.
Published: (2025)
by: Heineman, David, et al.
Published: (2025)
Evaluating Game Difficulty in Tetris Block Puzzle
by: Wang, Chun-Jui, et al.
Published: (2026)
by: Wang, Chun-Jui, et al.
Published: (2026)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Lower Difficulty and Better Robustness: A Bregman Divergence Perspective for Adversarial Training
by: Wu, Zihui, et al.
Published: (2022)
by: Wu, Zihui, et al.
Published: (2022)
The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
Reassessing the Validity of Spurious Correlations Benchmarks
by: Bell, Samuel J., et al.
Published: (2024)
by: Bell, Samuel J., et al.
Published: (2024)
Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs
by: Yadav, Suraj, et al.
Published: (2026)
by: Yadav, Suraj, et al.
Published: (2026)
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
by: Luo, Qin-Wen, et al.
Published: (2026)
by: Luo, Qin-Wen, et al.
Published: (2026)
Interpretability of Language Models via Task Spaces
by: Weber, Lucas, et al.
Published: (2024)
by: Weber, Lucas, et al.
Published: (2024)
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
Beyond Scaling Curves: Internal Dynamics of Neural Networks Through the NTK Lens
by: Nikolaou, Konstantin, et al.
Published: (2025)
by: Nikolaou, Konstantin, et al.
Published: (2025)
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
by: Chen, Tianhao, et al.
Published: (2025)
by: Chen, Tianhao, et al.
Published: (2025)
Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
by: Jandali, Yaman, et al.
Published: (2025)
by: Jandali, Yaman, et al.
Published: (2025)
Circuit Transformer: A Transformer That Preserves Logical Equivalence
by: Li, Xihan, et al.
Published: (2024)
by: Li, Xihan, et al.
Published: (2024)
Scaling Up Diffusion and Flow-based XGBoost Models
by: Cresswell, Jesse C., et al.
Published: (2024)
by: Cresswell, Jesse C., et al.
Published: (2024)
Identity-Link IRT for Label-Free LLM Evaluation: Preserving Additivity in TVD-MI Scores
by: Robertson, Zachary
Published: (2025)
by: Robertson, Zachary
Published: (2025)
Robust Mitigation of Age-Dependent Confounding Effects via Sample-Difficulty Decorrelation
by: Kurian, Nikhil Cherian, et al.
Published: (2026)
by: Kurian, Nikhil Cherian, et al.
Published: (2026)
D3: Diversity, Difficulty, and Dependability-Aware Data Selection for Sample-Efficient LLM Instruction Tuning
by: Zhang, Jia, et al.
Published: (2025)
by: Zhang, Jia, et al.
Published: (2025)
Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies
by: Schneider, Chris, et al.
Published: (2026)
by: Schneider, Chris, et al.
Published: (2026)
Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation
by: Kricheli, Joshua Shay, et al.
Published: (2026)
by: Kricheli, Joshua Shay, et al.
Published: (2026)
LLM-Assisted Logic Rule Learning: Scaling Human Expertise for Time Series Anomaly Detection
by: Zhang, Haoting, et al.
Published: (2026)
by: Zhang, Haoting, et al.
Published: (2026)
Similar Items
-
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
by: Roberts, Nicholas, et al.
Published: (2025) -
Quantifying Variance in Evaluation Benchmarks
by: Madaan, Lovish, et al.
Published: (2024) -
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
by: Mondorf, Philipp, et al.
Published: (2024) -
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
by: Hupkes, Dieuwke, et al.
Published: (2025) -
Brittlebench: Quantifying LLM robustness via prompt sensitivity
by: Romanou, Angelika, et al.
Published: (2026)