LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
Fuente:
arXiv
Guardado en:
| Autores principales: | Mondorf, Philipp, Bell, Samuel J., Dodge, Jesse, Hupkes, Dieuwke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
por: Roberts, Nicholas, et al.
Publicado: (2025)
por: Roberts, Nicholas, et al.
Publicado: (2025)
Quantifying Variance in Evaluation Benchmarks
por: Madaan, Lovish, et al.
Publicado: (2024)
por: Madaan, Lovish, et al.
Publicado: (2024)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
por: Hupkes, Dieuwke, et al.
Publicado: (2025)
por: Hupkes, Dieuwke, et al.
Publicado: (2025)
Brittlebench: Quantifying LLM robustness via prompt sensitivity
por: Romanou, Angelika, et al.
Publicado: (2026)
por: Romanou, Angelika, et al.
Publicado: (2026)
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale
por: Thomas, Noel
Publicado: (2026)
por: Thomas, Noel
Publicado: (2026)
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
por: Ohmer, Xenia, et al.
Publicado: (2024)
por: Ohmer, Xenia, et al.
Publicado: (2024)
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
por: Eichin, Florian, et al.
Publicado: (2025)
por: Eichin, Florian, et al.
Publicado: (2025)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Reason to Rote: Rethinking Memorization in Reasoning
por: Du, Yupei, et al.
Publicado: (2025)
por: Du, Yupei, et al.
Publicado: (2025)
Holistically Evaluating the Environmental Impact of Creating Language Models
por: Morrison, Jacob, et al.
Publicado: (2025)
por: Morrison, Jacob, et al.
Publicado: (2025)
LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion
por: Long, Yunbo, et al.
Publicado: (2025)
por: Long, Yunbo, et al.
Publicado: (2025)
Difficulties with Evaluating a Deception Detector for AIs
por: Smith, Lewis, et al.
Publicado: (2025)
por: Smith, Lewis, et al.
Publicado: (2025)
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
por: Fairoze, Jaiden, et al.
Publicado: (2025)
por: Fairoze, Jaiden, et al.
Publicado: (2025)
Fake News Detection After LLM Laundering: Measurement and Explanation
por: Das, Rupak Kumar, et al.
Publicado: (2025)
por: Das, Rupak Kumar, et al.
Publicado: (2025)
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment
por: Li, Cheryl, et al.
Publicado: (2025)
por: Li, Cheryl, et al.
Publicado: (2025)
Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
por: Heineman, David, et al.
Publicado: (2025)
por: Heineman, David, et al.
Publicado: (2025)
Evaluating Game Difficulty in Tetris Block Puzzle
por: Wang, Chun-Jui, et al.
Publicado: (2026)
por: Wang, Chun-Jui, et al.
Publicado: (2026)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
por: Haller, Patrick, et al.
Publicado: (2025)
por: Haller, Patrick, et al.
Publicado: (2025)
Lower Difficulty and Better Robustness: A Bregman Divergence Perspective for Adversarial Training
por: Wu, Zihui, et al.
Publicado: (2022)
por: Wu, Zihui, et al.
Publicado: (2022)
The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter
por: Bell, Samuel J., et al.
Publicado: (2024)
por: Bell, Samuel J., et al.
Publicado: (2024)
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
por: Liu, Siyuan, et al.
Publicado: (2026)
por: Liu, Siyuan, et al.
Publicado: (2026)
Reassessing the Validity of Spurious Correlations Benchmarks
por: Bell, Samuel J., et al.
Publicado: (2024)
por: Bell, Samuel J., et al.
Publicado: (2024)
Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs
por: Yadav, Suraj, et al.
Publicado: (2026)
por: Yadav, Suraj, et al.
Publicado: (2026)
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
por: Luo, Qin-Wen, et al.
Publicado: (2026)
por: Luo, Qin-Wen, et al.
Publicado: (2026)
Interpretability of Language Models via Task Spaces
por: Weber, Lucas, et al.
Publicado: (2024)
por: Weber, Lucas, et al.
Publicado: (2024)
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
por: Sun, Yifan, et al.
Publicado: (2025)
por: Sun, Yifan, et al.
Publicado: (2025)
Beyond Scaling Curves: Internal Dynamics of Neural Networks Through the NTK Lens
por: Nikolaou, Konstantin, et al.
Publicado: (2025)
por: Nikolaou, Konstantin, et al.
Publicado: (2025)
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
por: Chen, Tianhao, et al.
Publicado: (2025)
por: Chen, Tianhao, et al.
Publicado: (2025)
Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
por: Jandali, Yaman, et al.
Publicado: (2025)
por: Jandali, Yaman, et al.
Publicado: (2025)
Circuit Transformer: A Transformer That Preserves Logical Equivalence
por: Li, Xihan, et al.
Publicado: (2024)
por: Li, Xihan, et al.
Publicado: (2024)
Scaling Up Diffusion and Flow-based XGBoost Models
por: Cresswell, Jesse C., et al.
Publicado: (2024)
por: Cresswell, Jesse C., et al.
Publicado: (2024)
Identity-Link IRT for Label-Free LLM Evaluation: Preserving Additivity in TVD-MI Scores
por: Robertson, Zachary
Publicado: (2025)
por: Robertson, Zachary
Publicado: (2025)
Robust Mitigation of Age-Dependent Confounding Effects via Sample-Difficulty Decorrelation
por: Kurian, Nikhil Cherian, et al.
Publicado: (2026)
por: Kurian, Nikhil Cherian, et al.
Publicado: (2026)
D3: Diversity, Difficulty, and Dependability-Aware Data Selection for Sample-Efficient LLM Instruction Tuning
por: Zhang, Jia, et al.
Publicado: (2025)
por: Zhang, Jia, et al.
Publicado: (2025)
Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies
por: Schneider, Chris, et al.
Publicado: (2026)
por: Schneider, Chris, et al.
Publicado: (2026)
Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation
por: Kricheli, Joshua Shay, et al.
Publicado: (2026)
por: Kricheli, Joshua Shay, et al.
Publicado: (2026)
LLM-Assisted Logic Rule Learning: Scaling Human Expertise for Time Series Anomaly Detection
por: Zhang, Haoting, et al.
Publicado: (2026)
por: Zhang, Haoting, et al.
Publicado: (2026)
Ejemplares similares
-
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
por: Roberts, Nicholas, et al.
Publicado: (2025) -
Quantifying Variance in Evaluation Benchmarks
por: Madaan, Lovish, et al.
Publicado: (2024) -
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
por: Mondorf, Philipp, et al.
Publicado: (2024) -
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
por: Hupkes, Dieuwke, et al.
Publicado: (2025) -
Brittlebench: Quantifying LLM robustness via prompt sensitivity
por: Romanou, Angelika, et al.
Publicado: (2026)