Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Egashira, Kazuki, Vero, Mark, Dekoninck, Jasper, Dorner, Florian E., Staab, Robin, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
by: Zhan, Xiaohua, et al.
Published: (2026)
by: Zhan, Xiaohua, et al.
Published: (2026)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
Mind the Gap: A Practical Attack on GGUF Quantization
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
A Synthetic Dataset for Personal Attribute Inference
by: Yukhymenko, Hanna, et al.
Published: (2024)
by: Yukhymenko, Hanna, et al.
Published: (2024)
Beyond Memorization: Violating Privacy Via Inference with Large Language Models
by: Staab, Robin, et al.
Published: (2023)
by: Staab, Robin, et al.
Published: (2023)
Private Attribute Inference from Images with Vision-Language Models
by: Tömekçe, Batuhan, et al.
Published: (2024)
by: Tömekçe, Batuhan, et al.
Published: (2024)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
Back to the Drawing Board for Fair Representation Learning
by: Pouget, Angéline, et al.
Published: (2024)
by: Pouget, Angéline, et al.
Published: (2024)
Adaptive Generation of Bias-Eliciting Questions for LLMs
by: Staab, Robin, et al.
Published: (2025)
by: Staab, Robin, et al.
Published: (2025)
Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
Large Language Models are Advanced Anonymizers
by: Staab, Robin, et al.
Published: (2024)
by: Staab, Robin, et al.
Published: (2024)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Watermark Stealing in Large Language Models
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Discovering Spoofing Attempts on Language Model Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2024)
by: Gloaguen, Thibaud, et al.
Published: (2024)
Watermarking Diffusion Language Models
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
A Unified Framework for LLM Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
Ward: Provable RAG Dataset Inference via LLM Watermarks
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Instruction Tuning for Secure Code Generation
by: He, Jingxuan, et al.
Published: (2024)
by: He, Jingxuan, et al.
Published: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
by: von Arx, Tobias, et al.
Published: (2025)
by: von Arx, Tobias, et al.
Published: (2025)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
by: Kawakami, Tatsuki, et al.
Published: (2025)
by: Kawakami, Tatsuki, et al.
Published: (2025)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
by: Vero, Mark, et al.
Published: (2026)
by: Vero, Mark, et al.
Published: (2026)
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
by: Dékány, Csaba, et al.
Published: (2025)
by: Dékány, Csaba, et al.
Published: (2025)
Automated Classification of Model Errors on ImageNet
by: Peychev, Momchil, et al.
Published: (2023)
by: Peychev, Momchil, et al.
Published: (2023)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
by: Mündler, Niels, et al.
Published: (2025)
by: Mündler, Niels, et al.
Published: (2025)
BaxBench: Can LLMs Generate Correct and Secure Backends?
by: Vero, Mark, et al.
Published: (2025)
by: Vero, Mark, et al.
Published: (2025)
Training on the Test Task Confounds Evaluation and Emergence
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
CuTS: Customizable Tabular Synthetic Data Generation
by: Vero, Mark, et al.
Published: (2023)
by: Vero, Mark, et al.
Published: (2023)
Understanding Certified Training with Interval Bound Propagation
by: Mao, Yuhao, et al.
Published: (2023)
by: Mao, Yuhao, et al.
Published: (2023)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
by: Liu, Zhanyu, et al.
Published: (2026)
by: Liu, Zhanyu, et al.
Published: (2026)
The Cost of Relaxation: Evaluating the Error in Convex Neural Network Verification
by: Papamichail, Merkouris, et al.
Published: (2026)
by: Papamichail, Merkouris, et al.
Published: (2026)
CTBENCH: A Library and Benchmark for Certified Training
by: Mao, Yuhao, et al.
Published: (2024)
by: Mao, Yuhao, et al.
Published: (2024)
Similar Items
-
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
by: Zhan, Xiaohua, et al.
Published: (2026) -
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024) -
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
by: Egashira, Kazuki, et al.
Published: (2025) -
Mind the Gap: A Practical Attack on GGUF Quantization
by: Egashira, Kazuki, et al.
Published: (2025) -
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
by: Gloaguen, Thibaud, et al.
Published: (2025)