Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ding, Mucong, Deng, Chenghao, Choo, Jocelyn, Wu, Zichu, Agrawal, Aakriti, Schwarzschild, Avi, Zhou, Tianyi, Goldstein, Tom, Langford, John, Anandkumar, Anima, Huang, Furong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2024)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2024)
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
WAVES: Benchmarking the Robustness of Image Watermarks
von: An, Bang, et al.
Veröffentlicht: (2024)
von: An, Bang, et al.
Veröffentlicht: (2024)
Benchmarking ChatGPT on Algorithmic Reasoning
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
SAFLEX: Self-Adaptive Augmentation via Feature Label Extrapolation
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
VeriGate: Verifier-Gated Step-Level Supervision for GRPO
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2026)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2026)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
von: Levin, Roman, et al.
Veröffentlicht: (2025)
von: Levin, Roman, et al.
Veröffentlicht: (2025)
Spectral Greedy Coresets for Graph Neural Networks
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
Sketch-GNN: Scalable Graph Neural Networks with Sublinear Training Complexity
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
Diffusion State-Guided Projected Gradient for Inverse Problems
von: Zirvi, Rayhan, et al.
Veröffentlicht: (2024)
von: Zirvi, Rayhan, et al.
Veröffentlicht: (2024)
Lean Copilot: Large Language Models as Copilots for Theorem Proving in Lean
von: Song, Peiyang, et al.
Veröffentlicht: (2024)
von: Song, Peiyang, et al.
Veröffentlicht: (2024)
Calibrated Uncertainty Quantification for Operator Learning via Conformal Prediction
von: Ma, Ziqi, et al.
Veröffentlicht: (2024)
von: Ma, Ziqi, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
Fourier Neural Operators Explained: A Practical Perspective
von: Duruisseaux, Valentin, et al.
Veröffentlicht: (2025)
von: Duruisseaux, Valentin, et al.
Veröffentlicht: (2025)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
GATES: Self-Distillation under Privileged Context with Consensus Gating
von: Stein, Alex, et al.
Veröffentlicht: (2026)
von: Stein, Alex, et al.
Veröffentlicht: (2026)
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2026)
von: Luo, Qin-Wen, et al.
Veröffentlicht: (2026)
Solving Poisson Equations using Neural Walk-on-Spheres
von: Nam, Hong Chul, et al.
Veröffentlicht: (2024)
von: Nam, Hong Chul, et al.
Veröffentlicht: (2024)
MetaLint: Easy-to-Hard Generalization for Code Linting
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
Command-V: Pasting LLM Behaviors via Activation Profiles
von: Wang, Barry, et al.
Veröffentlicht: (2025)
von: Wang, Barry, et al.
Veröffentlicht: (2025)
Fully Attentional Networks with Self-emerging Token Labeling
von: Zhao, Bingyin, et al.
Veröffentlicht: (2024)
von: Zhao, Bingyin, et al.
Veröffentlicht: (2024)
Geometric Operator Learning with Optimal Transport
von: Li, Xinyi, et al.
Veröffentlicht: (2025)
von: Li, Xinyi, et al.
Veröffentlicht: (2025)
Fast Training of Diffusion Models with Masked Transformers
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
LeanProgress: Guiding Search for Neural Theorem Proving via Proof Progress Prediction
von: George, Robert Joseph, et al.
Veröffentlicht: (2025)
von: George, Robert Joseph, et al.
Veröffentlicht: (2025)
Fourier Neural Operator with Learned Deformations for PDEs on General Geometries
von: Li, Zongyi, et al.
Veröffentlicht: (2022)
von: Li, Zongyi, et al.
Veröffentlicht: (2022)
Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
von: Nasriddinov, Firdavs, et al.
Veröffentlicht: (2025)
von: Nasriddinov, Firdavs, et al.
Veröffentlicht: (2025)
Pure Event Semantics
von: Roger Schwarzschild
Veröffentlicht: (2024)
von: Roger Schwarzschild
Veröffentlicht: (2024)
The CLRS-Text Algorithmic Reasoning Language Benchmark
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
Extrapolation by Association: Length Generalization Transfer in Transformers
von: Cai, Ziyang, et al.
Veröffentlicht: (2025)
von: Cai, Ziyang, et al.
Veröffentlicht: (2025)
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
Forcing Diffuse Distributions out of Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
von: Kordi, Yeganeh, et al.
Veröffentlicht: (2025)
von: Kordi, Yeganeh, et al.
Veröffentlicht: (2025)
Operator Learning Using Weak Supervision from Walk-on-Spheres
von: Viswanath, Hrishikesh, et al.
Veröffentlicht: (2026)
von: Viswanath, Hrishikesh, et al.
Veröffentlicht: (2026)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
InRank: Incremental Low-Rank Learning
von: Zhao, Jiawei, et al.
Veröffentlicht: (2023)
von: Zhao, Jiawei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2024) -
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025) -
WAVES: Benchmarking the Robustness of Image Watermarks
von: An, Bang, et al.
Veröffentlicht: (2024) -
Benchmarking ChatGPT on Algorithmic Reasoning
von: McLeish, Sean, et al.
Veröffentlicht: (2024) -
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)