QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Jeremy, Andriushchenko, Maksym |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
by: Freeman, Joshua, et al.
Published: (2024)
by: Freeman, Joshua, et al.
Published: (2024)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
Decomposing and Measuring Evaluation Awareness
by: Li, Changling, et al.
Published: (2026)
by: Li, Changling, et al.
Published: (2026)
Robust Multi-Modal Forecasting: Integrating Static and Dynamic Features
by: Qin, Jeremy
Published: (2025)
by: Qin, Jeremy
Published: (2025)
Towards Trustworthy Vital Sign Forecasting: Leveraging Uncertainty for Prediction Intervals
by: Wang, Li Rong, et al.
Published: (2025)
by: Wang, Li Rong, et al.
Published: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
by: Goel, Shashwat, et al.
Published: (2026)
by: Goel, Shashwat, et al.
Published: (2026)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data
by: Cheng, Pu, et al.
Published: (2026)
by: Cheng, Pu, et al.
Published: (2026)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
by: Li, Pingzhi, et al.
Published: (2024)
by: Li, Pingzhi, et al.
Published: (2024)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Is More Context Always Better? Examining LLM Reasoning Capability for Time Interval Prediction
by: Cao, Yanan, et al.
Published: (2026)
by: Cao, Yanan, et al.
Published: (2026)
HindSight: Evaluating LLM-Generated Research Ideas via Future Impact
by: Jiang, Bo
Published: (2026)
by: Jiang, Bo
Published: (2026)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
by: Panfilov, Alexander, et al.
Published: (2026)
by: Panfilov, Alexander, et al.
Published: (2026)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025)
by: Terekhov, Mikhail, et al.
Published: (2025)
Enhancing Interval Type-2 Fuzzy Logic Systems: Learning for Precision and Prediction Intervals
by: Koklu, Ata, et al.
Published: (2024)
by: Koklu, Ata, et al.
Published: (2024)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
QuantFactor REINFORCE: Mining Steady Formulaic Alpha Factors with Variance-bounded REINFORCE
by: Zhao, Junjie, et al.
Published: (2024)
by: Zhao, Junjie, et al.
Published: (2024)
EF-LLM: Energy Forecasting LLM with AI-assisted Automation, Enhanced Sparse Prediction, Hallucination Detection
by: Qiu, Zihang, et al.
Published: (2024)
by: Qiu, Zihang, et al.
Published: (2024)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
Deep Language Geometry: Constructing a Metric Space from LLM Weights
by: Shamrai, Maksym, et al.
Published: (2025)
by: Shamrai, Maksym, et al.
Published: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
CONTINA: Confidence Interval for Traffic Demand Prediction with Coverage Guarantee
by: Yang, Chao, et al.
Published: (2025)
by: Yang, Chao, et al.
Published: (2025)
Tube Loss: A Novel Approach for Prediction Interval Estimation
by: Anand, Pritam, et al.
Published: (2024)
by: Anand, Pritam, et al.
Published: (2024)
Unveil Sources of Uncertainty: Feature Contribution to Conformal Prediction Intervals
by: Idrissi, Marouane Il, et al.
Published: (2025)
by: Idrissi, Marouane Il, et al.
Published: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting
by: Feng, Wanjin, et al.
Published: (2025)
by: Feng, Wanjin, et al.
Published: (2025)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
by: Karger, Ezra, et al.
Published: (2024)
by: Karger, Ezra, et al.
Published: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction
by: Yang, Zhiwen, et al.
Published: (2026)
by: Yang, Zhiwen, et al.
Published: (2026)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale
by: Thomas, Noel
Published: (2026)
by: Thomas, Noel
Published: (2026)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
by: Xu, Bingxin, et al.
Published: (2025)
by: Xu, Bingxin, et al.
Published: (2025)
Similar Items
-
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026) -
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024) -
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024) -
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
by: Freeman, Joshua, et al.
Published: (2024) -
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)