On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Siddiqui, Shoaib Ahmed, Chen, Yanzhi, Heo, Juyeon, Xia, Menglin, Weller, Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Capability Control Should be a Separate Goal From Alignment
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
Estimation of Concept Explanations Should be Uncertainty Aware
by: Piratla, Vihari, et al.
Published: (2023)
by: Piratla, Vihari, et al.
Published: (2023)
A deeper look at depth pruning of LLMs
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
The Topological Trouble With Transformers
by: Mozer, Michael C., et al.
Published: (2026)
by: Mozer, Michael C., et al.
Published: (2026)
Permissive Information-Flow Analysis for Large Language Models
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
Do Concept Bottleneck Models Respect Localities?
by: Raman, Naveen, et al.
Published: (2024)
by: Raman, Naveen, et al.
Published: (2024)
LLMs Judging LLMs: A Simplex Perspective
by: Vossler, Patrick, et al.
Published: (2025)
by: Vossler, Patrick, et al.
Published: (2025)
Exploring the design space of deep-learning-based weather forecasting systems
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
Evaluating LLMs Capabilities Towards Understanding Social Dynamics
by: Tahir, Anique, et al.
Published: (2024)
by: Tahir, Anique, et al.
Published: (2024)
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
Enhancing Neural Function Approximation: The XNet Outperforming KAN
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
FSP-Laplace: Function-Space Priors for the Laplace Approximation in Bayesian Deep Learning
by: Cinquin, Tristan, et al.
Published: (2024)
by: Cinquin, Tristan, et al.
Published: (2024)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
by: Han, Dongge, et al.
Published: (2025)
by: Han, Dongge, et al.
Published: (2025)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
by: Li, Haoming, et al.
Published: (2025)
by: Li, Haoming, et al.
Published: (2025)
ALVIN: Active Learning Via INterpolation
by: Korakakis, Michalis, et al.
Published: (2024)
by: Korakakis, Michalis, et al.
Published: (2024)
Mitigating Shortcut Learning with InterpoLated Learning
by: Korakakis, Michalis, et al.
Published: (2025)
by: Korakakis, Michalis, et al.
Published: (2025)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
by: Skelic, Lejla, et al.
Published: (2025)
by: Skelic, Lejla, et al.
Published: (2025)
Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving
by: Didolkar, Aniket, et al.
Published: (2024)
by: Didolkar, Aniket, et al.
Published: (2024)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
by: Shi, Yichen, et al.
Published: (2025)
by: Shi, Yichen, et al.
Published: (2025)
Unsupervised Outlier Detection using Random Subspace and Subsampling Ensembles of Dirichlet Process Mixtures
by: Kim, Dongwook, et al.
Published: (2024)
by: Kim, Dongwook, et al.
Published: (2024)
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
by: Feng, Zhangying, et al.
Published: (2025)
by: Feng, Zhangying, et al.
Published: (2025)
Binary structured physics-informed neural networks for solving equations with rapidly changing solutions
by: Liu, Yanzhi, et al.
Published: (2024)
by: Liu, Yanzhi, et al.
Published: (2024)
On the Weaknesses of Backdoor-based Model Watermarking: An Information-theoretic Perspective
by: Hu, Aoting, et al.
Published: (2024)
by: Hu, Aoting, et al.
Published: (2024)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
by: Yao, Zhiyuan, et al.
Published: (2026)
by: Yao, Zhiyuan, et al.
Published: (2026)
Tubular Riemannian Laplace Approximations for Bayesian Neural Networks
by: David, Rodrigo Pereira
Published: (2025)
by: David, Rodrigo Pereira
Published: (2025)
Evaluating Large Language Models for Security Bug Report Prediction
by: Soltaniani, Farnaz, et al.
Published: (2026)
by: Soltaniani, Farnaz, et al.
Published: (2026)
Batch Acquisition Function Evaluations and Decouple Optimizer Updates for Faster Bayesian Optimization
by: Irie, Kaichi, et al.
Published: (2025)
by: Irie, Kaichi, et al.
Published: (2025)
Experiment Planning with Function Approximation
by: Pacchiano, Aldo, et al.
Published: (2024)
by: Pacchiano, Aldo, et al.
Published: (2024)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
by: Hu, Haiquan, et al.
Published: (2025)
by: Hu, Haiquan, et al.
Published: (2025)
Geometric Scaling of Bayesian Inference in LLMs
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Text2Chart31: Instruction Tuning for Chart Generation with Automatic Feedback
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2024)
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2024)
The Elicitation Game: Evaluating Capability Elicitation Techniques
by: Hofstätter, Felix, et al.
Published: (2025)
by: Hofstätter, Felix, et al.
Published: (2025)
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
by: Miao, Tingjia, et al.
Published: (2026)
by: Miao, Tingjia, et al.
Published: (2026)
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
by: Wang, Weixin, et al.
Published: (2025)
by: Wang, Weixin, et al.
Published: (2025)
HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules
by: He, Neil, et al.
Published: (2025)
by: He, Neil, et al.
Published: (2025)
Interaction-Aware Influence Functions for Group Attribution
by: Heo, Jaeseung, et al.
Published: (2026)
by: Heo, Jaeseung, et al.
Published: (2026)
From Tokenizer Bias to Backbone Capability: A Controlled Study of LLMs for Time Series Forecasting
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Similar Items
-
Position: Capability Control Should be a Separate Goal From Alignment
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026) -
Estimation of Concept Explanations Should be Uncertainty Aware
by: Piratla, Vihari, et al.
Published: (2023) -
A deeper look at depth pruning of LLMs
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024) -
The Topological Trouble With Transformers
by: Mozer, Michael C., et al.
Published: (2026) -
Permissive Information-Flow Analysis for Large Language Models
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)