Quantifying the Capabilities of LLMs across Scale and Precision
Fuente:
arXiv
Salvato in:
| Autori principali: | Badshah, Sher, Sajjad, Hassan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
di: Badshah, Sher, et al.
Pubblicazione: (2024)
di: Badshah, Sher, et al.
Pubblicazione: (2024)
TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
di: Badshah, Sher, et al.
Pubblicazione: (2025)
di: Badshah, Sher, et al.
Pubblicazione: (2025)
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
di: Feng, Guhao, et al.
Pubblicazione: (2024)
di: Feng, Guhao, et al.
Pubblicazione: (2024)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
di: Badshah, Sher, et al.
Pubblicazione: (2025)
di: Badshah, Sher, et al.
Pubblicazione: (2025)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
di: Lin, Xiaofeng, et al.
Pubblicazione: (2026)
di: Lin, Xiaofeng, et al.
Pubblicazione: (2026)
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
di: Badshah, Sher, et al.
Pubblicazione: (2026)
di: Badshah, Sher, et al.
Pubblicazione: (2026)
LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering
di: Wong, Sing Hieng, et al.
Pubblicazione: (2026)
di: Wong, Sing Hieng, et al.
Pubblicazione: (2026)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
di: Garg, Ankur, et al.
Pubblicazione: (2025)
di: Garg, Ankur, et al.
Pubblicazione: (2025)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
di: Arbabi, Alireza, et al.
Pubblicazione: (2025)
di: Arbabi, Alireza, et al.
Pubblicazione: (2025)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
di: Deng, Wenhao, et al.
Pubblicazione: (2025)
di: Deng, Wenhao, et al.
Pubblicazione: (2025)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
di: Turk, Matt
Pubblicazione: (2026)
di: Turk, Matt
Pubblicazione: (2026)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
di: Ramezanali, Mohammad, et al.
Pubblicazione: (2025)
di: Ramezanali, Mohammad, et al.
Pubblicazione: (2025)
Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
di: Haider, Muhammad Umair, et al.
Pubblicazione: (2025)
di: Haider, Muhammad Umair, et al.
Pubblicazione: (2025)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
di: Hu, Haiquan, et al.
Pubblicazione: (2025)
di: Hu, Haiquan, et al.
Pubblicazione: (2025)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
Quantifying Document Impact in RAG-LLMs
di: Gerami, Armin, et al.
Pubblicazione: (2025)
di: Gerami, Armin, et al.
Pubblicazione: (2025)
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
di: Ball, Thomas, et al.
Pubblicazione: (2024)
di: Ball, Thomas, et al.
Pubblicazione: (2024)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
di: DeepSeek-AI, et al.
Pubblicazione: (2025)
di: DeepSeek-AI, et al.
Pubblicazione: (2025)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
di: Rizvi, Md Imbesat Hassan, et al.
Pubblicazione: (2024)
di: Rizvi, Md Imbesat Hassan, et al.
Pubblicazione: (2024)
ALKAFI-LLAMA3: Fine-Tuning LLMs for Precise Legal Understanding in Palestine
di: Qasem, Rabee, et al.
Pubblicazione: (2024)
di: Qasem, Rabee, et al.
Pubblicazione: (2024)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
di: Mao, Yujun, et al.
Pubblicazione: (2024)
di: Mao, Yujun, et al.
Pubblicazione: (2024)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
di: Tian, Yijun, et al.
Pubblicazione: (2024)
di: Tian, Yijun, et al.
Pubblicazione: (2024)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
di: Dong, Yihong, et al.
Pubblicazione: (2025)
di: Dong, Yihong, et al.
Pubblicazione: (2025)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
Early Signs of Steganographic Capabilities in Frontier LLMs
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2025)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2025)
Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning
di: Chen, Nan, et al.
Pubblicazione: (2026)
di: Chen, Nan, et al.
Pubblicazione: (2026)
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
di: Jan, Essa, et al.
Pubblicazione: (2025)
di: Jan, Essa, et al.
Pubblicazione: (2025)
DLO: Dynamic Layer Operation for Efficient Vertical Scaling of LLMs
di: Tan, Zhen, et al.
Pubblicazione: (2024)
di: Tan, Zhen, et al.
Pubblicazione: (2024)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
LLMs cannot find reasoning errors, but can correct them given the error location
di: Tyen, Gladys, et al.
Pubblicazione: (2023)
di: Tyen, Gladys, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
di: Badshah, Sher, et al.
Pubblicazione: (2024) -
TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
di: Badshah, Sher, et al.
Pubblicazione: (2025) -
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025) -
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
di: Feng, Guhao, et al.
Pubblicazione: (2024) -
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
di: Badshah, Sher, et al.
Pubblicazione: (2025)