On Speeding Up Language Model Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Jin Peng, Belardi, Christian K., Wu, Ruihan, Zhang, Travis, Gomes, Carla P., Sun, Wen, Weinberger, Kilian Q. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pre-training Limited Memory Language Models with Internal and External Knowledge
von: Zhao, Linxi, et al.
Veröffentlicht: (2025)
von: Zhao, Linxi, et al.
Veröffentlicht: (2025)
Orchestrating LLMs with Different Personalizations
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024)
INPROVF: Leveraging Large Language Models to Repair High-level Robot Controllers from Assumption Violations
von: Meng, Qian, et al.
Veröffentlicht: (2025)
von: Meng, Qian, et al.
Veröffentlicht: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024)
Correction with Backtracking Reduces Hallucination in Summarization
von: Liu, Zhenzhen, et al.
Veröffentlicht: (2023)
von: Liu, Zhenzhen, et al.
Veröffentlicht: (2023)
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
von: Gong, Albert, et al.
Veröffentlicht: (2025)
von: Gong, Albert, et al.
Veröffentlicht: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
von: Jin, Ruihan, et al.
Veröffentlicht: (2025)
von: Jin, Ruihan, et al.
Veröffentlicht: (2025)
Learning from Synthetic Data Improves Multi-hop Reasoning
von: Kabra, Anmol, et al.
Veröffentlicht: (2026)
von: Kabra, Anmol, et al.
Veröffentlicht: (2026)
Fake News Detection and Manipulation Reasoning via Large Vision-Language Models
von: Jin, Ruihan, et al.
Veröffentlicht: (2024)
von: Jin, Ruihan, et al.
Veröffentlicht: (2024)
AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?
von: Press, Ori, et al.
Veröffentlicht: (2025)
von: Press, Ori, et al.
Veröffentlicht: (2025)
Stop-Think-AutoRegress: Language Modeling with Latent Diffusion Planning
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
von: He, Nan, et al.
Veröffentlicht: (2024)
von: He, Nan, et al.
Veröffentlicht: (2024)
Adaptive Moments are Surprisingly Effective for Plug-and-Play Diffusion Sampling
von: Belardi, Christian, et al.
Veröffentlicht: (2026)
von: Belardi, Christian, et al.
Veröffentlicht: (2026)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
Prescriptive Scaling Laws for Data Constrained Training
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
DReSS: Data-driven Regularized Structured Streamlining for Large Language Models
von: Feng, Mingkuan, et al.
Veröffentlicht: (2025)
von: Feng, Mingkuan, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models on Financial Report Summarization: An Empirical Study
von: Yang, Xinqi, et al.
Veröffentlicht: (2024)
von: Yang, Xinqi, et al.
Veröffentlicht: (2024)
A Review of Multi-Modal Large Language and Vision Models
von: Carolan, Kilian, et al.
Veröffentlicht: (2024)
von: Carolan, Kilian, et al.
Veröffentlicht: (2024)
From Implicit to Explicit: Enhancing Self-Recognition in Large Language Models
von: Zhou, Yinghan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinghan, et al.
Veröffentlicht: (2025)
Are Language Models Up to Sequential Optimization Problems? From Evaluation to a Hegelian-Inspired Enhancement
von: Abbasloo, Soheil
Veröffentlicht: (2025)
von: Abbasloo, Soheil
Veröffentlicht: (2025)
ImF: Implicit Fingerprint for Large Language Models
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
SentenceVAE: Enable Next-sentence Prediction for Large Language Models with Faster Speed, Higher Accuracy and Longer Context
von: An, Hongjun, et al.
Veröffentlicht: (2024)
von: An, Hongjun, et al.
Veröffentlicht: (2024)
Reasoning Up the Instruction Ladder for Controllable Language Models
von: Zheng, Zishuo, et al.
Veröffentlicht: (2025)
von: Zheng, Zishuo, et al.
Veröffentlicht: (2025)
Language Models can Evaluate Themselves via Probability Discrepancy
von: Xia, Tingyu, et al.
Veröffentlicht: (2024)
von: Xia, Tingyu, et al.
Veröffentlicht: (2024)
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models
von: Fodor, James
Veröffentlicht: (2025)
von: Fodor, James
Veröffentlicht: (2025)
DiffScore: Text Evaluation Beyond Autoregressive Likelihood
von: Lai, Wen, et al.
Veröffentlicht: (2026)
von: Lai, Wen, et al.
Veröffentlicht: (2026)
Speed and Conversational Large Language Models: Not All Is About Tokens per Second
von: Conde, Javier, et al.
Veröffentlicht: (2025)
von: Conde, Javier, et al.
Veröffentlicht: (2025)
ARE: Scaling Up Agent Environments and Evaluations
von: Froger, Romain, et al.
Veröffentlicht: (2025)
von: Froger, Romain, et al.
Veröffentlicht: (2025)
Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration
von: Zhong, Linhao, et al.
Veröffentlicht: (2026)
von: Zhong, Linhao, et al.
Veröffentlicht: (2026)
LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought
von: Qi, Ruiyan, et al.
Veröffentlicht: (2025)
von: Qi, Ruiyan, et al.
Veröffentlicht: (2025)
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
von: Sun, Peng, et al.
Veröffentlicht: (2026)
von: Sun, Peng, et al.
Veröffentlicht: (2026)
Evaluating Ill-Defined Tasks in Large Language Models
von: Zhou, Yi, et al.
Veröffentlicht: (2026)
von: Zhou, Yi, et al.
Veröffentlicht: (2026)
LawGPT: A Chinese Legal Knowledge-Enhanced Large Language Model
von: Zhou, Zhi, et al.
Veröffentlicht: (2024)
von: Zhou, Zhi, et al.
Veröffentlicht: (2024)
Copy-Paste to Mitigate Large Language Model Hallucinations
von: Long, Yongchao, et al.
Veröffentlicht: (2025)
von: Long, Yongchao, et al.
Veröffentlicht: (2025)
Follow-Up Questions Improve Documents Generated by Large Language Models
von: Tix, Bernadette J
Veröffentlicht: (2024)
von: Tix, Bernadette J
Veröffentlicht: (2024)
Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
von: Li, Boyi, et al.
Veröffentlicht: (2022)
von: Li, Boyi, et al.
Veröffentlicht: (2022)
OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology
von: Zhou, Chengfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Chengfeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pre-training Limited Memory Language Models with Internal and External Knowledge
von: Zhao, Linxi, et al.
Veröffentlicht: (2025) -
Orchestrating LLMs with Different Personalizations
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024) -
INPROVF: Leveraging Large Language Models to Repair High-level Robot Controllers from Assumption Violations
von: Meng, Qian, et al.
Veröffentlicht: (2025) -
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024) -
Correction with Backtracking Reduces Hallucination in Summarization
von: Liu, Zhenzhen, et al.
Veröffentlicht: (2023)