BSBench: will your LLM find the largest prime number?
Fuente:
arXiv
Salvato in:
| Autore principale: | Erziev, K. O. T. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
À la recherche du sens perdu: your favourite LLM might have more to say than you can understand
di: Erziev, K. O. T.
Pubblicazione: (2025)
di: Erziev, K. O. T.
Pubblicazione: (2025)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
di: Łazuka, Małgorzata, et al.
Pubblicazione: (2024)
di: Łazuka, Małgorzata, et al.
Pubblicazione: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
di: Zhang, Boning, et al.
Pubblicazione: (2024)
di: Zhang, Boning, et al.
Pubblicazione: (2024)
Let your LLM generate a few tokens and you will reduce the need for retrieval
di: Déjean, Hervé
Pubblicazione: (2024)
di: Déjean, Hervé
Pubblicazione: (2024)
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
di: Cheng, Yu-Ang, et al.
Pubblicazione: (2025)
di: Cheng, Yu-Ang, et al.
Pubblicazione: (2025)
Data-Prep-Kit: getting your data ready for LLM application development
di: Wood, David, et al.
Pubblicazione: (2024)
di: Wood, David, et al.
Pubblicazione: (2024)
On the largest prime factors of shifted semiprime numbers
di: Tam, Do Duc
Pubblicazione: (2025)
di: Tam, Do Duc
Pubblicazione: (2025)
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs
di: Haq, Saiful, et al.
Pubblicazione: (2025)
di: Haq, Saiful, et al.
Pubblicazione: (2025)
Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
di: Zhao, Shu, et al.
Pubblicazione: (2025)
di: Zhao, Shu, et al.
Pubblicazione: (2025)
A hierarchical Bayesian model for syntactic priming
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
Is your multimodal large language model a good science tutor?
di: Liu, Ming, et al.
Pubblicazione: (2025)
di: Liu, Ming, et al.
Pubblicazione: (2025)
Making the Most of your Model: Methods for Finetuning and Applying Pretrained Transformers
di: Yoshida, Davis
Pubblicazione: (2024)
di: Yoshida, Davis
Pubblicazione: (2024)
Getting the most out of your tokenizer for pre-training and domain adaptation
di: Dagan, Gautier, et al.
Pubblicazione: (2024)
di: Dagan, Gautier, et al.
Pubblicazione: (2024)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
Finding your MUSE: Mining Unexpected Solutions Engine
di: Sweed, Nir, et al.
Pubblicazione: (2025)
di: Sweed, Nir, et al.
Pubblicazione: (2025)
"Image, Tell me your story!" Predicting the original meta-context of visual misinformation
di: Tonglet, Jonathan, et al.
Pubblicazione: (2024)
di: Tonglet, Jonathan, et al.
Pubblicazione: (2024)
In your own words: computationally identifying interpretable themes in free-text survey data
di: Wang, Jenny S, et al.
Pubblicazione: (2026)
di: Wang, Jenny S, et al.
Pubblicazione: (2026)
Progressing beyond Art Masterpieces or Touristic Clichés: how to assess your LLMs for cultural alignment?
di: Branco, António, et al.
Pubblicazione: (2026)
di: Branco, António, et al.
Pubblicazione: (2026)
When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
di: Bafna, Niyati, et al.
Pubblicazione: (2023)
di: Bafna, Niyati, et al.
Pubblicazione: (2023)
Source-primed Multi-turn Conversation Helps Large Language Models Translate Documents
di: Hu, Hanxu, et al.
Pubblicazione: (2025)
di: Hu, Hanxu, et al.
Pubblicazione: (2025)
Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely
di: Zhao, Siyun, et al.
Pubblicazione: (2024)
di: Zhao, Siyun, et al.
Pubblicazione: (2024)
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
di: Chen, Zhiliang, et al.
Pubblicazione: (2025)
di: Chen, Zhiliang, et al.
Pubblicazione: (2025)
Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval
di: Ma, Guangyuan, et al.
Pubblicazione: (2024)
di: Ma, Guangyuan, et al.
Pubblicazione: (2024)
A survey of neural-network-based methods utilising comparable data for finding translation equivalents
di: Denisová, Michaela, et al.
Pubblicazione: (2024)
di: Denisová, Michaela, et al.
Pubblicazione: (2024)
CELL your Model: Contrastive Explanations for Large Language Models
di: Luss, Ronny, et al.
Pubblicazione: (2024)
di: Luss, Ronny, et al.
Pubblicazione: (2024)
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
di: Yugeswardeenoo, Dharunish, et al.
Pubblicazione: (2024)
di: Yugeswardeenoo, Dharunish, et al.
Pubblicazione: (2024)
Don't lie to your friends: Learning what you know from collaborative self-play
di: Eisenstein, Jacob, et al.
Pubblicazione: (2025)
di: Eisenstein, Jacob, et al.
Pubblicazione: (2025)
Does your data spark joy? Performance gains from domain upsampling at the end of training
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
Generics are puzzling. Can language models find the missing piece?
di: Calderón, Gustavo Cilleruelo, et al.
Pubblicazione: (2024)
di: Calderón, Gustavo Cilleruelo, et al.
Pubblicazione: (2024)
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
di: Shin, Jisu, et al.
Pubblicazione: (2024)
di: Shin, Jisu, et al.
Pubblicazione: (2024)
The Impact of Annotator Personas on LLM Behavior Across the Perspectivism Spectrum
di: Sarumi, Olufunke O., et al.
Pubblicazione: (2025)
di: Sarumi, Olufunke O., et al.
Pubblicazione: (2025)
Continual Pre-training of MoEs: How robust is your router?
di: Thérien, Benjamin, et al.
Pubblicazione: (2025)
di: Thérien, Benjamin, et al.
Pubblicazione: (2025)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
di: Chen, Peter, et al.
Pubblicazione: (2025)
di: Chen, Peter, et al.
Pubblicazione: (2025)
Jabuticaba: The largest commercial corpus for LLMs in Portuguese
di: Amadeus, Marcellus, et al.
Pubblicazione: (2025)
di: Amadeus, Marcellus, et al.
Pubblicazione: (2025)
QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks
di: Lundy, Taylor, et al.
Pubblicazione: (2026)
di: Lundy, Taylor, et al.
Pubblicazione: (2026)
Data-Efficient Domain Adaptation for LLM-based MT using Contrastive Preference Optimization
di: Vieira, Inacio, et al.
Pubblicazione: (2025)
di: Vieira, Inacio, et al.
Pubblicazione: (2025)
QuaLLM-Health: An Adaptation of an LLM-Based Framework for Quantitative Data Extraction from Online Health Discussions
di: Kouzy, Ramez, et al.
Pubblicazione: (2024)
di: Kouzy, Ramez, et al.
Pubblicazione: (2024)
GaelEval: Benchmarking LLM Performance for Scottish Gaelic
di: Devine, Peter, et al.
Pubblicazione: (2026)
di: Devine, Peter, et al.
Pubblicazione: (2026)
DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's Disease
di: Mo, Tingyu, et al.
Pubblicazione: (2025)
di: Mo, Tingyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
À la recherche du sens perdu: your favourite LLM might have more to say than you can understand
di: Erziev, K. O. T.
Pubblicazione: (2025) -
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
di: Łazuka, Małgorzata, et al.
Pubblicazione: (2024) -
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
di: Zhang, Boning, et al.
Pubblicazione: (2024) -
Let your LLM generate a few tokens and you will reduce the need for retrieval
di: Déjean, Hervé
Pubblicazione: (2024) -
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
di: Cheng, Yu-Ang, et al.
Pubblicazione: (2025)