100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
Fuente:
arXiv
Salvato in:
| Autori principali: | Pacchiardi, Lorenzo, Cheke, Lucy G., Hernández-Orallo, José |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2024)
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2024)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
di: Testini, Irene, et al.
Pubblicazione: (2025)
di: Testini, Irene, et al.
Pubblicazione: (2025)
PredictaBoard: Benchmarking LLM Score Predictability
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2025)
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2025)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
di: Burden, John, et al.
Pubblicazione: (2025)
di: Burden, John, et al.
Pubblicazione: (2025)
Collaboration is all you need: LLM Assisted Safe Code Translation
di: Karanjai, Rabimba, et al.
Pubblicazione: (2025)
di: Karanjai, Rabimba, et al.
Pubblicazione: (2025)
Large Language Models aren't all that you need
di: Holla, Kiran Voderhobli, et al.
Pubblicazione: (2024)
di: Holla, Kiran Voderhobli, et al.
Pubblicazione: (2024)
Inferring Capabilities from Task Performance with Bayesian Triangulation
di: Burden, John, et al.
Pubblicazione: (2023)
di: Burden, John, et al.
Pubblicazione: (2023)
Leveraging LLMs to support co-evolution between definitions and instances of textual DSLs
di: Zhang, Weixing, et al.
Pubblicazione: (2025)
di: Zhang, Weixing, et al.
Pubblicazione: (2025)
Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
di: Ren, Qingyu, et al.
Pubblicazione: (2025)
di: Ren, Qingyu, et al.
Pubblicazione: (2025)
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
di: Rutar, Danaja, et al.
Pubblicazione: (2025)
di: Rutar, Danaja, et al.
Pubblicazione: (2025)
Multilingual Natural Language Processing Model for Radiology Reports -- The Summary is all you need!
di: Lindo, Mariana, et al.
Pubblicazione: (2023)
di: Lindo, Mariana, et al.
Pubblicazione: (2023)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
di: Balkır, Esma, et al.
Pubblicazione: (2026)
di: Balkır, Esma, et al.
Pubblicazione: (2026)
Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation framework
di: Chang, Kaiyan, et al.
Pubblicazione: (2024)
di: Chang, Kaiyan, et al.
Pubblicazione: (2024)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
di: Cencerrado, Iván Vicente Moreno, et al.
Pubblicazione: (2025)
di: Cencerrado, Iván Vicente Moreno, et al.
Pubblicazione: (2025)
Visuospatial Perspective Taking in Multimodal Language Models
di: Prunty, Jonathan, et al.
Pubblicazione: (2026)
di: Prunty, Jonathan, et al.
Pubblicazione: (2026)
Conversational Complexity for Assessing Risk in Large Language Models
di: Burden, John, et al.
Pubblicazione: (2024)
di: Burden, John, et al.
Pubblicazione: (2024)
Effects of structure on reasoning in instance-level Self-Discover
di: Gunasekara, Sachith, et al.
Pubblicazione: (2025)
di: Gunasekara, Sachith, et al.
Pubblicazione: (2025)
Propaganda is all you need
di: Kronlund-Drouault, Paul
Pubblicazione: (2024)
di: Kronlund-Drouault, Paul
Pubblicazione: (2024)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
di: Fitz, Stephen, et al.
Pubblicazione: (2025)
di: Fitz, Stephen, et al.
Pubblicazione: (2025)
REL: Working out is all you need
di: Simonds, Toby, et al.
Pubblicazione: (2024)
di: Simonds, Toby, et al.
Pubblicazione: (2024)
Bringing Comparative Cognition To Computers
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2025)
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2025)
Where the Really Hard Quadratic Assignment Problems Are: the QAP-SAT instances
di: Verel, Sébastien, et al.
Pubblicazione: (2024)
di: Verel, Sébastien, et al.
Pubblicazione: (2024)
Compression is all you need: Modeling Mathematics
di: Aksenov, Vitaly, et al.
Pubblicazione: (2026)
di: Aksenov, Vitaly, et al.
Pubblicazione: (2026)
Learning county from pixels: corn yield prediction with attention-weighted multiple instance learning
di: Wang, Xiaoyu, et al.
Pubblicazione: (2023)
di: Wang, Xiaoyu, et al.
Pubblicazione: (2023)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
di: Cao, Yang Trista, et al.
Pubblicazione: (2023)
di: Cao, Yang Trista, et al.
Pubblicazione: (2023)
CTDGSI: A comprehensive exploitation of instance selection methods for automatic text classification. VII Concurso de Teses, Dissertações e Trabalhos de Graduação em SI -- XXI Simpósio Brasileiro de Sistemas de Informação
di: Cunha, Washington, et al.
Pubblicazione: (2025)
di: Cunha, Washington, et al.
Pubblicazione: (2025)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
di: Zhou, Lexin, et al.
Pubblicazione: (2025)
di: Zhou, Lexin, et al.
Pubblicazione: (2025)
Tabular Data: Is Deep Learning all you need?
di: Zabërgja, Guri, et al.
Pubblicazione: (2024)
di: Zabërgja, Guri, et al.
Pubblicazione: (2024)
A Two-Stage Algorithm for Cost-Efficient Multi-instance Counterfactual Explanations
di: Artelt, André, et al.
Pubblicazione: (2024)
di: Artelt, André, et al.
Pubblicazione: (2024)
Towards LLM-based optimization compilers. Can LLMs learn how to apply a single peephole optimization? Reasoning is all LLMs need!
di: Fang, Xiangxin, et al.
Pubblicazione: (2024)
di: Fang, Xiangxin, et al.
Pubblicazione: (2024)
Multi-agent AI systems outperform human teams in creativity
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
Seven simple steps for log analysis in AI systems
di: Dubois, Magda, et al.
Pubblicazione: (2026)
di: Dubois, Magda, et al.
Pubblicazione: (2026)
ScrapeGraphAI-100k: Dataset for Schema-Constrained LLM Generation
di: Brach, William, et al.
Pubblicazione: (2026)
di: Brach, William, et al.
Pubblicazione: (2026)
Let your LLM generate a few tokens and you will reduce the need for retrieval
di: Déjean, Hervé
Pubblicazione: (2024)
di: Déjean, Hervé
Pubblicazione: (2024)
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
di: Meincke, Lennart, et al.
Pubblicazione: (2025)
di: Meincke, Lennart, et al.
Pubblicazione: (2025)
Attention-based multiple instance learning for predominant growth pattern prediction in lung adenocarcinoma wsi using foundation models
di: Perez-Herrera, Laura Valeria, et al.
Pubblicazione: (2026)
di: Perez-Herrera, Laura Valeria, et al.
Pubblicazione: (2026)
Layered LA-MAPF: a decomposition of large agent MAPF instance to accelerate solving without compromising solvability
di: Yao, Zhuo
Pubblicazione: (2024)
di: Yao, Zhuo
Pubblicazione: (2024)
Neural Operator: Is data all you need to model the world? An insight into the paradigm of data-driven scientific ML
di: Viswanath, Hrishikesh, et al.
Pubblicazione: (2023)
di: Viswanath, Hrishikesh, et al.
Pubblicazione: (2023)
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
di: Bašaragin, Bojana, et al.
Pubblicazione: (2024)
di: Bašaragin, Bojana, et al.
Pubblicazione: (2024)
How new data permeates LLM knowledge and how to dilute it
di: Sun, Chen, et al.
Pubblicazione: (2025)
di: Sun, Chen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2024) -
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
di: Testini, Irene, et al.
Pubblicazione: (2025) -
PredictaBoard: Benchmarking LLM Score Predictability
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2025) -
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
di: Burden, John, et al.
Pubblicazione: (2025) -
Collaboration is all you need: LLM Assisted Safe Code Translation
di: Karanjai, Rabimba, et al.
Pubblicazione: (2025)