TEL'M: Test and Evaluation of Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Cybenko, George, Ackerman, Joshua, Lintilhac, Paul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Sharper Picture of Generalization in Transformers
di: Lintilhac, Paul, et al.
Pubblicazione: (2026)
di: Lintilhac, Paul, et al.
Pubblicazione: (2026)
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
di: Ackerman, Christopher
Pubblicazione: (2026)
di: Ackerman, Christopher
Pubblicazione: (2026)
Survey on Reasoning Capabilities and Accessibility of Large Language Models Using Biology-related Questions
di: Ackerman, Michael
Pubblicazione: (2024)
di: Ackerman, Michael
Pubblicazione: (2024)
Evidence for Limited Metacognition in LLMs
di: Ackerman, Christopher
Pubblicazione: (2025)
di: Ackerman, Christopher
Pubblicazione: (2025)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
di: Ackerman, Samuel, et al.
Pubblicazione: (2025)
di: Ackerman, Samuel, et al.
Pubblicazione: (2025)
Application and Evaluation of Large Language Models for Forecasting the Impact of Traffic Incidents
di: Jagadeesh, George, et al.
Pubblicazione: (2025)
di: Jagadeesh, George, et al.
Pubblicazione: (2025)
Evaluating Language Models' Evaluations of Games
di: Collins, Katherine M., et al.
Pubblicazione: (2025)
di: Collins, Katherine M., et al.
Pubblicazione: (2025)
Evaluation of Causal Reasoning for Large Language Models in Contextualized Clinical Scenarios of Laboratory Test Interpretation
di: Bhasuran, Balu, et al.
Pubblicazione: (2025)
di: Bhasuran, Balu, et al.
Pubblicazione: (2025)
Large Language Models as Test Case Generators: Performance Evaluation and Enhancement
di: Li, Kefan, et al.
Pubblicazione: (2024)
di: Li, Kefan, et al.
Pubblicazione: (2024)
LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation
di: Wang, Zhijie, et al.
Pubblicazione: (2024)
di: Wang, Zhijie, et al.
Pubblicazione: (2024)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
di: Ackerman, Christopher, et al.
Pubblicazione: (2024)
di: Ackerman, Christopher, et al.
Pubblicazione: (2024)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
di: Kour, George, et al.
Pubblicazione: (2025)
di: Kour, George, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for the Generation of Unit Tests with Equivalence Partitions and Boundary Values
di: Rodríguez, Martín, et al.
Pubblicazione: (2025)
di: Rodríguez, Martín, et al.
Pubblicazione: (2025)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
di: Jackson, Declan, et al.
Pubblicazione: (2025)
di: Jackson, Declan, et al.
Pubblicazione: (2025)
Adversarial Moral Stress Testing of Large Language Models
di: Jamshidi, Saeid, et al.
Pubblicazione: (2026)
di: Jamshidi, Saeid, et al.
Pubblicazione: (2026)
LocateBench: Evaluating the Locating Ability of Vision Language Models
di: Chiang, Ting-Rui, et al.
Pubblicazione: (2024)
di: Chiang, Ting-Rui, et al.
Pubblicazione: (2024)
Mitigating Many-Shot Jailbreaking
di: Ackerman, Christopher M., et al.
Pubblicazione: (2025)
di: Ackerman, Christopher M., et al.
Pubblicazione: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
di: Wang, Wenxuan
Pubblicazione: (2024)
di: Wang, Wenxuan
Pubblicazione: (2024)
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
di: Dsouza, Amanda, et al.
Pubblicazione: (2024)
di: Dsouza, Amanda, et al.
Pubblicazione: (2024)
Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking
di: Zheng, Tianpeng, et al.
Pubblicazione: (2026)
di: Zheng, Tianpeng, et al.
Pubblicazione: (2026)
The Workflow as Medium: A Framework for Navigating Human-AI Co-Creation
di: Ackerman, Lee
Pubblicazione: (2025)
di: Ackerman, Lee
Pubblicazione: (2025)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
di: Saad-Falcon, Jon, et al.
Pubblicazione: (2024)
di: Saad-Falcon, Jon, et al.
Pubblicazione: (2024)
Probabilistic Medical Predictions of Large Language Models
di: Gu, Bowen, et al.
Pubblicazione: (2024)
di: Gu, Bowen, et al.
Pubblicazione: (2024)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
di: Holtermann, Carolin, et al.
Pubblicazione: (2024)
di: Holtermann, Carolin, et al.
Pubblicazione: (2024)
Evaluating Language Models for Generating and Judging Programming Feedback
di: Koutcheme, Charles, et al.
Pubblicazione: (2024)
di: Koutcheme, Charles, et al.
Pubblicazione: (2024)
Comparison of Static Application Security Testing Tools and Large Language Models for Repo-level Vulnerability Detection
di: Zhou, Xin, et al.
Pubblicazione: (2024)
di: Zhou, Xin, et al.
Pubblicazione: (2024)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
di: Röttger, Paul, et al.
Pubblicazione: (2023)
di: Röttger, Paul, et al.
Pubblicazione: (2023)
Metamorphic Testing for Fairness Evaluation in Large Language Models: Identifying Intersectional Bias in LLaMA and GPT
di: Reddy, Harishwar, et al.
Pubblicazione: (2025)
di: Reddy, Harishwar, et al.
Pubblicazione: (2025)
Should We Really Edit Language Models? On the Evaluation of Edited Language Models
di: Li, Qi, et al.
Pubblicazione: (2024)
di: Li, Qi, et al.
Pubblicazione: (2024)
IMU-1: Sample-Efficient Pre-training of Small Language Models
di: Grigorev, George
Pubblicazione: (2026)
di: Grigorev, George
Pubblicazione: (2026)
Test-Time Adaptation for Tactile-Vision-Language Models
di: Ye, Chuyang, et al.
Pubblicazione: (2026)
di: Ye, Chuyang, et al.
Pubblicazione: (2026)
LaajMeter: A Framework for LaaJ Evaluation
di: Ackerman, Samuel, et al.
Pubblicazione: (2025)
di: Ackerman, Samuel, et al.
Pubblicazione: (2025)
Active Testing of Large Language Models via Approximate Neyman Allocation
di: Liu, Zeli, et al.
Pubblicazione: (2026)
di: Liu, Zeli, et al.
Pubblicazione: (2026)
TMIQ: Quantifying Test and Measurement Domain Intelligence in Large Language Models
di: Olowe, Emmanuel A., et al.
Pubblicazione: (2025)
di: Olowe, Emmanuel A., et al.
Pubblicazione: (2025)
Generating Symbolic World Models via Test-time Scaling of Large Language Models
di: Yu, Zhouliang, et al.
Pubblicazione: (2025)
di: Yu, Zhouliang, et al.
Pubblicazione: (2025)
Using Combinatorial Optimization to Design a High quality LLM Solution
di: Ackerman, Samuel, et al.
Pubblicazione: (2024)
di: Ackerman, Samuel, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Fair and Reliable Organ Allocation
di: Kim, Brian Hyeongseok, et al.
Pubblicazione: (2025)
di: Kim, Brian Hyeongseok, et al.
Pubblicazione: (2025)
Metamorphic Testing of Large Language Models for Natural Language Processing
di: Cho, Steven, et al.
Pubblicazione: (2025)
di: Cho, Steven, et al.
Pubblicazione: (2025)
Enterprise Large Language Model Evaluation Benchmark
di: Wang, Liya, et al.
Pubblicazione: (2025)
di: Wang, Liya, et al.
Pubblicazione: (2025)
Large Language Models Assisting Ontology Evaluation
di: Lippolis, Anna Sofia, et al.
Pubblicazione: (2025)
di: Lippolis, Anna Sofia, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Sharper Picture of Generalization in Transformers
di: Lintilhac, Paul, et al.
Pubblicazione: (2026) -
Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
di: Ackerman, Christopher
Pubblicazione: (2026) -
Survey on Reasoning Capabilities and Accessibility of Large Language Models Using Biology-related Questions
di: Ackerman, Michael
Pubblicazione: (2024) -
Evidence for Limited Metacognition in LLMs
di: Ackerman, Christopher
Pubblicazione: (2025) -
Evaluating perturbation robustness of generative systems that use COBOL code inputs
di: Ackerman, Samuel, et al.
Pubblicazione: (2025)