ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
Fuente:
arXiv
Salvato in:
| Autori principali: | Surana, Risha, Ye, Qinyuan, Swayamdipta, Swabha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Every Language Model Has a Forgery-Resistant Signature
di: Finlayson, Matthew, et al.
Pubblicazione: (2025)
di: Finlayson, Matthew, et al.
Pubblicazione: (2025)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
di: Diddee, Harshita, et al.
Pubblicazione: (2026)
di: Diddee, Harshita, et al.
Pubblicazione: (2026)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
di: Ethayarajh, Kawin, et al.
Pubblicazione: (2021)
di: Ethayarajh, Kawin, et al.
Pubblicazione: (2021)
Logits of API-Protected LLMs Leak Proprietary Information
di: Finlayson, Matthew, et al.
Pubblicazione: (2024)
di: Finlayson, Matthew, et al.
Pubblicazione: (2024)
MixRx: Predicting Drug Combination Interactions with LLMs
di: Surana, Risha, et al.
Pubblicazione: (2025)
di: Surana, Risha, et al.
Pubblicazione: (2025)
CNSight: Evaluation of Clinical Note Segmentation Tools
di: Surana, Risha, et al.
Pubblicazione: (2025)
di: Surana, Risha, et al.
Pubblicazione: (2025)
Side-by-side Comparison Amplifies Dialect Bias in Language Models
di: Kondapally, Kritee, et al.
Pubblicazione: (2026)
di: Kondapally, Kritee, et al.
Pubblicazione: (2026)
Annotating FrameNet via Structure-Conditioned Language Generation
di: Cui, Xinyue, et al.
Pubblicazione: (2024)
di: Cui, Xinyue, et al.
Pubblicazione: (2024)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
di: Xu, Xiaoyue, et al.
Pubblicazione: (2024)
di: Xu, Xiaoyue, et al.
Pubblicazione: (2024)
How Reliable is Language Model Micro-Benchmarking?
di: Yauney, Gregory, et al.
Pubblicazione: (2025)
di: Yauney, Gregory, et al.
Pubblicazione: (2025)
Compare without Despair: Reliable Preference Evaluation with Generation Separability
di: Ghosh, Sayan, et al.
Pubblicazione: (2024)
di: Ghosh, Sayan, et al.
Pubblicazione: (2024)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
di: Kaplan, Guy, et al.
Pubblicazione: (2026)
di: Kaplan, Guy, et al.
Pubblicazione: (2026)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
di: Guo, Yipin, et al.
Pubblicazione: (2024)
di: Guo, Yipin, et al.
Pubblicazione: (2024)
Design and Realization of a Benchmarking Testbed for Evaluating Autonomous Platooning Algorithms
di: Shaham, Michael, et al.
Pubblicazione: (2024)
di: Shaham, Michael, et al.
Pubblicazione: (2024)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
di: Ye, Qinyuan, et al.
Pubblicazione: (2025)
di: Ye, Qinyuan, et al.
Pubblicazione: (2025)
Disentangling Geometry, Performance, and Training in Language Models
di: Kulkarni, Atharva, et al.
Pubblicazione: (2026)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2026)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
di: Wang, Siyin, et al.
Pubblicazione: (2024)
di: Wang, Siyin, et al.
Pubblicazione: (2024)
The Emergence of Social Science of Large Language Models
di: Jia, Xiao, et al.
Pubblicazione: (2025)
di: Jia, Xiao, et al.
Pubblicazione: (2025)
Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification
di: Liu, Joseph, et al.
Pubblicazione: (2025)
di: Liu, Joseph, et al.
Pubblicazione: (2025)
The Emergence of Altruism in Large-Language-Model Agents Society
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
Command A: An Enterprise-Ready Large Language Model
di: Cohere, Team, et al.
Pubblicazione: (2025)
di: Cohere, Team, et al.
Pubblicazione: (2025)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
Quantifying Semantic Emergence in Language Models
di: Chen, Hang, et al.
Pubblicazione: (2024)
di: Chen, Hang, et al.
Pubblicazione: (2024)
Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
di: Zhao, Jianpeng, et al.
Pubblicazione: (2025)
di: Zhao, Jianpeng, et al.
Pubblicazione: (2025)
Prompt Engineering a Prompt Engineer
di: Ye, Qinyuan, et al.
Pubblicazione: (2023)
di: Ye, Qinyuan, et al.
Pubblicazione: (2023)
Efficient Fusion and Task Guided Embedding for End-to-end Autonomous Driving
di: Guo, Yipin, et al.
Pubblicazione: (2024)
di: Guo, Yipin, et al.
Pubblicazione: (2024)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
di: Cui, Xinyue, et al.
Pubblicazione: (2025)
di: Cui, Xinyue, et al.
Pubblicazione: (2025)
Street-Level AI: Are Large Language Models Ready for Real-World Judgments?
di: Pokharel, Gaurab, et al.
Pubblicazione: (2025)
di: Pokharel, Gaurab, et al.
Pubblicazione: (2025)
Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering
di: Sammour, Farouq, et al.
Pubblicazione: (2024)
di: Sammour, Farouq, et al.
Pubblicazione: (2024)
AI Data Readiness Inspector (AIDRIN) for Quantitative Assessment of Data Readiness for AI
di: Hiniduma, Kaveen, et al.
Pubblicazione: (2024)
di: Hiniduma, Kaveen, et al.
Pubblicazione: (2024)
Structured Program Synthesis using LLMs: Results and Insights from the IPARC Challenge
di: Surana, Shraddha, et al.
Pubblicazione: (2025)
di: Surana, Shraddha, et al.
Pubblicazione: (2025)
Is Your VLM for Autonomous Driving Safety-Ready? A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
di: Meng, Xianhui, et al.
Pubblicazione: (2025)
di: Meng, Xianhui, et al.
Pubblicazione: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
di: Juvekar, Kush, et al.
Pubblicazione: (2025)
di: Juvekar, Kush, et al.
Pubblicazione: (2025)
Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
di: Jia, Zheng, et al.
Pubblicazione: (2025)
di: Jia, Zheng, et al.
Pubblicazione: (2025)
Ready2Unlearn: A Learning-Time Approach for Preparing Models with Future Unlearning Readiness
di: Duan, Hanyu, et al.
Pubblicazione: (2025)
di: Duan, Hanyu, et al.
Pubblicazione: (2025)
Ethics Readiness of Artificial Intelligence: A Practical Evaluation Method
di: Adomaitis, Laurynas, et al.
Pubblicazione: (2025)
di: Adomaitis, Laurynas, et al.
Pubblicazione: (2025)
Emergence of Human to Robot Transfer in Vision-Language-Action Models
di: Kareer, Simar, et al.
Pubblicazione: (2025)
di: Kareer, Simar, et al.
Pubblicazione: (2025)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
di: Lyu, Xinxi, et al.
Pubblicazione: (2024)
di: Lyu, Xinxi, et al.
Pubblicazione: (2024)
The Emergence of Abstract Thought in Large Language Models Beyond Any Language
di: Chen, Yuxin, et al.
Pubblicazione: (2025)
di: Chen, Yuxin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Every Language Model Has a Forgery-Resistant Signature
di: Finlayson, Matthew, et al.
Pubblicazione: (2025) -
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
di: Diddee, Harshita, et al.
Pubblicazione: (2026) -
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
di: Ethayarajh, Kawin, et al.
Pubblicazione: (2021) -
Logits of API-Protected LLMs Leak Proprietary Information
di: Finlayson, Matthew, et al.
Pubblicazione: (2024) -
MixRx: Predicting Drug Combination Interactions with LLMs
di: Surana, Risha, et al.
Pubblicazione: (2025)