Can We Reliably Rank Model Performance across Domains without Labeled Data?
Fuente:
arXiv
Guardado en:
| Autores principales: | Rammouz, Veronica, Gonzalez, Aaron, Cruzportillo, Carlos, Tan, Adrian, Beebe, Nicole, Rios, Anthony |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Telling Speculative Stories to Help Humans Imagine the Harms of Healthcare AI
por: Zhao, Xingmeng, et al.
Publicado: (2025)
por: Zhao, Xingmeng, et al.
Publicado: (2025)
Can We Predict Performance of Large Models across Vision-Language Tasks?
por: Zhao, Qinyu, et al.
Publicado: (2024)
por: Zhao, Qinyu, et al.
Publicado: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
Can We Trust Machine Learning? The Reliability of Features from Open-Source Speech Analysis Tools for Speech Modeling
por: Chowdhury, Tahiya, et al.
Publicado: (2025)
por: Chowdhury, Tahiya, et al.
Publicado: (2025)
Crossing Domains without Labels: Distant Supervision for Term Extraction
por: Senger, Elena, et al.
Publicado: (2025)
por: Senger, Elena, et al.
Publicado: (2025)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
por: Chatrath, Veronica, et al.
Publicado: (2024)
por: Chatrath, Veronica, et al.
Publicado: (2024)
Can We Afford The Perfect Prompt? Balancing Cost and Accuracy with the Economical Prompting Index
por: McDonald, Tyler, et al.
Publicado: (2024)
por: McDonald, Tyler, et al.
Publicado: (2024)
LLM Compression: How Far Can We Go in Balancing Size and Performance?
por: Sk, Sahil, et al.
Publicado: (2025)
por: Sk, Sahil, et al.
Publicado: (2025)
How Much of Your Data Can Suck? Thresholds for Domain Performance and Emergent Misalignment in LLMs
por: Ouyang, Jian, et al.
Publicado: (2025)
por: Ouyang, Jian, et al.
Publicado: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
por: Badawi, Abeer, et al.
Publicado: (2025)
por: Badawi, Abeer, et al.
Publicado: (2025)
Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
por: Dedhia, Bhishma, et al.
Publicado: (2025)
por: Dedhia, Bhishma, et al.
Publicado: (2025)
Can Language Models Represent the Past without Anachronism?
por: Underwood, Ted, et al.
Publicado: (2025)
por: Underwood, Ted, et al.
Publicado: (2025)
Instruction-tuned Large Language Models for Machine Translation in the Medical Domain
por: Rios, Miguel
Publicado: (2024)
por: Rios, Miguel
Publicado: (2024)
LLMs as Data Annotators: How Close Are We to Human Performance
por: Haq, Muhammad Uzair Ul, et al.
Publicado: (2025)
por: Haq, Muhammad Uzair Ul, et al.
Publicado: (2025)
Lateral Phishing With Large Language Models: A Large Organization Comparative Study
por: Bethany, Mazal, et al.
Publicado: (2024)
por: Bethany, Mazal, et al.
Publicado: (2024)
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
por: He, Jianfeng, et al.
Publicado: (2024)
por: He, Jianfeng, et al.
Publicado: (2024)
Based on Data Balancing and Model Improvement for Multi-Label Sentiment Classification Performance Enhancement
por: Su, Zijin, et al.
Publicado: (2025)
por: Su, Zijin, et al.
Publicado: (2025)
Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data
por: Li, Rumeng, et al.
Publicado: (2023)
por: Li, Rumeng, et al.
Publicado: (2023)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
por: Fang, Qingkai, et al.
Publicado: (2024)
por: Fang, Qingkai, et al.
Publicado: (2024)
How Much Can We Forget about Data Contamination?
por: Bordt, Sebastian, et al.
Publicado: (2024)
por: Bordt, Sebastian, et al.
Publicado: (2024)
Can We Evaluate Domain Adaptation Models Without Target-Domain Labels?
por: Yang, Jianfei, et al.
Publicado: (2023)
por: Yang, Jianfei, et al.
Publicado: (2023)
Can We Trust LLM Detectors?
por: Sandhan, Jivnesh, et al.
Publicado: (2026)
por: Sandhan, Jivnesh, et al.
Publicado: (2026)
How Far Can We Extract Diverse Perspectives from Large Language Models?
por: Hayati, Shirley Anugrah, et al.
Publicado: (2023)
por: Hayati, Shirley Anugrah, et al.
Publicado: (2023)
LabelCoRank: Revolutionizing Long Tail Multi-Label Classification with Co-Occurrence Reranking
por: Yan, Yan, et al.
Publicado: (2025)
por: Yan, Yan, et al.
Publicado: (2025)
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
por: Doll, Niclas, et al.
Publicado: (2026)
por: Doll, Niclas, et al.
Publicado: (2026)
Can Humans Identify Domains?
por: Barrett, Maria, et al.
Publicado: (2024)
por: Barrett, Maria, et al.
Publicado: (2024)
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
por: Wang, Shaobo, et al.
Publicado: (2025)
por: Wang, Shaobo, et al.
Publicado: (2025)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
por: Zverev, Egor, et al.
Publicado: (2024)
por: Zverev, Egor, et al.
Publicado: (2024)
Compare without Despair: Reliable Preference Evaluation with Generation Separability
por: Ghosh, Sayan, et al.
Publicado: (2024)
por: Ghosh, Sayan, et al.
Publicado: (2024)
Can We Infer Confidential Properties of Training Data from LLMs?
por: Huang, Pengrun, et al.
Publicado: (2025)
por: Huang, Pengrun, et al.
Publicado: (2025)
Ranking Large Language Models without Ground Truth
por: Dhurandhar, Amit, et al.
Publicado: (2024)
por: Dhurandhar, Amit, et al.
Publicado: (2024)
A Comprehensive Study of Gender Bias in Chemical Named Entity Recognition Models
por: Zhao, Xingmeng, et al.
Publicado: (2022)
por: Zhao, Xingmeng, et al.
Publicado: (2022)
Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data
por: Zhang, Jihong, et al.
Publicado: (2025)
por: Zhang, Jihong, et al.
Publicado: (2025)
Generative Pseudo-Labeling for Pre-Ranking with LLMs
por: Bi, Junyu, et al.
Publicado: (2026)
por: Bi, Junyu, et al.
Publicado: (2026)
Exploring the Performance of ML/DL Architectures on the MNIST-1D Dataset
por: Beebe, Michael, et al.
Publicado: (2026)
por: Beebe, Michael, et al.
Publicado: (2026)
Can We Locate and Prevent Stereotypes in LLMs?
por: D'Souza, Alex
Publicado: (2026)
por: D'Souza, Alex
Publicado: (2026)
CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?
por: Chen, Ruirui, et al.
Publicado: (2026)
por: Chen, Ruirui, et al.
Publicado: (2026)
Extracting Biomedical Entities from Noisy Audio Transcripts
por: Ebadi, Nima, et al.
Publicado: (2024)
por: Ebadi, Nima, et al.
Publicado: (2024)
Improving Expert Radiology Report Summarization by Prompting Large Language Models with a Layperson Summary
por: Zhao, Xingmeng, et al.
Publicado: (2024)
por: Zhao, Xingmeng, et al.
Publicado: (2024)
Ejemplares similares
-
Telling Speculative Stories to Help Humans Imagine the Harms of Healthcare AI
por: Zhao, Xingmeng, et al.
Publicado: (2025) -
Can We Predict Performance of Large Models across Vision-Language Tasks?
por: Zhao, Qinyu, et al.
Publicado: (2024) -
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025) -
Can We Trust Machine Learning? The Reliability of Features from Open-Source Speech Analysis Tools for Speech Modeling
por: Chowdhury, Tahiya, et al.
Publicado: (2025) -
Crossing Domains without Labels: Distant Supervision for Term Extraction
por: Senger, Elena, et al.
Publicado: (2025)