Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses
Fuente:
arXiv
Guardado en:
| Autores principales: | Ilia, Evgenia, Aziz, Wilker |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
por: Ilia, Evgenia, et al.
Publicado: (2024)
por: Ilia, Evgenia, et al.
Publicado: (2024)
Teaching Language Models to Faithfully Express their Uncertainty
por: Eikema, Bryan, et al.
Publicado: (2025)
por: Eikema, Bryan, et al.
Publicado: (2025)
Explanation Regularisation through the Lens of Attributions
por: Ferreira, Pedro, et al.
Publicado: (2024)
por: Ferreira, Pedro, et al.
Publicado: (2024)
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
por: Ferreira, Pedro, et al.
Publicado: (2025)
por: Ferreira, Pedro, et al.
Publicado: (2025)
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
por: Mishra, Nishant, et al.
Publicado: (2025)
por: Mishra, Nishant, et al.
Publicado: (2025)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
por: Baan, Joris, et al.
Publicado: (2024)
por: Baan, Joris, et al.
Publicado: (2024)
Learning to vary: Teaching LMs to reproduce human linguistic variability in next-word prediction
por: Groot, Tobias, et al.
Publicado: (2025)
por: Groot, Tobias, et al.
Publicado: (2025)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
por: Baan, Joris, et al.
Publicado: (2026)
por: Baan, Joris, et al.
Publicado: (2026)
Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models
por: Gonen, Hila, et al.
Publicado: (2024)
por: Gonen, Hila, et al.
Publicado: (2024)
Tailoring AI-Driven Reading Scaffolds to the Distinct Needs of Neurodiverse Learners
por: Jhilal, Soufiane, et al.
Publicado: (2026)
por: Jhilal, Soufiane, et al.
Publicado: (2026)
Correlation Does Not Imply Compensation: Complexity and Irregularity in the Lexicon
por: Doucette, Amanda, et al.
Publicado: (2024)
por: Doucette, Amanda, et al.
Publicado: (2024)
Pragmatic Theories Enhance Understanding of Implied Meanings in LLMs
por: Sato, Takuma, et al.
Publicado: (2025)
por: Sato, Takuma, et al.
Publicado: (2025)
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
por: Aljaafari, Nura, et al.
Publicado: (2026)
por: Aljaafari, Nura, et al.
Publicado: (2026)
Watermarking Needs Input Repetition Masking
por: Khachaturov, David, et al.
Publicado: (2025)
por: Khachaturov, David, et al.
Publicado: (2025)
The Ability of Large Language Models to Evaluate Constraint-satisfaction in Agent Responses to Open-ended Requests
por: Madmoni, Lior, et al.
Publicado: (2024)
por: Madmoni, Lior, et al.
Publicado: (2024)
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
por: Taghavi, Zeinab Sadat, et al.
Publicado: (2025)
por: Taghavi, Zeinab Sadat, et al.
Publicado: (2025)
Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses
por: Dai, Chongyuan, et al.
Publicado: (2026)
por: Dai, Chongyuan, et al.
Publicado: (2026)
Evaluating LLMs at Detecting Errors in LLM Responses
por: Kamoi, Ryo, et al.
Publicado: (2024)
por: Kamoi, Ryo, et al.
Publicado: (2024)
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
por: Vasudev, Rakshith, et al.
Publicado: (2026)
por: Vasudev, Rakshith, et al.
Publicado: (2026)
Learning Wisdom from Errors: Promoting LLM's Continual Relation Learning through Exploiting Error Cases
por: Yin, Shaozhe, et al.
Publicado: (2025)
por: Yin, Shaozhe, et al.
Publicado: (2025)
My Words Imply Your Opinion: Reader Agent-based Propagation Enhancement for Personalized Implicit Emotion Analysis
por: Liao, Jian, et al.
Publicado: (2024)
por: Liao, Jian, et al.
Publicado: (2024)
CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models
por: Xiang, Bajian, et al.
Publicado: (2026)
por: Xiang, Bajian, et al.
Publicado: (2026)
Leveraging Open-Source Large Language Models for Native Language Identification
por: Ng, Yee Man, et al.
Publicado: (2024)
por: Ng, Yee Man, et al.
Publicado: (2024)
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
por: Brook, Joshua Wolfe, et al.
Publicado: (2025)
por: Brook, Joshua Wolfe, et al.
Publicado: (2025)
Beyond Accuracy: Investigating Error Types in GPT-4 Responses to USMLE Questions
por: Roy, Soumyadeep, et al.
Publicado: (2024)
por: Roy, Soumyadeep, et al.
Publicado: (2024)
A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning
por: Hong, Guan Zhe, et al.
Publicado: (2024)
por: Hong, Guan Zhe, et al.
Publicado: (2024)
On the Distinctive Co-occurrence Characteristics of Antonymy
por: Cao, Zhihan, et al.
Publicado: (2025)
por: Cao, Zhihan, et al.
Publicado: (2025)
Do We Need Language-Specific Fact-Checking Models? The Case of Chinese
por: Zhang, Caiqi, et al.
Publicado: (2024)
por: Zhang, Caiqi, et al.
Publicado: (2024)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
por: Ponkshe, Kaustubh, et al.
Publicado: (2025)
por: Ponkshe, Kaustubh, et al.
Publicado: (2025)
Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study
por: Alhafni, Bashar, et al.
Publicado: (2025)
por: Alhafni, Bashar, et al.
Publicado: (2025)
Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
por: Ruan, Qian, et al.
Publicado: (2024)
por: Ruan, Qian, et al.
Publicado: (2024)
Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
por: Ruan, Qian, et al.
Publicado: (2024)
por: Ruan, Qian, et al.
Publicado: (2024)
Identifying Aspects in Peer Reviews
por: Lu, Sheng, et al.
Publicado: (2025)
por: Lu, Sheng, et al.
Publicado: (2025)
Variable-Length Semantic IDs for Recommender Systems
por: Khrylchenko, Kirill
Publicado: (2026)
por: Khrylchenko, Kirill
Publicado: (2026)
Are LLMs Models of Distributional Semantics? A Case Study on Quantifiers
por: Enyan, Zhang, et al.
Publicado: (2024)
por: Enyan, Zhang, et al.
Publicado: (2024)
Automated Essay Scoring Using Grammatical Variety and Errors with Multi-Task Learning and Item Response Theory
por: Doi, Kosuke, et al.
Publicado: (2024)
por: Doi, Kosuke, et al.
Publicado: (2024)
Grammatical Error Correction for Low-Resource Languages: The Case of Zarma
por: Keita, Mamadou K., et al.
Publicado: (2024)
por: Keita, Mamadou K., et al.
Publicado: (2024)
Model Unlearning Objectives Vary for Distinct Language Functions
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Persona Switch: Mixing Distinct Perspectives in Decoding Time
por: Kim, Junseok, et al.
Publicado: (2026)
por: Kim, Junseok, et al.
Publicado: (2026)
Ejemplares similares
-
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
por: Ilia, Evgenia, et al.
Publicado: (2024) -
Teaching Language Models to Faithfully Express their Uncertainty
por: Eikema, Bryan, et al.
Publicado: (2025) -
Explanation Regularisation through the Lens of Attributions
por: Ferreira, Pedro, et al.
Publicado: (2024) -
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
por: Ferreira, Pedro, et al.
Publicado: (2025) -
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
por: Mishra, Nishant, et al.
Publicado: (2025)