Guardado en:
| Autor principal: | Coronado-Blázquez, Javier |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.19965 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
por: Awad, Samer, et al.
Publicado: (2026)
por: Awad, Samer, et al.
Publicado: (2026)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
por: Mavi, John, et al.
Publicado: (2024)
por: Mavi, John, et al.
Publicado: (2024)
Exploring the psychology of LLMs' Moral and Legal Reasoning
por: Almeida, Guilherme F. C. F., et al.
Publicado: (2023)
por: Almeida, Guilherme F. C. F., et al.
Publicado: (2023)
Evaluating book summaries from internal knowledge in Large Language Models: a cross-model and semantic consistency approach
por: Coronado-Blázquez, Javier
Publicado: (2025)
por: Coronado-Blázquez, Javier
Publicado: (2025)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
por: Mavi, John, et al.
Publicado: (2025)
por: Mavi, John, et al.
Publicado: (2025)
A NLP Approach to "Review Bombing" in Metacritic PC Videogames User Ratings
por: Coronado-Blázquez, Javier
Publicado: (2024)
por: Coronado-Blázquez, Javier
Publicado: (2024)
Redefining "Hallucination" in LLMs: Towards a psychology-informed framework for mitigating misinformation
por: Berberette, Elijah, et al.
Publicado: (2024)
por: Berberette, Elijah, et al.
Publicado: (2024)
PsyMem: Fine-grained psychological alignment and Explicit Memory Control for Advanced Role-Playing LLMs
por: Cheng, Xilong, et al.
Publicado: (2025)
por: Cheng, Xilong, et al.
Publicado: (2025)
ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs
por: Muhamad, Wicaksono Leksono, et al.
Publicado: (2026)
por: Muhamad, Wicaksono Leksono, et al.
Publicado: (2026)
Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice
por: Ravenda, Federico, et al.
Publicado: (2025)
por: Ravenda, Federico, et al.
Publicado: (2025)
A Geometric Taxonomy of Hallucinations in LLMs
por: Marín, Javier
Publicado: (2026)
por: Marín, Javier
Publicado: (2026)
Empirical Characterization of Temporal Constraint Processing in LLMs
por: Marín, Javier
Publicado: (2025)
por: Marín, Javier
Publicado: (2025)
Unifying Ontology Construction and Semantic Alignment for Deterministic Enterprise Reasoning at Scale
por: Zhu, Hongyin
Publicado: (2026)
por: Zhu, Hongyin
Publicado: (2026)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
por: Pletenev, Sergey, et al.
Publicado: (2025)
por: Pletenev, Sergey, et al.
Publicado: (2025)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
por: Miralles-González, Pablo, et al.
Publicado: (2025)
por: Miralles-González, Pablo, et al.
Publicado: (2025)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
por: Sands, Brendan, et al.
Publicado: (2025)
por: Sands, Brendan, et al.
Publicado: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
por: Piot, Paloma, et al.
Publicado: (2025)
por: Piot, Paloma, et al.
Publicado: (2025)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
por: Arcadinho, Samuel, et al.
Publicado: (2024)
por: Arcadinho, Samuel, et al.
Publicado: (2024)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
por: Maar, Jim, et al.
Publicado: (2026)
por: Maar, Jim, et al.
Publicado: (2026)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
por: Tiwari, Aman, et al.
Publicado: (2024)
por: Tiwari, Aman, et al.
Publicado: (2024)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
por: Mushtaq, Abdullah, et al.
Publicado: (2025)
por: Mushtaq, Abdullah, et al.
Publicado: (2025)
CogBench: a large language model walks into a psychology lab
por: Coda-Forno, Julian, et al.
Publicado: (2024)
por: Coda-Forno, Julian, et al.
Publicado: (2024)
Non-Determinism of "Deterministic" LLM Settings
por: Atil, Berk, et al.
Publicado: (2024)
por: Atil, Berk, et al.
Publicado: (2024)
Retrieval-augmented generation in multilingual settings
por: Chirkova, Nadezhda, et al.
Publicado: (2024)
por: Chirkova, Nadezhda, et al.
Publicado: (2024)
Sentiment analysis and random forest to classify LLM versus human source applied to Scientific Texts
por: Sanchez-Medina, Javier J.
Publicado: (2024)
por: Sanchez-Medina, Javier J.
Publicado: (2024)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
por: Tie, Guiyao, et al.
Publicado: (2025)
por: Tie, Guiyao, et al.
Publicado: (2025)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
por: Bui, Anh Thu Maria, et al.
Publicado: (2024)
por: Bui, Anh Thu Maria, et al.
Publicado: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
A validity-guided workflow for robust large language model research in psychology
por: Lin, Zhicheng
Publicado: (2025)
por: Lin, Zhicheng
Publicado: (2025)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
por: Papi, Sara, et al.
Publicado: (2025)
por: Papi, Sara, et al.
Publicado: (2025)
Ranking LLMs by compression
por: Guo, Peijia, et al.
Publicado: (2024)
por: Guo, Peijia, et al.
Publicado: (2024)
Densing Law of LLMs
por: Xiao, Chaojun, et al.
Publicado: (2024)
por: Xiao, Chaojun, et al.
Publicado: (2024)
The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth
por: Lissak, Shir, et al.
Publicado: (2024)
por: Lissak, Shir, et al.
Publicado: (2024)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
por: Kwon, Deuksin, et al.
Publicado: (2024)
por: Kwon, Deuksin, et al.
Publicado: (2024)
Benchmark of stylistic variation in LLM-generated texts
por: Milička, Jiří, et al.
Publicado: (2025)
por: Milička, Jiří, et al.
Publicado: (2025)
Are generative AI text annotations systematically biased?
por: Stolwijk, Sjoerd B., et al.
Publicado: (2025)
por: Stolwijk, Sjoerd B., et al.
Publicado: (2025)
Gender Bias in LLM-generated Interview Responses
por: Kong, Haein, et al.
Publicado: (2024)
por: Kong, Haein, et al.
Publicado: (2024)
Raply: A profanity-mitigated rap generator
por: Bendali, Omar Manil, et al.
Publicado: (2024)
por: Bendali, Omar Manil, et al.
Publicado: (2024)
Ejemplares similares
-
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
por: Awad, Samer, et al.
Publicado: (2026) -
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
por: Mavi, John, et al.
Publicado: (2024) -
Exploring the psychology of LLMs' Moral and Legal Reasoning
por: Almeida, Guilherme F. C. F., et al.
Publicado: (2023) -
Evaluating book summaries from internal knowledge in Large Language Models: a cross-model and semantic consistency approach
por: Coronado-Blázquez, Javier
Publicado: (2025) -
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
por: Mavi, John, et al.
Publicado: (2025)