Enregistré dans:
| Auteur principal: | Coronado-Blázquez, Javier |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2502.19965 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
par: Awad, Samer, et autres
Publié: (2026)
par: Awad, Samer, et autres
Publié: (2026)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
par: Mavi, John, et autres
Publié: (2024)
par: Mavi, John, et autres
Publié: (2024)
Exploring the psychology of LLMs' Moral and Legal Reasoning
par: Almeida, Guilherme F. C. F., et autres
Publié: (2023)
par: Almeida, Guilherme F. C. F., et autres
Publié: (2023)
Evaluating book summaries from internal knowledge in Large Language Models: a cross-model and semantic consistency approach
par: Coronado-Blázquez, Javier
Publié: (2025)
par: Coronado-Blázquez, Javier
Publié: (2025)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
par: Mavi, John, et autres
Publié: (2025)
par: Mavi, John, et autres
Publié: (2025)
A NLP Approach to "Review Bombing" in Metacritic PC Videogames User Ratings
par: Coronado-Blázquez, Javier
Publié: (2024)
par: Coronado-Blázquez, Javier
Publié: (2024)
Redefining "Hallucination" in LLMs: Towards a psychology-informed framework for mitigating misinformation
par: Berberette, Elijah, et autres
Publié: (2024)
par: Berberette, Elijah, et autres
Publié: (2024)
PsyMem: Fine-grained psychological alignment and Explicit Memory Control for Advanced Role-Playing LLMs
par: Cheng, Xilong, et autres
Publié: (2025)
par: Cheng, Xilong, et autres
Publié: (2025)
ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs
par: Muhamad, Wicaksono Leksono, et autres
Publié: (2026)
par: Muhamad, Wicaksono Leksono, et autres
Publié: (2026)
Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice
par: Ravenda, Federico, et autres
Publié: (2025)
par: Ravenda, Federico, et autres
Publié: (2025)
A Geometric Taxonomy of Hallucinations in LLMs
par: Marín, Javier
Publié: (2026)
par: Marín, Javier
Publié: (2026)
Empirical Characterization of Temporal Constraint Processing in LLMs
par: Marín, Javier
Publié: (2025)
par: Marín, Javier
Publié: (2025)
Unifying Ontology Construction and Semantic Alignment for Deterministic Enterprise Reasoning at Scale
par: Zhu, Hongyin
Publié: (2026)
par: Zhu, Hongyin
Publié: (2026)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
par: Pletenev, Sergey, et autres
Publié: (2025)
par: Pletenev, Sergey, et autres
Publié: (2025)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
par: Miralles-González, Pablo, et autres
Publié: (2025)
par: Miralles-González, Pablo, et autres
Publié: (2025)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
par: Sands, Brendan, et autres
Publié: (2025)
par: Sands, Brendan, et autres
Publié: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
par: Piot, Paloma, et autres
Publié: (2025)
par: Piot, Paloma, et autres
Publié: (2025)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
par: Conde, Javier, et autres
Publié: (2025)
par: Conde, Javier, et autres
Publié: (2025)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
par: Arcadinho, Samuel, et autres
Publié: (2024)
par: Arcadinho, Samuel, et autres
Publié: (2024)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
par: Maar, Jim, et autres
Publié: (2026)
par: Maar, Jim, et autres
Publié: (2026)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
par: Fu, Tairan, et autres
Publié: (2025)
par: Fu, Tairan, et autres
Publié: (2025)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
par: Tiwari, Aman, et autres
Publié: (2024)
par: Tiwari, Aman, et autres
Publié: (2024)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
par: Mushtaq, Abdullah, et autres
Publié: (2025)
par: Mushtaq, Abdullah, et autres
Publié: (2025)
CogBench: a large language model walks into a psychology lab
par: Coda-Forno, Julian, et autres
Publié: (2024)
par: Coda-Forno, Julian, et autres
Publié: (2024)
Non-Determinism of "Deterministic" LLM Settings
par: Atil, Berk, et autres
Publié: (2024)
par: Atil, Berk, et autres
Publié: (2024)
Retrieval-augmented generation in multilingual settings
par: Chirkova, Nadezhda, et autres
Publié: (2024)
par: Chirkova, Nadezhda, et autres
Publié: (2024)
Sentiment analysis and random forest to classify LLM versus human source applied to Scientific Texts
par: Sanchez-Medina, Javier J.
Publié: (2024)
par: Sanchez-Medina, Javier J.
Publié: (2024)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
par: Tie, Guiyao, et autres
Publié: (2025)
par: Tie, Guiyao, et autres
Publié: (2025)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
par: Bui, Anh Thu Maria, et autres
Publié: (2024)
par: Bui, Anh Thu Maria, et autres
Publié: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
par: Alam, Firoj, et autres
Publié: (2026)
par: Alam, Firoj, et autres
Publié: (2026)
A validity-guided workflow for robust large language model research in psychology
par: Lin, Zhicheng
Publié: (2025)
par: Lin, Zhicheng
Publié: (2025)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
par: Papi, Sara, et autres
Publié: (2025)
par: Papi, Sara, et autres
Publié: (2025)
Ranking LLMs by compression
par: Guo, Peijia, et autres
Publié: (2024)
par: Guo, Peijia, et autres
Publié: (2024)
Densing Law of LLMs
par: Xiao, Chaojun, et autres
Publié: (2024)
par: Xiao, Chaojun, et autres
Publié: (2024)
The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth
par: Lissak, Shir, et autres
Publié: (2024)
par: Lissak, Shir, et autres
Publié: (2024)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
par: Kwon, Deuksin, et autres
Publié: (2024)
par: Kwon, Deuksin, et autres
Publié: (2024)
Benchmark of stylistic variation in LLM-generated texts
par: Milička, Jiří, et autres
Publié: (2025)
par: Milička, Jiří, et autres
Publié: (2025)
Are generative AI text annotations systematically biased?
par: Stolwijk, Sjoerd B., et autres
Publié: (2025)
par: Stolwijk, Sjoerd B., et autres
Publié: (2025)
Gender Bias in LLM-generated Interview Responses
par: Kong, Haein, et autres
Publié: (2024)
par: Kong, Haein, et autres
Publié: (2024)
Raply: A profanity-mitigated rap generator
par: Bendali, Omar Manil, et autres
Publié: (2024)
par: Bendali, Omar Manil, et autres
Publié: (2024)
Documents similaires
-
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
par: Awad, Samer, et autres
Publié: (2026) -
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
par: Mavi, John, et autres
Publié: (2024) -
Exploring the psychology of LLMs' Moral and Legal Reasoning
par: Almeida, Guilherme F. C. F., et autres
Publié: (2023) -
Evaluating book summaries from internal knowledge in Large Language Models: a cross-model and semantic consistency approach
par: Coronado-Blázquez, Javier
Publié: (2025) -
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
par: Mavi, John, et autres
Publié: (2025)