Detecting and Steering LLMs' Empathy in Action
Fuente:
arXiv
Guardado en:
| Autor principal: | Cadile, Juan P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Detecting Data Contamination in LLMs via In-Context Learning
por: Zawalski, Michał, et al.
Publicado: (2025)
por: Zawalski, Michał, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
por: Tereshchenko, Yehor, et al.
Publicado: (2025)
por: Tereshchenko, Yehor, et al.
Publicado: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
por: Wu, Zhengxuan, et al.
Publicado: (2025)
por: Wu, Zhengxuan, et al.
Publicado: (2025)
Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
por: Bucher, Martin Juan José, et al.
Publicado: (2024)
por: Bucher, Martin Juan José, et al.
Publicado: (2024)
Action-Item-Driven Summarization of Long Meeting Transcripts
por: Golia, Logan, et al.
Publicado: (2023)
por: Golia, Logan, et al.
Publicado: (2023)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
ExpressivityBench: Can LLMs Communicate Implicitly?
por: Tint, Joshua, et al.
Publicado: (2024)
por: Tint, Joshua, et al.
Publicado: (2024)
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
por: Mueller, Felix B, et al.
Publicado: (2024)
por: Mueller, Felix B, et al.
Publicado: (2024)
Training LLMs to Recognize Hedges in Spontaneous Narratives
por: Paige, Amie J., et al.
Publicado: (2024)
por: Paige, Amie J., et al.
Publicado: (2024)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
Multilingual jailbreaking of LLMs using low-resource languages
por: Marx, Dylan, et al.
Publicado: (2026)
por: Marx, Dylan, et al.
Publicado: (2026)
From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations
por: Baskar, Sarvesh, et al.
Publicado: (2025)
por: Baskar, Sarvesh, et al.
Publicado: (2025)
Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction
por: Guo, Xu
Publicado: (2025)
por: Guo, Xu
Publicado: (2025)
LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
por: Gu, Yu, et al.
Publicado: (2024)
por: Gu, Yu, et al.
Publicado: (2024)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
por: Iqbal, Hasan, et al.
Publicado: (2024)
por: Iqbal, Hasan, et al.
Publicado: (2024)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
por: Tahir, Munief Hassan, et al.
Publicado: (2024)
por: Tahir, Munief Hassan, et al.
Publicado: (2024)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
por: Saha, Soumadeep, et al.
Publicado: (2025)
por: Saha, Soumadeep, et al.
Publicado: (2025)
Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
por: Kocbek, Primoz, et al.
Publicado: (2025)
por: Kocbek, Primoz, et al.
Publicado: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
por: Simhi, Adi, et al.
Publicado: (2026)
por: Simhi, Adi, et al.
Publicado: (2026)
Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment
por: Su, Ruoxi, et al.
Publicado: (2026)
por: Su, Ruoxi, et al.
Publicado: (2026)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2025)
por: Cho, Seonglae, et al.
Publicado: (2025)
Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
por: Paulsen, Norman
Publicado: (2025)
por: Paulsen, Norman
Publicado: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
por: Naeem, Numaan, et al.
Publicado: (2025)
por: Naeem, Numaan, et al.
Publicado: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
por: Er, Yakup Abrek, et al.
Publicado: (2025)
por: Er, Yakup Abrek, et al.
Publicado: (2025)
Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Identifying Bias in Machine-generated Text Detection
por: Stowe, Kevin, et al.
Publicado: (2025)
por: Stowe, Kevin, et al.
Publicado: (2025)
Detecting AI-Generated Texts in Cross-Domains
por: Zhou, You, et al.
Publicado: (2024)
por: Zhou, You, et al.
Publicado: (2024)
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
por: Mehenni, Gaya, et al.
Publicado: (2025)
por: Mehenni, Gaya, et al.
Publicado: (2025)
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
por: Stowe, Kevin, et al.
Publicado: (2026)
por: Stowe, Kevin, et al.
Publicado: (2026)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
por: Michail, Andrianos, et al.
Publicado: (2024)
por: Michail, Andrianos, et al.
Publicado: (2024)
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
por: Mascarell, Laura, et al.
Publicado: (2024)
por: Mascarell, Laura, et al.
Publicado: (2024)
On the Effectiveness of LLM-Specific Fine-Tuning for Detecting AI-Generated Text
por: Gromadzki, Michał, et al.
Publicado: (2026)
por: Gromadzki, Michał, et al.
Publicado: (2026)
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Historical Ink: Semantic Shift Detection for 19th Century Spanish
por: Montes, Tony, et al.
Publicado: (2024)
por: Montes, Tony, et al.
Publicado: (2024)
NER4all or Context is All You Need: Using LLMs for low-effort, high-performance NER on historical texts. A humanities informed approach
por: Hiltmann, Torsten, et al.
Publicado: (2025)
por: Hiltmann, Torsten, et al.
Publicado: (2025)
Ejemplares similares
-
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025) -
Detecting Data Contamination in LLMs via In-Context Learning
por: Zawalski, Michał, et al.
Publicado: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025) -
Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
por: Tereshchenko, Yehor, et al.
Publicado: (2025) -
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
por: Wu, Zhengxuan, et al.
Publicado: (2025)