Evaluating Alignment of Behavioral Dispositions in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taubenfeld, Amir, Gekhman, Zorik, Nezry, Lior, Feldman, Omri, Harris, Natalie, Reddy, Shashir, Stella, Romina, Goldstein, Ariel, Croak, Marian, Matias, Yossi, Feder, Amir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Confidence Improves Self-Consistency in LLMs
von: Taubenfeld, Amir, et al.
Veröffentlicht: (2025)
von: Taubenfeld, Amir, et al.
Veröffentlicht: (2025)
Can LLMs Learn Macroeconomic Narratives from Social Media?
von: Gueta, Almog, et al.
Veröffentlicht: (2024)
von: Gueta, Almog, et al.
Veröffentlicht: (2024)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning
von: Shalev, Yuval, et al.
Veröffentlicht: (2024)
von: Shalev, Yuval, et al.
Veröffentlicht: (2024)
Systematic Biases in LLM Simulations of Debates
von: Taubenfeld, Amir, et al.
Veröffentlicht: (2024)
von: Taubenfeld, Amir, et al.
Veröffentlicht: (2024)
Causal Effect Estimation with Latent Textual Treatments
von: Feldman, Omri, et al.
Veröffentlicht: (2026)
von: Feldman, Omri, et al.
Veröffentlicht: (2026)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
von: Wagner, Eitan, et al.
Veröffentlicht: (2024)
von: Wagner, Eitan, et al.
Veröffentlicht: (2024)
SAUCE: Synchronous and Asynchronous User-Customizable Environment for Multi-Agent LLM Interaction
von: Neuberger, Shlomo, et al.
Veröffentlicht: (2024)
von: Neuberger, Shlomo, et al.
Veröffentlicht: (2024)
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
von: Peisakhovsky, Yehonatan, et al.
Veröffentlicht: (2025)
von: Peisakhovsky, Yehonatan, et al.
Veröffentlicht: (2025)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2026)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2026)
Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
von: Trainin, Itamar, et al.
Veröffentlicht: (2026)
von: Trainin, Itamar, et al.
Veröffentlicht: (2026)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
von: Calderon, Nitay, et al.
Veröffentlicht: (2026)
von: Calderon, Nitay, et al.
Veröffentlicht: (2026)
Inside-Out: Hidden Factual Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
NL-Eye: Abductive NLI for Images
von: Ventura, Mor, et al.
Veröffentlicht: (2024)
von: Ventura, Mor, et al.
Veröffentlicht: (2024)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
Multi-environment Topic Models
von: Sobhani, Dominic, et al.
Veröffentlicht: (2024)
von: Sobhani, Dominic, et al.
Veröffentlicht: (2024)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
von: Roh, Jaechul, et al.
Veröffentlicht: (2026)
von: Roh, Jaechul, et al.
Veröffentlicht: (2026)
Inelastic decay from integrability
von: Burshtein, Amir, et al.
Veröffentlicht: (2023)
von: Burshtein, Amir, et al.
Veröffentlicht: (2023)
Quantum simulation of the microscopic to macroscopic crossover using superconducting quantum impurities
von: Burshtein, Amir, et al.
Veröffentlicht: (2024)
von: Burshtein, Amir, et al.
Veröffentlicht: (2024)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
AI in Action: Accelerating Progress Towards the Sustainable Development Goals
von: Gosselink, Brigitte Hoyer, et al.
Veröffentlicht: (2024)
von: Gosselink, Brigitte Hoyer, et al.
Veröffentlicht: (2024)
Prompt Repetition Improves Non-Reasoning LLMs
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2025)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2025)
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
von: Wu, Ke, et al.
Veröffentlicht: (2026)
von: Wu, Ke, et al.
Veröffentlicht: (2026)
Measuring the Robustness of NLP Models to Domain Shifts
von: Calderon, Nitay, et al.
Veröffentlicht: (2023)
von: Calderon, Nitay, et al.
Veröffentlicht: (2023)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
Dynamic Alignment as a Statistical Survival Effect
von: Jafari, Amir
Veröffentlicht: (2026)
von: Jafari, Amir
Veröffentlicht: (2026)
Phase-Only Beam Shaping for Transmitting Array Antennas in Radar Applications
von: Maman, Lior, et al.
Veröffentlicht: (2024)
von: Maman, Lior, et al.
Veröffentlicht: (2024)
Ethnography in educational policy research
von: Feldman, Ariel
Veröffentlicht: (2026)
von: Feldman, Ariel
Veröffentlicht: (2026)
Como pano de fundo ao Império. A trajetória do Fundamento Histórico, de sua produção a sua publicação na imprensa joanina (1773-1819)
von: Ariel Feldman
Veröffentlicht: (2013)
von: Ariel Feldman
Veröffentlicht: (2013)
O IMAGINÁRIO SOCIAL DE DEMOCRACIA NO PROCESSO DE MUNICIPALIZAÇÃO DO ENSINO FUNDAMENTAL NO BRASIL (1985-1990)
von: Ariel Feldman
Veröffentlicht: (2018)
von: Ariel Feldman
Veröffentlicht: (2018)
A mesma independência: a atuação pública de um unitário pernambucano (1822-1823)
von: Ariel Feldman
Veröffentlicht: (2014)
von: Ariel Feldman
Veröffentlicht: (2014)
STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts
von: Bamberger, Zachary, et al.
Veröffentlicht: (2026)
von: Bamberger, Zachary, et al.
Veröffentlicht: (2026)
Entanglement corner dependence in two-dimensional systems: A tensor network perspective
von: Feldman, Noa, et al.
Veröffentlicht: (2025)
von: Feldman, Noa, et al.
Veröffentlicht: (2025)
Las elites y las derechas en oposición al gobierno de Pedro Castillo en Perú
von: Ariel Goldstein
Veröffentlicht: (2022)
von: Ariel Goldstein
Veröffentlicht: (2022)
La Prensa Brasileña y sus “Cruzadas Morales”: Un Análisis de los Casos del Segundo Gobierno de Getúlio Vargas y el Primer Gobierno de Lula da Silva
von: Ariel Goldstein
Veröffentlicht: (2017)
von: Ariel Goldstein
Veröffentlicht: (2017)
De la expresión corporativa a la lucha por la hegemonía: las oposiciones políticas en Argentina, Brasil y Venezuela
von: Ariel Goldstein
Veröffentlicht: (2015)
von: Ariel Goldstein
Veröffentlicht: (2015)
In-Context Representation Hijacking
von: Yona, Itay, et al.
Veröffentlicht: (2025)
von: Yona, Itay, et al.
Veröffentlicht: (2025)
Beyond Behavior: Why AI Evaluation Needs a Cognitive Revolution
von: Konigsberg, Amir
Veröffentlicht: (2026)
von: Konigsberg, Amir
Veröffentlicht: (2026)
HACK: Hallucinations Along Certainty and Knowledge Axes
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Confidence Improves Self-Consistency in LLMs
von: Taubenfeld, Amir, et al.
Veröffentlicht: (2025) -
Can LLMs Learn Macroeconomic Narratives from Social Media?
von: Gueta, Almog, et al.
Veröffentlicht: (2024) -
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024) -
Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning
von: Shalev, Yuval, et al.
Veröffentlicht: (2024) -
Systematic Biases in LLM Simulations of Debates
von: Taubenfeld, Amir, et al.
Veröffentlicht: (2024)