Enregistré dans:
| Auteurs principaux: | Petrova, Nora, Burden, John |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.20813 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
par: Petrova, Nora, et autres
Publié: (2026)
par: Petrova, Nora, et autres
Publié: (2026)
Evaluating AI Evaluation: Perils and Prospects
par: Burden, John
Publié: (2024)
par: Burden, John
Publié: (2024)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
par: Petrova, Nora, et autres
Publié: (2026)
par: Petrova, Nora, et autres
Publié: (2026)
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
par: Burden, John, et autres
Publié: (2025)
par: Burden, John, et autres
Publié: (2025)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
par: Burden, John, et autres
Publié: (2025)
par: Burden, John, et autres
Publié: (2025)
Framing the Game: How Context Shapes LLM Decision-Making
par: Robinson, Isaac, et autres
Publié: (2025)
par: Robinson, Isaac, et autres
Publié: (2025)
Empirical Evaluation of the Implicit Hitting Set Approach for Weighted CSPs
par: Petrova, Aleksandra, et autres
Publié: (2025)
par: Petrova, Aleksandra, et autres
Publié: (2025)
Stress-Testing Model Specs Reveals Character Differences among Language Models
par: Zhang, Jifan, et autres
Publié: (2025)
par: Zhang, Jifan, et autres
Publié: (2025)
Conversational Complexity for Assessing Risk in Large Language Models
par: Burden, John, et autres
Publié: (2024)
par: Burden, John, et autres
Publié: (2024)
Token Alignment via Character Matching for Subword Completion
par: Athiwaratkun, Ben, et autres
Publié: (2024)
par: Athiwaratkun, Ben, et autres
Publié: (2024)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
par: Zhang, Jiawei, et autres
Publié: (2025)
par: Zhang, Jiawei, et autres
Publié: (2025)
Behavioural Analysis of Alignment Faking
par: Hadida, Nathaniel Mitrani, et autres
Publié: (2026)
par: Hadida, Nathaniel Mitrani, et autres
Publié: (2026)
From Abstract to Actionable: Pairwise Shapley Values for Explainable AI
par: Xu, Jiaxin, et autres
Publié: (2025)
par: Xu, Jiaxin, et autres
Publié: (2025)
Inferring Capabilities from Task Performance with Bayesian Triangulation
par: Burden, John, et autres
Publié: (2023)
par: Burden, John, et autres
Publié: (2023)
BehAVE: Behaviour Alignment of Video Game Encodings
par: Rašajski, Nemanja, et autres
Publié: (2024)
par: Rašajski, Nemanja, et autres
Publié: (2024)
Anytime Cooperative Implicit Hitting Set Solving
par: Rollón, Emma, et autres
Publié: (2025)
par: Rollón, Emma, et autres
Publié: (2025)
Towards Generalisable Imitation Learning Through Conditioned Transition Estimation and Online Behaviour Alignment
par: Gavenski, Nathan, et autres
Publié: (2026)
par: Gavenski, Nathan, et autres
Publié: (2026)
Evaluating Stability of Unreflective Alignment
par: Lucassen, James, et autres
Publié: (2024)
par: Lucassen, James, et autres
Publié: (2024)
From Stories to Cities to Games: A Qualitative Evaluation of Behaviour Planning
par: Abdelwahed, Mustafa F., et autres
Publié: (2026)
par: Abdelwahed, Mustafa F., et autres
Publié: (2026)
Entropy-Aware Structural Alignment for Zero-Shot Handwritten Chinese Character Recognition
par: Luo, Qiuming, et autres
Publié: (2026)
par: Luo, Qiuming, et autres
Publié: (2026)
Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
par: Gu, Zhuojun, et autres
Publié: (2025)
par: Gu, Zhuojun, et autres
Publié: (2025)
A Revealed Preference Framework for AI Alignment
par: Suleymanov, Elchin
Publié: (2026)
par: Suleymanov, Elchin
Publié: (2026)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
par: Wang, Lei, et autres
Publié: (2024)
par: Wang, Lei, et autres
Publié: (2024)
Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector Alignment
par: Haas, Nathanaël, et autres
Publié: (2026)
par: Haas, Nathanaël, et autres
Publié: (2026)
Geometric Analysis of Token Selection in Multi-Head Attention
par: Mudarisov, Timur, et autres
Publié: (2026)
par: Mudarisov, Timur, et autres
Publié: (2026)
Interactive AI Alignment: Specification, Process, and Evaluation Alignment
par: Terry, Michael, et autres
Publié: (2023)
par: Terry, Michael, et autres
Publié: (2023)
Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
par: Ali, Dalia, et autres
Publié: (2025)
par: Ali, Dalia, et autres
Publié: (2025)
Evaluating Cognitive Age Alignment in Interactive AI Agents
par: Shen, Yifan, et autres
Publié: (2026)
par: Shen, Yifan, et autres
Publié: (2026)
Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models
par: Nair, Inderjeet, et autres
Publié: (2026)
par: Nair, Inderjeet, et autres
Publié: (2026)
Learning to Align: Addressing Character Frequency Distribution Shifts in Handwritten Text Recognition
par: Kaliosis, Panagiotis, et autres
Publié: (2025)
par: Kaliosis, Panagiotis, et autres
Publié: (2025)
Behaviour Distillation
par: Lupu, Andrei, et autres
Publié: (2024)
par: Lupu, Andrei, et autres
Publié: (2024)
An Evaluation of Cultural Value Alignment in LLM
par: Sukiennik, Nicholas, et autres
Publié: (2025)
par: Sukiennik, Nicholas, et autres
Publié: (2025)
Pluralistic Off-policy Evaluation and Alignment
par: Huang, Chengkai, et autres
Publié: (2025)
par: Huang, Chengkai, et autres
Publié: (2025)
Character-Adapter: Prompt-Guided Region Control for High-Fidelity Character Customization
par: Ma, Yuhang, et autres
Publié: (2024)
par: Ma, Yuhang, et autres
Publié: (2024)
StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall
par: Wu, Yerong, et autres
Publié: (2026)
par: Wu, Yerong, et autres
Publié: (2026)
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
par: Fan, Siqi, et autres
Publié: (2025)
par: Fan, Siqi, et autres
Publié: (2025)
From Language Models over Tokens to Language Models over Characters
par: Vieira, Tim, et autres
Publié: (2024)
par: Vieira, Tim, et autres
Publié: (2024)
Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
par: Galatolo, Alessio, et autres
Publié: (2025)
par: Galatolo, Alessio, et autres
Publié: (2025)
Assessing Domain-Level Susceptibility to Emergent Misalignment from Narrow Finetuning
par: Mishra, Abhishek, et autres
Publié: (2026)
par: Mishra, Abhishek, et autres
Publié: (2026)
Can Brain Signals Reveal Inner Alignment with Human Languages?
par: Han, William, et autres
Publié: (2022)
par: Han, William, et autres
Publié: (2022)
Documents similaires
-
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
par: Petrova, Nora, et autres
Publié: (2026) -
Evaluating AI Evaluation: Perils and Prospects
par: Burden, John
Publié: (2024) -
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
par: Petrova, Nora, et autres
Publié: (2026) -
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
par: Burden, John, et autres
Publié: (2025) -
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
par: Burden, John, et autres
Publié: (2025)