Salvato in:
| Autori principali: | Ward, Francis Rhys, Yang, Zejia, Jackson, Alex, Brown, Randy, Smith, Chandler, Colverd, Grace, Thomson, Louis, Douglas, Raymond, Bartak, Patrik, Rowan, Andrew |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2410.04272 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
di: Chen, Runjin, et al.
Pubblicazione: (2025)
di: Chen, Runjin, et al.
Pubblicazione: (2025)
Towards a Theory of AI Personhood
di: Ward, Francis Rhys
Pubblicazione: (2025)
di: Ward, Francis Rhys
Pubblicazione: (2025)
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
di: You, Zejia, et al.
Pubblicazione: (2026)
di: You, Zejia, et al.
Pubblicazione: (2026)
NEBULA: A National Scale Dataset for Neighbourhood-Level Urban Building Energy Modelling for England and Wales
di: Colverd, Grace, et al.
Pubblicazione: (2025)
di: Colverd, Grace, et al.
Pubblicazione: (2025)
Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
di: Yuan, Xinfeng, et al.
Pubblicazione: (2024)
di: Yuan, Xinfeng, et al.
Pubblicazione: (2024)
Large Language Models Lack Understanding of Character Composition of Words
di: Shin, Andrew, et al.
Pubblicazione: (2024)
di: Shin, Andrew, et al.
Pubblicazione: (2024)
Evaluating Computational Representations of Character: An Austen Character Similarity Benchmark
di: Yang, Funing, et al.
Pubblicazione: (2024)
di: Yang, Funing, et al.
Pubblicazione: (2024)
CharacterBench: Benchmarking Character Customization of Large Language Models
di: Zhou, Jinfeng, et al.
Pubblicazione: (2024)
di: Zhou, Jinfeng, et al.
Pubblicazione: (2024)
Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual Errors
di: Chandler, Alex, et al.
Pubblicazione: (2024)
di: Chandler, Alex, et al.
Pubblicazione: (2024)
Evaluating Prompt-Based and Fine-Tuned Approaches to Czech Anaphora Resolution
di: Stano, Patrik, et al.
Pubblicazione: (2025)
di: Stano, Patrik, et al.
Pubblicazione: (2025)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
di: Jackson, Declan, et al.
Pubblicazione: (2025)
di: Jackson, Declan, et al.
Pubblicazione: (2025)
Domain-Specific Tensor Languages
di: Bernardy, Jean-Philippe, et al.
Pubblicazione: (2023)
di: Bernardy, Jean-Philippe, et al.
Pubblicazione: (2023)
Single Character Perturbations Break LLM Alignment
di: Lin, Leon, et al.
Pubblicazione: (2024)
di: Lin, Leon, et al.
Pubblicazione: (2024)
ChainNet: Structured Metaphor and Metonymy in WordNet
di: Maudslay, Rowan Hall, et al.
Pubblicazione: (2024)
di: Maudslay, Rowan Hall, et al.
Pubblicazione: (2024)
Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations
di: Ma, Xiaoxu, et al.
Pubblicazione: (2026)
di: Ma, Xiaoxu, et al.
Pubblicazione: (2026)
Zero-Shot Satellite Image Retrieval through Joint Embeddings: Application to Crisis Response
di: Walsh, James, et al.
Pubblicazione: (2026)
di: Walsh, James, et al.
Pubblicazione: (2026)
Programming of Cellular Automata in C and C++
di: Christen, Patrik
Pubblicazione: (2024)
di: Christen, Patrik
Pubblicazione: (2024)
Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
di: Deas, Nicholas, et al.
Pubblicazione: (2025)
di: Deas, Nicholas, et al.
Pubblicazione: (2025)
Incorporating Different Verbal Cues to Improve Text-Based Computer-Delivered Health Messaging
di: Cox, Samuel Rhys
Pubblicazione: (2024)
di: Cox, Samuel Rhys
Pubblicazione: (2024)
Theory of Mind and Self-Disclosure to CUIs
di: Cox, Samuel Rhys
Pubblicazione: (2025)
di: Cox, Samuel Rhys
Pubblicazione: (2025)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
di: Williams, Kai, et al.
Pubblicazione: (2025)
di: Williams, Kai, et al.
Pubblicazione: (2025)
Higher-Order Belief in Incomplete Information MAIDs
di: Foxabbott, Jack, et al.
Pubblicazione: (2025)
di: Foxabbott, Jack, et al.
Pubblicazione: (2025)
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
di: Ahn, Jaewoo, et al.
Pubblicazione: (2024)
di: Ahn, Jaewoo, et al.
Pubblicazione: (2024)
Robust Bias Detection in MLMs and its Application to Human Trait Ratings
di: Shrestha, Ingroj, et al.
Pubblicazione: (2025)
di: Shrestha, Ingroj, et al.
Pubblicazione: (2025)
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
di: Cohen, Vanya, et al.
Pubblicazione: (2025)
di: Cohen, Vanya, et al.
Pubblicazione: (2025)
HEDS 3.0: The Human Evaluation Data Sheet Version 3.0
di: Belz, Anya, et al.
Pubblicazione: (2024)
di: Belz, Anya, et al.
Pubblicazione: (2024)
Generation, Distillation and Evaluation of Motivational Interviewing-Style Reflections with a Foundational Language Model
di: Brown, Andrew, et al.
Pubblicazione: (2024)
di: Brown, Andrew, et al.
Pubblicazione: (2024)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
di: Uzan, Omri, et al.
Pubblicazione: (2025)
di: Uzan, Omri, et al.
Pubblicazione: (2025)
BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing
di: Roy, Subhro, et al.
Pubblicazione: (2022)
di: Roy, Subhro, et al.
Pubblicazione: (2022)
Tree Species Classification using Machine Learning and 3D Tomographic SAR -- a case study in Northern Europe
di: Grace, Colverd, et al.
Pubblicazione: (2024)
di: Grace, Colverd, et al.
Pubblicazione: (2024)
LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation
di: Byun, Grace, et al.
Pubblicazione: (2025)
di: Byun, Grace, et al.
Pubblicazione: (2025)
Evaluating The Impact of Stimulus Quality in Investigations of LLM Language Performance
di: Pistotti, Timothy, et al.
Pubblicazione: (2025)
di: Pistotti, Timothy, et al.
Pubblicazione: (2025)
How Do Language Models Acquire Character-Level Information?
di: Sato, Soma, et al.
Pubblicazione: (2026)
di: Sato, Soma, et al.
Pubblicazione: (2026)
Improving Language and Modality Transfer in Translation by Character-level Modeling
di: Tsiamas, Ioannis, et al.
Pubblicazione: (2025)
di: Tsiamas, Ioannis, et al.
Pubblicazione: (2025)
Mitigating the Influence of Distractor Tasks in LMs with Prior-Aware Decoding
di: Douglas, Raymond, et al.
Pubblicazione: (2024)
di: Douglas, Raymond, et al.
Pubblicazione: (2024)
Eliciting Personality Traits in Large Language Models
di: Hilliard, Airlie, et al.
Pubblicazione: (2024)
di: Hilliard, Airlie, et al.
Pubblicazione: (2024)
Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference
di: Proebsting, Grace, et al.
Pubblicazione: (2024)
di: Proebsting, Grace, et al.
Pubblicazione: (2024)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
di: Rivera, Juan-Pablo, et al.
Pubblicazione: (2024)
di: Rivera, Juan-Pablo, et al.
Pubblicazione: (2024)
From Language Models over Tokens to Language Models over Characters
di: Vieira, Tim, et al.
Pubblicazione: (2024)
di: Vieira, Tim, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
di: van der Weij, Teun, et al.
Pubblicazione: (2024) -
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
di: Chen, Runjin, et al.
Pubblicazione: (2025) -
Towards a Theory of AI Personhood
di: Ward, Francis Rhys
Pubblicazione: (2025) -
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
di: You, Zejia, et al.
Pubblicazione: (2026) -
NEBULA: A National Scale Dataset for Neighbourhood-Level Urban Building Energy Modelling for England and Wales
di: Colverd, Grace, et al.
Pubblicazione: (2025)