What do your logits know? (The answer may surprise you!)
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Fedzechkina, Masha, Gualdoni, Eleonora, Ramos, Rita, Williamson, Sinead |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ExpertLens: Activation steering features are highly interpretable
par: Fedzechkina, Masha, et autres
Publié: (2025)
par: Fedzechkina, Masha, et autres
Publié: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
par: Sundar, Anirudh, et autres
Publié: (2025)
par: Sundar, Anirudh, et autres
Publié: (2025)
Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning
par: Nascimento, Pedro Gomes do, et autres
Publié: (2024)
par: Nascimento, Pedro Gomes do, et autres
Publié: (2024)
How well do you know your shoalmate? The answer could be life or death
par: William Bernard Perry
Publié: (2026)
par: William Bernard Perry
Publié: (2026)
HyperTransport: Amortized Conditioning of T2I Generative Models
par: Maiorca, Valentino, et autres
Publié: (2026)
par: Maiorca, Valentino, et autres
Publié: (2026)
OntView: What you See is What you Meant
par: Bobed, Carlos, et autres
Publié: (2025)
par: Bobed, Carlos, et autres
Publié: (2025)
What role, if any, should phonics play in a middle school or high school? The answer may surprise you
par: Timothy Shanahan
Publié: (2024)
par: Timothy Shanahan
Publié: (2024)
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
par: Bašaragin, Bojana, et autres
Publié: (2024)
par: Bašaragin, Bojana, et autres
Publié: (2024)
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
par: Chen, Chacha, et autres
Publié: (2026)
par: Chen, Chacha, et autres
Publié: (2026)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
par: Eshuijs, Leon, et autres
Publié: (2025)
par: Eshuijs, Leon, et autres
Publié: (2025)
What killed the cat? Towards a logical formalization of curiosity (and suspense, and surprise) in narratives
par: de Saint-Cyr, Florence Dupin, et autres
Publié: (2024)
par: de Saint-Cyr, Florence Dupin, et autres
Publié: (2024)
What do professional software developers need to know to succeed in an age of Artificial Intelligence?
par: Kam, Matthew, et autres
Publié: (2025)
par: Kam, Matthew, et autres
Publié: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
par: Rodriguez, Pau, et autres
Publié: (2025)
par: Rodriguez, Pau, et autres
Publié: (2025)
From task structures to world models: What do LLMs know?
par: Yildirim, Ilker, et autres
Publié: (2023)
par: Yildirim, Ilker, et autres
Publié: (2023)
What you get is what you see: Decomposing Epistemic Planning using Functional STRIPS
par: Hu, Guang, et autres
Publié: (2019)
par: Hu, Guang, et autres
Publié: (2019)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
par: Devic, Siddartha, et autres
Publié: (2025)
par: Devic, Siddartha, et autres
Publié: (2025)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
par: Elhassan, Fay, et autres
Publié: (2025)
par: Elhassan, Fay, et autres
Publié: (2025)
Why do objects have many names? A study on word informativeness in language use and lexical systems
par: Gualdoni, Eleonora, et autres
Publié: (2024)
par: Gualdoni, Eleonora, et autres
Publié: (2024)
"Accessibility people, you go work on that thing of yours over there": Addressing Disability Inclusion in AI Product Organizations
par: Moharana, Sanika, et autres
Publié: (2025)
par: Moharana, Sanika, et autres
Publié: (2025)
Five questions and answers about artificial intelligence
par: Prieto, Alberto, et autres
Publié: (2024)
par: Prieto, Alberto, et autres
Publié: (2024)
Why is plausibility surprisingly problematic as an XAI criterion?
par: Jin, Weina, et autres
Publié: (2023)
par: Jin, Weina, et autres
Publié: (2023)
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
par: Gan, Xingwei, et autres
Publié: (2026)
par: Gan, Xingwei, et autres
Publié: (2026)
CUPID: Curating Data your Robot Loves with Influence Functions
par: Agia, Christopher, et autres
Publié: (2025)
par: Agia, Christopher, et autres
Publié: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
par: Maar, Jim, et autres
Publié: (2026)
par: Maar, Jim, et autres
Publié: (2026)
The surprising efficiency of temporal difference learning for rare event prediction
par: Cheng, Xiaoou, et autres
Publié: (2024)
par: Cheng, Xiaoou, et autres
Publié: (2024)
Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter
par: Blaschke, Verena, et autres
Publié: (2025)
par: Blaschke, Verena, et autres
Publié: (2025)
Metric assessment protocol in the context of answer fluctuation on MCQ tasks
par: Goliakova, Ekaterina, et autres
Publié: (2025)
par: Goliakova, Ekaterina, et autres
Publié: (2025)
What are you sinking? A geometric approach on attention sink
par: Ruscio, Valeria, et autres
Publié: (2025)
par: Ruscio, Valeria, et autres
Publié: (2025)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
par: Friel, Robert, et autres
Publié: (2024)
par: Friel, Robert, et autres
Publié: (2024)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
par: Kanai, Sekitoshi, et autres
Publié: (2025)
par: Kanai, Sekitoshi, et autres
Publié: (2025)
What you know or who you know? The role of intellectual and social capital in opportunity recognition
par: Ramos-Rodriguez, Antonio Rafael, et autres
Publié: (2024)
par: Ramos-Rodriguez, Antonio Rafael, et autres
Publié: (2024)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
par: Liu, Tong, et autres
Publié: (2023)
par: Liu, Tong, et autres
Publié: (2023)
What would you feel if a robot performed actions on your behalf? The case of Artificial Intelligence Review Assistant (AIRA) and the cobotization of the peer-review process
par: Martinovich, Viviana, et autres
Publié: (2023)
par: Martinovich, Viviana, et autres
Publié: (2023)
If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems
par: Chang, Jiamin, et autres
Publié: (2026)
par: Chang, Jiamin, et autres
Publié: (2026)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
par: Chang, Yapei, et autres
Publié: (2025)
par: Chang, Yapei, et autres
Publié: (2025)
Can MLLMs generate human-like feedback in grading multimodal short answers?
par: Sil, Pritam, et autres
Publié: (2024)
par: Sil, Pritam, et autres
Publié: (2024)
Online inductive learning from answer sets for efficient reinforcement learning exploration
par: Veronese, Celeste, et autres
Publié: (2025)
par: Veronese, Celeste, et autres
Publié: (2025)
Hypothetical answers to continuous queries over data streams
par: Cruz-Filipe, Luís, et autres
Publié: (2019)
par: Cruz-Filipe, Luís, et autres
Publié: (2019)
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
par: Meincke, Lennart, et autres
Publié: (2025)
par: Meincke, Lennart, et autres
Publié: (2025)
Buggy rule diagnosis for combined steps through final answer evaluation in stepwise tasks
par: van der Hoek, Gerben, et autres
Publié: (2025)
par: van der Hoek, Gerben, et autres
Publié: (2025)
Documents similaires
-
ExpertLens: Activation steering features are highly interpretable
par: Fedzechkina, Masha, et autres
Publié: (2025) -
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
par: Sundar, Anirudh, et autres
Publié: (2025) -
Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning
par: Nascimento, Pedro Gomes do, et autres
Publié: (2024) -
How well do you know your shoalmate? The answer could be life or death
par: William Bernard Perry
Publié: (2026) -
HyperTransport: Amortized Conditioning of T2I Generative Models
par: Maiorca, Valentino, et autres
Publié: (2026)