What do your logits know? (The answer may surprise you!)
Fuente:
arXiv
Guardado en:
| Autores principales: | Fedzechkina, Masha, Gualdoni, Eleonora, Ramos, Rita, Williamson, Sinead |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ExpertLens: Activation steering features are highly interpretable
por: Fedzechkina, Masha, et al.
Publicado: (2025)
por: Fedzechkina, Masha, et al.
Publicado: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
por: Sundar, Anirudh, et al.
Publicado: (2025)
por: Sundar, Anirudh, et al.
Publicado: (2025)
Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning
por: Nascimento, Pedro Gomes do, et al.
Publicado: (2024)
por: Nascimento, Pedro Gomes do, et al.
Publicado: (2024)
How well do you know your shoalmate? The answer could be life or death
por: William Bernard Perry
Publicado: (2026)
por: William Bernard Perry
Publicado: (2026)
HyperTransport: Amortized Conditioning of T2I Generative Models
por: Maiorca, Valentino, et al.
Publicado: (2026)
por: Maiorca, Valentino, et al.
Publicado: (2026)
OntView: What you See is What you Meant
por: Bobed, Carlos, et al.
Publicado: (2025)
por: Bobed, Carlos, et al.
Publicado: (2025)
What role, if any, should phonics play in a middle school or high school? The answer may surprise you
por: Timothy Shanahan
Publicado: (2024)
por: Timothy Shanahan
Publicado: (2024)
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
por: Bašaragin, Bojana, et al.
Publicado: (2024)
por: Bašaragin, Bojana, et al.
Publicado: (2024)
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
por: Chen, Chacha, et al.
Publicado: (2026)
por: Chen, Chacha, et al.
Publicado: (2026)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
por: Eshuijs, Leon, et al.
Publicado: (2025)
por: Eshuijs, Leon, et al.
Publicado: (2025)
What killed the cat? Towards a logical formalization of curiosity (and suspense, and surprise) in narratives
por: de Saint-Cyr, Florence Dupin, et al.
Publicado: (2024)
por: de Saint-Cyr, Florence Dupin, et al.
Publicado: (2024)
What do professional software developers need to know to succeed in an age of Artificial Intelligence?
por: Kam, Matthew, et al.
Publicado: (2025)
por: Kam, Matthew, et al.
Publicado: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
por: Rodriguez, Pau, et al.
Publicado: (2025)
por: Rodriguez, Pau, et al.
Publicado: (2025)
From task structures to world models: What do LLMs know?
por: Yildirim, Ilker, et al.
Publicado: (2023)
por: Yildirim, Ilker, et al.
Publicado: (2023)
What you get is what you see: Decomposing Epistemic Planning using Functional STRIPS
por: Hu, Guang, et al.
Publicado: (2019)
por: Hu, Guang, et al.
Publicado: (2019)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
por: Devic, Siddartha, et al.
Publicado: (2025)
por: Devic, Siddartha, et al.
Publicado: (2025)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
por: Elhassan, Fay, et al.
Publicado: (2025)
por: Elhassan, Fay, et al.
Publicado: (2025)
Why do objects have many names? A study on word informativeness in language use and lexical systems
por: Gualdoni, Eleonora, et al.
Publicado: (2024)
por: Gualdoni, Eleonora, et al.
Publicado: (2024)
"Accessibility people, you go work on that thing of yours over there": Addressing Disability Inclusion in AI Product Organizations
por: Moharana, Sanika, et al.
Publicado: (2025)
por: Moharana, Sanika, et al.
Publicado: (2025)
Five questions and answers about artificial intelligence
por: Prieto, Alberto, et al.
Publicado: (2024)
por: Prieto, Alberto, et al.
Publicado: (2024)
Why is plausibility surprisingly problematic as an XAI criterion?
por: Jin, Weina, et al.
Publicado: (2023)
por: Jin, Weina, et al.
Publicado: (2023)
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
por: Gan, Xingwei, et al.
Publicado: (2026)
por: Gan, Xingwei, et al.
Publicado: (2026)
CUPID: Curating Data your Robot Loves with Influence Functions
por: Agia, Christopher, et al.
Publicado: (2025)
por: Agia, Christopher, et al.
Publicado: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
por: Maar, Jim, et al.
Publicado: (2026)
por: Maar, Jim, et al.
Publicado: (2026)
The surprising efficiency of temporal difference learning for rare event prediction
por: Cheng, Xiaoou, et al.
Publicado: (2024)
por: Cheng, Xiaoou, et al.
Publicado: (2024)
Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter
por: Blaschke, Verena, et al.
Publicado: (2025)
por: Blaschke, Verena, et al.
Publicado: (2025)
Metric assessment protocol in the context of answer fluctuation on MCQ tasks
por: Goliakova, Ekaterina, et al.
Publicado: (2025)
por: Goliakova, Ekaterina, et al.
Publicado: (2025)
What are you sinking? A geometric approach on attention sink
por: Ruscio, Valeria, et al.
Publicado: (2025)
por: Ruscio, Valeria, et al.
Publicado: (2025)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
por: Friel, Robert, et al.
Publicado: (2024)
por: Friel, Robert, et al.
Publicado: (2024)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
por: Kanai, Sekitoshi, et al.
Publicado: (2025)
por: Kanai, Sekitoshi, et al.
Publicado: (2025)
What you know or who you know? The role of intellectual and social capital in opportunity recognition
por: Ramos-Rodriguez, Antonio Rafael, et al.
Publicado: (2024)
por: Ramos-Rodriguez, Antonio Rafael, et al.
Publicado: (2024)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
por: Liu, Tong, et al.
Publicado: (2023)
por: Liu, Tong, et al.
Publicado: (2023)
What would you feel if a robot performed actions on your behalf? The case of Artificial Intelligence Review Assistant (AIRA) and the cobotization of the peer-review process
por: Martinovich, Viviana, et al.
Publicado: (2023)
por: Martinovich, Viviana, et al.
Publicado: (2023)
If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems
por: Chang, Jiamin, et al.
Publicado: (2026)
por: Chang, Jiamin, et al.
Publicado: (2026)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
por: Chang, Yapei, et al.
Publicado: (2025)
por: Chang, Yapei, et al.
Publicado: (2025)
Can MLLMs generate human-like feedback in grading multimodal short answers?
por: Sil, Pritam, et al.
Publicado: (2024)
por: Sil, Pritam, et al.
Publicado: (2024)
Online inductive learning from answer sets for efficient reinforcement learning exploration
por: Veronese, Celeste, et al.
Publicado: (2025)
por: Veronese, Celeste, et al.
Publicado: (2025)
Hypothetical answers to continuous queries over data streams
por: Cruz-Filipe, Luís, et al.
Publicado: (2019)
por: Cruz-Filipe, Luís, et al.
Publicado: (2019)
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
por: Meincke, Lennart, et al.
Publicado: (2025)
por: Meincke, Lennart, et al.
Publicado: (2025)
Buggy rule diagnosis for combined steps through final answer evaluation in stepwise tasks
por: van der Hoek, Gerben, et al.
Publicado: (2025)
por: van der Hoek, Gerben, et al.
Publicado: (2025)
Ejemplares similares
-
ExpertLens: Activation steering features are highly interpretable
por: Fedzechkina, Masha, et al.
Publicado: (2025) -
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
por: Sundar, Anirudh, et al.
Publicado: (2025) -
Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning
por: Nascimento, Pedro Gomes do, et al.
Publicado: (2024) -
How well do you know your shoalmate? The answer could be life or death
por: William Bernard Perry
Publicado: (2026) -
HyperTransport: Amortized Conditioning of T2I Generative Models
por: Maiorca, Valentino, et al.
Publicado: (2026)