What do your logits know? (The answer may surprise you!)
Fuente:
arXiv
Salvato in:
| Autori principali: | Fedzechkina, Masha, Gualdoni, Eleonora, Ramos, Rita, Williamson, Sinead |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ExpertLens: Activation steering features are highly interpretable
di: Fedzechkina, Masha, et al.
Pubblicazione: (2025)
di: Fedzechkina, Masha, et al.
Pubblicazione: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
di: Sundar, Anirudh, et al.
Pubblicazione: (2025)
di: Sundar, Anirudh, et al.
Pubblicazione: (2025)
Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning
di: Nascimento, Pedro Gomes do, et al.
Pubblicazione: (2024)
di: Nascimento, Pedro Gomes do, et al.
Pubblicazione: (2024)
How well do you know your shoalmate? The answer could be life or death
di: William Bernard Perry
Pubblicazione: (2026)
di: William Bernard Perry
Pubblicazione: (2026)
HyperTransport: Amortized Conditioning of T2I Generative Models
di: Maiorca, Valentino, et al.
Pubblicazione: (2026)
di: Maiorca, Valentino, et al.
Pubblicazione: (2026)
OntView: What you See is What you Meant
di: Bobed, Carlos, et al.
Pubblicazione: (2025)
di: Bobed, Carlos, et al.
Pubblicazione: (2025)
What role, if any, should phonics play in a middle school or high school? The answer may surprise you
di: Timothy Shanahan
Pubblicazione: (2024)
di: Timothy Shanahan
Pubblicazione: (2024)
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
di: Bašaragin, Bojana, et al.
Pubblicazione: (2024)
di: Bašaragin, Bojana, et al.
Pubblicazione: (2024)
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
di: Chen, Chacha, et al.
Pubblicazione: (2026)
di: Chen, Chacha, et al.
Pubblicazione: (2026)
But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
What killed the cat? Towards a logical formalization of curiosity (and suspense, and surprise) in narratives
di: de Saint-Cyr, Florence Dupin, et al.
Pubblicazione: (2024)
di: de Saint-Cyr, Florence Dupin, et al.
Pubblicazione: (2024)
What do professional software developers need to know to succeed in an age of Artificial Intelligence?
di: Kam, Matthew, et al.
Pubblicazione: (2025)
di: Kam, Matthew, et al.
Pubblicazione: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
di: Rodriguez, Pau, et al.
Pubblicazione: (2025)
di: Rodriguez, Pau, et al.
Pubblicazione: (2025)
From task structures to world models: What do LLMs know?
di: Yildirim, Ilker, et al.
Pubblicazione: (2023)
di: Yildirim, Ilker, et al.
Pubblicazione: (2023)
What you get is what you see: Decomposing Epistemic Planning using Functional STRIPS
di: Hu, Guang, et al.
Pubblicazione: (2019)
di: Hu, Guang, et al.
Pubblicazione: (2019)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
di: Devic, Siddartha, et al.
Pubblicazione: (2025)
di: Devic, Siddartha, et al.
Pubblicazione: (2025)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
di: Elhassan, Fay, et al.
Pubblicazione: (2025)
di: Elhassan, Fay, et al.
Pubblicazione: (2025)
Why do objects have many names? A study on word informativeness in language use and lexical systems
di: Gualdoni, Eleonora, et al.
Pubblicazione: (2024)
di: Gualdoni, Eleonora, et al.
Pubblicazione: (2024)
"Accessibility people, you go work on that thing of yours over there": Addressing Disability Inclusion in AI Product Organizations
di: Moharana, Sanika, et al.
Pubblicazione: (2025)
di: Moharana, Sanika, et al.
Pubblicazione: (2025)
Five questions and answers about artificial intelligence
di: Prieto, Alberto, et al.
Pubblicazione: (2024)
di: Prieto, Alberto, et al.
Pubblicazione: (2024)
Why is plausibility surprisingly problematic as an XAI criterion?
di: Jin, Weina, et al.
Pubblicazione: (2023)
di: Jin, Weina, et al.
Pubblicazione: (2023)
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
di: Gan, Xingwei, et al.
Pubblicazione: (2026)
di: Gan, Xingwei, et al.
Pubblicazione: (2026)
CUPID: Curating Data your Robot Loves with Influence Functions
di: Agia, Christopher, et al.
Pubblicazione: (2025)
di: Agia, Christopher, et al.
Pubblicazione: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
di: Maar, Jim, et al.
Pubblicazione: (2026)
di: Maar, Jim, et al.
Pubblicazione: (2026)
The surprising efficiency of temporal difference learning for rare event prediction
di: Cheng, Xiaoou, et al.
Pubblicazione: (2024)
di: Cheng, Xiaoou, et al.
Pubblicazione: (2024)
Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter
di: Blaschke, Verena, et al.
Pubblicazione: (2025)
di: Blaschke, Verena, et al.
Pubblicazione: (2025)
Metric assessment protocol in the context of answer fluctuation on MCQ tasks
di: Goliakova, Ekaterina, et al.
Pubblicazione: (2025)
di: Goliakova, Ekaterina, et al.
Pubblicazione: (2025)
What are you sinking? A geometric approach on attention sink
di: Ruscio, Valeria, et al.
Pubblicazione: (2025)
di: Ruscio, Valeria, et al.
Pubblicazione: (2025)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
di: Friel, Robert, et al.
Pubblicazione: (2024)
di: Friel, Robert, et al.
Pubblicazione: (2024)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
di: Kanai, Sekitoshi, et al.
Pubblicazione: (2025)
di: Kanai, Sekitoshi, et al.
Pubblicazione: (2025)
What you know or who you know? The role of intellectual and social capital in opportunity recognition
di: Ramos-Rodriguez, Antonio Rafael, et al.
Pubblicazione: (2024)
di: Ramos-Rodriguez, Antonio Rafael, et al.
Pubblicazione: (2024)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
di: Liu, Tong, et al.
Pubblicazione: (2023)
di: Liu, Tong, et al.
Pubblicazione: (2023)
What would you feel if a robot performed actions on your behalf? The case of Artificial Intelligence Review Assistant (AIRA) and the cobotization of the peer-review process
di: Martinovich, Viviana, et al.
Pubblicazione: (2023)
di: Martinovich, Viviana, et al.
Pubblicazione: (2023)
If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems
di: Chang, Jiamin, et al.
Pubblicazione: (2026)
di: Chang, Jiamin, et al.
Pubblicazione: (2026)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
di: Chang, Yapei, et al.
Pubblicazione: (2025)
di: Chang, Yapei, et al.
Pubblicazione: (2025)
Can MLLMs generate human-like feedback in grading multimodal short answers?
di: Sil, Pritam, et al.
Pubblicazione: (2024)
di: Sil, Pritam, et al.
Pubblicazione: (2024)
Online inductive learning from answer sets for efficient reinforcement learning exploration
di: Veronese, Celeste, et al.
Pubblicazione: (2025)
di: Veronese, Celeste, et al.
Pubblicazione: (2025)
Hypothetical answers to continuous queries over data streams
di: Cruz-Filipe, Luís, et al.
Pubblicazione: (2019)
di: Cruz-Filipe, Luís, et al.
Pubblicazione: (2019)
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
di: Meincke, Lennart, et al.
Pubblicazione: (2025)
di: Meincke, Lennart, et al.
Pubblicazione: (2025)
Buggy rule diagnosis for combined steps through final answer evaluation in stepwise tasks
di: van der Hoek, Gerben, et al.
Pubblicazione: (2025)
di: van der Hoek, Gerben, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ExpertLens: Activation steering features are highly interpretable
di: Fedzechkina, Masha, et al.
Pubblicazione: (2025) -
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
di: Sundar, Anirudh, et al.
Pubblicazione: (2025) -
Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning
di: Nascimento, Pedro Gomes do, et al.
Pubblicazione: (2024) -
How well do you know your shoalmate? The answer could be life or death
di: William Bernard Perry
Pubblicazione: (2026) -
HyperTransport: Amortized Conditioning of T2I Generative Models
di: Maiorca, Valentino, et al.
Pubblicazione: (2026)