Reference-Free Rating of LLM Responses via Latent Information
Fuente:
arXiv
Saved in:
| Main Authors: | Girrbach, Leander, Su, Chi-Ping, Saanum, Tankred, Socher, Richard, Schulz, Eric, Akata, Zeynep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
by: Spohn, Philipp, et al.
Published: (2025)
by: Spohn, Philipp, et al.
Published: (2025)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
A Systematic Study of In-the-Wild Model Merging for Large Language Models
by: Hitit, Oğuz Kağan, et al.
Published: (2025)
by: Hitit, Oğuz Kağan, et al.
Published: (2025)
Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
by: von Recum, Alexander, et al.
Published: (2026)
by: von Recum, Alexander, et al.
Published: (2026)
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings
by: Spohn, Philipp, et al.
Published: (2026)
by: Spohn, Philipp, et al.
Published: (2026)
Simplifying Latent Dynamics with Softly State-Invariant World Models
by: Saanum, Tankred, et al.
Published: (2024)
by: Saanum, Tankred, et al.
Published: (2024)
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
Building, Reusing, and Generalizing Abstract Representations from Concrete Sequences
by: Wu, Shuchen, et al.
Published: (2024)
by: Wu, Shuchen, et al.
Published: (2024)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
Discovering Chunks in Neural Embeddings for Interpretability
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)
by: Girrbach, Leander, et al.
Published: (2024)
by: Girrbach, Leander, et al.
Published: (2024)
A circuit for predicting hierarchical structure in-context in Large Language Models
by: Saanum, Tankred, et al.
Published: (2025)
by: Saanum, Tankred, et al.
Published: (2025)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
SemioLLM: Evaluating Large Language Models for Diagnostic Reasoning from Unstructured Clinical Narratives in Epilepsy
by: Dani, Meghal, et al.
Published: (2024)
by: Dani, Meghal, et al.
Published: (2024)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
Internalizing LLM Reasoning via Discovery and Replay of Latent Actions
by: Shi, Zhenning, et al.
Published: (2026)
by: Shi, Zhenning, et al.
Published: (2026)
Adaptive Elicitation of Latent Information Using Natural Language
by: Wang, Jimmy, et al.
Published: (2025)
by: Wang, Jimmy, et al.
Published: (2025)
Machine Psychology
by: Hagendorff, Thilo, et al.
Published: (2023)
by: Hagendorff, Thilo, et al.
Published: (2023)
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
by: Pach, Mateusz, et al.
Published: (2026)
by: Pach, Mateusz, et al.
Published: (2026)
Next state prediction gives rise to entangled, yet compositional representations of objects
by: Saanum, Tankred, et al.
Published: (2024)
by: Saanum, Tankred, et al.
Published: (2024)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
by: Roschmann, Simon, et al.
Published: (2026)
by: Roschmann, Simon, et al.
Published: (2026)
Dynamic Latent Routing
by: Yu, Fangyuan, et al.
Published: (2026)
by: Yu, Fangyuan, et al.
Published: (2026)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
by: Demircan, Can, et al.
Published: (2024)
by: Demircan, Can, et al.
Published: (2024)
Steer LLM Latents for Hallucination Detection
by: Park, Seongheon, et al.
Published: (2025)
by: Park, Seongheon, et al.
Published: (2025)
REFA: Reference Free Alignment for multi-preference optimization
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
References Improve LLM Alignment in Non-Verifiable Domains
by: Shi, Kejian, et al.
Published: (2026)
by: Shi, Kejian, et al.
Published: (2026)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
by: Nam, Hyunji, et al.
Published: (2026)
by: Nam, Hyunji, et al.
Published: (2026)
SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
by: Mooney, James, et al.
Published: (2025)
by: Mooney, James, et al.
Published: (2025)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)
by: Karlekar, Sweta, et al.
Published: (2026)
Emotion is Not Just a Label: Latent Emotional Factors in LLM Processing
by: Reichman, Benjamin, et al.
Published: (2026)
by: Reichman, Benjamin, et al.
Published: (2026)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
Similar Items
-
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025) -
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
by: Spohn, Philipp, et al.
Published: (2025) -
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
by: Bini, Massimo, et al.
Published: (2025) -
A Systematic Study of In-the-Wild Model Merging for Large Language Models
by: Hitit, Oğuz Kağan, et al.
Published: (2025) -
Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
by: von Recum, Alexander, et al.
Published: (2026)