Isotropy Matters: Soft-ZCA Whitening of Embeddings for Semantic Code Search
Fuente:
arXiv
Saved in:
| Main Authors: | Diera, Andor, Galke, Lukas, Scherp, Ansgar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Language Models Encode Semantic Relations? Probing and Sparse Feature Analysis
by: Diera, Andor, et al.
Published: (2026)
by: Diera, Andor, et al.
Published: (2026)
Efficient Continual Learning for Small Language Models with a Discrete Key-Value Bottleneck
by: Diera, Andor, et al.
Published: (2024)
by: Diera, Andor, et al.
Published: (2024)
Memorization of Named Entities in Fine-tuned BERT Models
by: Diera, Andor, et al.
Published: (2022)
by: Diera, Andor, et al.
Published: (2022)
Are We Really Making Much Progress in Text Classification? A Comparative Review
by: Galke, Lukas, et al.
Published: (2022)
by: Galke, Lukas, et al.
Published: (2022)
Multi-View Structural Graph Summaries
by: Frank, Jonatan, et al.
Published: (2024)
by: Frank, Jonatan, et al.
Published: (2024)
Semantic Source Code Segmentation using Small and Large Language Models
by: Dahou, Abdelhalim, et al.
Published: (2025)
by: Dahou, Abdelhalim, et al.
Published: (2025)
CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents
by: Karl, Fabian, et al.
Published: (2025)
by: Karl, Fabian, et al.
Published: (2025)
POWN: Prototypical Open-World Node Classification
by: Hoffmann, Marcel, et al.
Published: (2024)
by: Hoffmann, Marcel, et al.
Published: (2024)
Gumbel-MPNN: Graph Rewiring with Gumbel-Softmax
by: Hoffmann, Marcel, et al.
Published: (2025)
by: Hoffmann, Marcel, et al.
Published: (2025)
A Transformer-based Autoregressive Decoder Architecture for Hierarchical Text Classification
by: Yousef, Younes, et al.
Published: (2025)
by: Yousef, Younes, et al.
Published: (2025)
Your Extreme Multi-label Classifier is Secretly a Hierarchical Text Classifier for Free
by: Bertalis, Nerijus, et al.
Published: (2024)
by: Bertalis, Nerijus, et al.
Published: (2024)
On the Anatomy of Real-World R Code for Static Analysis
by: Sihler, Florian, et al.
Published: (2024)
by: Sihler, Florian, et al.
Published: (2024)
Text Role Classification in Scientific Charts Using Multimodal Transformers
by: Kim, Hye Jin, et al.
Published: (2024)
by: Kim, Hye Jin, et al.
Published: (2024)
Isolating Culture Neurons in Multilingual Large Language Models
by: Namazifard, Danial, et al.
Published: (2025)
by: Namazifard, Danial, et al.
Published: (2025)
Learning and communication pressures in neural networks: Lessons from emergent communication
by: Galke, Lukas, et al.
Published: (2024)
by: Galke, Lukas, et al.
Published: (2024)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2025)
by: Tsukagoshi, Hayato, et al.
Published: (2025)
Isotropy-Optimized Contrastive Learning for Semantic Course Recommendation
by: Khreis, Ali, et al.
Published: (2026)
by: Khreis, Ali, et al.
Published: (2026)
Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences
by: Seidlmayer, Eva, et al.
Published: (2025)
by: Seidlmayer, Eva, et al.
Published: (2025)
What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure
by: Galke, Lukas, et al.
Published: (2023)
by: Galke, Lukas, et al.
Published: (2023)
Zipfian Whitening
by: Yokoi, Sho, et al.
Published: (2024)
by: Yokoi, Sho, et al.
Published: (2024)
Isotropy, Clusters, and Classifiers
by: Mickus, Timothee, et al.
Published: (2024)
by: Mickus, Timothee, et al.
Published: (2024)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
by: Nielsen, Jacob, et al.
Published: (2024)
by: Nielsen, Jacob, et al.
Published: (2024)
Chain of Summaries: Summarization Through Iterative Questioning
by: Brach, William, et al.
Published: (2025)
by: Brach, William, et al.
Published: (2025)
iN2V: Bringing Transductive Node Embeddings to Inductive Graphs
by: Lell, Nicolas, et al.
Published: (2025)
by: Lell, Nicolas, et al.
Published: (2025)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
by: Beltoft, Stine, et al.
Published: (2025)
by: Beltoft, Stine, et al.
Published: (2025)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
by: Dang, Thao Anh, et al.
Published: (2024)
by: Dang, Thao Anh, et al.
Published: (2024)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
by: Torrielli, Federico, et al.
Published: (2026)
by: Torrielli, Federico, et al.
Published: (2026)
Sõnajaht: Definition Embeddings and Semantic Search for Reverse Dictionary Creation
by: Dorkin, Aleksei, et al.
Published: (2024)
by: Dorkin, Aleksei, et al.
Published: (2024)
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
by: Barmina, Gianluca, et al.
Published: (2025)
by: Barmina, Gianluca, et al.
Published: (2025)
Whitening Reveals Cluster Commitment as the Geometric Separator of Hallucination Types
by: Korun, Matic
Published: (2026)
by: Korun, Matic
Published: (2026)
A Multi-Perspective Architecture for Semantic Code Search
by: Haldar, Rajarshi, et al.
Published: (2020)
by: Haldar, Rajarshi, et al.
Published: (2020)
Whitening Not Recommended for Classification Tasks in LLMs
by: Forooghi, Ali, et al.
Published: (2024)
by: Forooghi, Ali, et al.
Published: (2024)
SDUs DAISY: A Benchmark for Danish Culture
by: Nielsen, Jacob, et al.
Published: (2026)
by: Nielsen, Jacob, et al.
Published: (2026)
SENSE: Semantic Embedding Navigation with Soft-gated Evaluation for Retrieval-based Speculative Decoding
by: Chen, Shaowen, et al.
Published: (2026)
by: Chen, Shaowen, et al.
Published: (2026)
LLM Agents Improve Semantic Code Search
by: Jain, Sarthak, et al.
Published: (2024)
by: Jain, Sarthak, et al.
Published: (2024)
HyperAggregation: Aggregating over Graph Edges with Hypernetworks
by: Lell, Nicolas, et al.
Published: (2024)
by: Lell, Nicolas, et al.
Published: (2024)
GLaMoR: Consistency Checking of OWL Ontologies using Graph Language Models
by: Mücke, Justin, et al.
Published: (2025)
by: Mücke, Justin, et al.
Published: (2025)
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
by: Deguchi, Hiroyuki, et al.
Published: (2025)
by: Deguchi, Hiroyuki, et al.
Published: (2025)
Training Language Models to Use Prolog as a Tool
by: Mellgren, Niklas, et al.
Published: (2025)
by: Mellgren, Niklas, et al.
Published: (2025)
Similar Items
-
Do Language Models Encode Semantic Relations? Probing and Sparse Feature Analysis
by: Diera, Andor, et al.
Published: (2026) -
Efficient Continual Learning for Small Language Models with a Discrete Key-Value Bottleneck
by: Diera, Andor, et al.
Published: (2024) -
Memorization of Named Entities in Fine-tuned BERT Models
by: Diera, Andor, et al.
Published: (2022) -
Are We Really Making Much Progress in Text Classification? A Comparative Review
by: Galke, Lukas, et al.
Published: (2022) -
Multi-View Structural Graph Summaries
by: Frank, Jonatan, et al.
Published: (2024)