An Analysis of Embedding Layers and Similarity Scores using Siamese Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Bingi, Yash, Yin, Yiqiao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
by: Sharma, Kartik, et al.
Published: (2025)
by: Sharma, Kartik, et al.
Published: (2025)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
Efficient Prompt Caching via Embedding Similarity
by: Zhu, Hanlin, et al.
Published: (2024)
by: Zhu, Hanlin, et al.
Published: (2024)
Refining GPT-3 Embeddings with a Siamese Structure for Technical Post Duplicate Detection
by: Wu, Xingfang, et al.
Published: (2023)
by: Wu, Xingfang, et al.
Published: (2023)
Approximate Attributions for Off-the-Shelf Siamese Transformers
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
Pseudo-Siamese Network for Planning in Target-Oriented Proactive Dialogues
by: Kang, Xinyue, et al.
Published: (2026)
by: Kang, Xinyue, et al.
Published: (2026)
Scaling Embedding Layers in Language Models
by: Yu, Da, et al.
Published: (2025)
by: Yu, Da, et al.
Published: (2025)
Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity
by: Hinostroza, Cristian, et al.
Published: (2026)
by: Hinostroza, Cristian, et al.
Published: (2026)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
Mechanistic Insights into Grokking from the Embedding Layer
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
by: Das, Sourya Dipta, et al.
Published: (2024)
by: Das, Sourya Dipta, et al.
Published: (2024)
TransGAT: Transformer-Based Graph Neural Networks for Multi-Dimensional Automated Essay Scoring
by: Aljuaid, Hind, et al.
Published: (2025)
by: Aljuaid, Hind, et al.
Published: (2025)
Can Embedding Similarity Predict Cross-Lingual Transfer? A Systematic Study on African Languages
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
by: Idris, Tewodros Kederalah, et al.
Published: (2026)
Position Information Emerges in Causal Transformers Without Positional Encodings via Similarity of Nearby Embeddings
by: Zuo, Chunsheng, et al.
Published: (2024)
by: Zuo, Chunsheng, et al.
Published: (2024)
How does Multi-Task Training Affect Transformer In-Context Capabilities? Investigations with Function Classes
by: Bhasin, Harmon, et al.
Published: (2024)
by: Bhasin, Harmon, et al.
Published: (2024)
High-Stakes Personalization: Rethinking LLM Customization for Individual Investor Decision-Making
by: Sawant, Yash Ganpat
Published: (2026)
by: Sawant, Yash Ganpat
Published: (2026)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Adaptable Embeddings Network (AEN)
by: Loosmore, Stan, et al.
Published: (2024)
by: Loosmore, Stan, et al.
Published: (2024)
Next Word Suggestion using Graph Neural Network
by: Magar, Abisha Thapa, et al.
Published: (2025)
by: Magar, Abisha Thapa, et al.
Published: (2025)
Out-of-distribution generalization via composition: a lens through induction heads in Transformers
by: Song, Jiajun, et al.
Published: (2024)
by: Song, Jiajun, et al.
Published: (2024)
Compute Where it Counts: Self Optimizing Language Models
by: Akhauri, Yash, et al.
Published: (2026)
by: Akhauri, Yash, et al.
Published: (2026)
Intelligent Neural Networks: From Layered Architectures to Graph-Organized Intelligence
by: Salomon, Antoine
Published: (2025)
by: Salomon, Antoine
Published: (2025)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
by: Mistry, Deven Mahesh, et al.
Published: (2025)
by: Mistry, Deven Mahesh, et al.
Published: (2025)
Investigating Automatic Scoring and Feedback using Large Language Models
by: Katuka, Gloria Ashiya, et al.
Published: (2024)
by: Katuka, Gloria Ashiya, et al.
Published: (2024)
Phrase-Level Adversarial Training for Mitigating Bias in Neural Network-based Automatic Essay Scoring
by: Philip, Haddad, et al.
Published: (2024)
by: Philip, Haddad, et al.
Published: (2024)
Till the Layers Collapse: Compressing a Deep Neural Network through the Lenses of Batch Normalization Layers
by: Liao, Zhu, et al.
Published: (2024)
by: Liao, Zhu, et al.
Published: (2024)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
by: Gan, Chunjing, et al.
Published: (2024)
by: Gan, Chunjing, et al.
Published: (2024)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language
by: Sharma, Yash, et al.
Published: (2024)
by: Sharma, Yash, et al.
Published: (2024)
Similarity-Distance-Magnitude Activations
by: Schmaltz, Allen
Published: (2025)
by: Schmaltz, Allen
Published: (2025)
Judgement Citation Retrieval using Contextual Similarity
by: Dasula, Akshat Mohan, et al.
Published: (2024)
by: Dasula, Akshat Mohan, et al.
Published: (2024)
The Unreasonable Effectiveness of Random Target Embeddings for Continuous-Output Neural Machine Translation
by: Tokarchuk, Evgeniia, et al.
Published: (2023)
by: Tokarchuk, Evgeniia, et al.
Published: (2023)
Attamba: Attending To Multi-Token States
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Neural Network Graph Similarity Computation Based on Graph Fusion
by: Chang, Zenghui, et al.
Published: (2025)
by: Chang, Zenghui, et al.
Published: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
by: Chang, Chi-Chih, et al.
Published: (2025)
by: Chang, Chi-Chih, et al.
Published: (2025)
TRACE: TRansformer-based Attribution using Contrastive Embeddings in LLMs
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
Explaining Text Similarity in Transformer Models
by: Vasileiou, Alexandros, et al.
Published: (2024)
by: Vasileiou, Alexandros, et al.
Published: (2024)
Similarity-Distance-Magnitude Universal Verification
by: Schmaltz, Allen
Published: (2025)
by: Schmaltz, Allen
Published: (2025)
Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity
by: Jiang, Jiachen, et al.
Published: (2024)
by: Jiang, Jiachen, et al.
Published: (2024)
StrAE: Autoencoding for Pre-Trained Embeddings using Explicit Structure
by: Opper, Mattia, et al.
Published: (2023)
by: Opper, Mattia, et al.
Published: (2023)
Similar Items
-
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
by: Sharma, Kartik, et al.
Published: (2025) -
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
by: Sengupta, Ayan, et al.
Published: (2025) -
Efficient Prompt Caching via Embedding Similarity
by: Zhu, Hanlin, et al.
Published: (2024) -
Refining GPT-3 Embeddings with a Siamese Structure for Technical Post Duplicate Detection
by: Wu, Xingfang, et al.
Published: (2023) -
Approximate Attributions for Off-the-Shelf Siamese Transformers
by: Möller, Lucas, et al.
Published: (2024)