Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
Fuente:
arXiv
Salvato in:
| Autori principali: | Myntti, Amanda, Kanerva, Jenna, Laippala, Veronika, Ginter, Filip |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
di: Myntti, Amanda, et al.
Pubblicazione: (2025)
di: Myntti, Amanda, et al.
Pubblicazione: (2025)
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
di: Kanerva, Jenna, et al.
Pubblicazione: (2025)
di: Kanerva, Jenna, et al.
Pubblicazione: (2025)
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs
di: Laato, Joonatan, et al.
Pubblicazione: (2025)
di: Laato, Joonatan, et al.
Pubblicazione: (2025)
Automatic register identification for the open web using multilingual deep learning
di: Henriksson, Erik, et al.
Pubblicazione: (2024)
di: Henriksson, Erik, et al.
Pubblicazione: (2024)
Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs
di: Laato, Joonatan, et al.
Pubblicazione: (2026)
di: Laato, Joonatan, et al.
Pubblicazione: (2026)
Semantic Search as Extractive Paraphrase Span Detection
di: Kanerva, Jenna, et al.
Pubblicazione: (2021)
di: Kanerva, Jenna, et al.
Pubblicazione: (2021)
Hopes and Fears -- Emotion Distribution in the Topic Landscape of Finnish Parliamentary Speech 2000-2020
di: Ristilä, Anna, et al.
Pubblicazione: (2026)
di: Ristilä, Anna, et al.
Pubblicazione: (2026)
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering
di: Henriksson, Erik, et al.
Pubblicazione: (2025)
di: Henriksson, Erik, et al.
Pubblicazione: (2025)
Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations
di: Nuutinen, Emil, et al.
Pubblicazione: (2025)
di: Nuutinen, Emil, et al.
Pubblicazione: (2025)
Creating a Historical Migration Dataset from Finnish Church Records, 1800-1920
di: Vesalainen, Ari, et al.
Pubblicazione: (2025)
di: Vesalainen, Ari, et al.
Pubblicazione: (2025)
Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke
di: Wu, Yu, et al.
Pubblicazione: (2026)
di: Wu, Yu, et al.
Pubblicazione: (2026)
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
di: Burchell, Laurie, et al.
Pubblicazione: (2025)
di: Burchell, Laurie, et al.
Pubblicazione: (2025)
ChemTEB: Chemical Text Embedding Benchmark, an Overview of Embedding Models Performance & Efficiency on a Specific Domain
di: Kasmaee, Ali Shiraee, et al.
Pubblicazione: (2024)
di: Kasmaee, Ali Shiraee, et al.
Pubblicazione: (2024)
Structured Token Retention and Computational Memory Paths in Large Language Models
di: Delena, Jonathan, et al.
Pubblicazione: (2025)
di: Delena, Jonathan, et al.
Pubblicazione: (2025)
Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space
di: Wu, Si, et al.
Pubblicazione: (2025)
di: Wu, Si, et al.
Pubblicazione: (2025)
One ruler to measure them all: Benchmarking multilingual long-context language models
di: Kim, Yekyung, et al.
Pubblicazione: (2025)
di: Kim, Yekyung, et al.
Pubblicazione: (2025)
Space Decomposition for Sentence Embedding
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2024)
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2024)
HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
di: Oepen, Stephan, et al.
Pubblicazione: (2025)
di: Oepen, Stephan, et al.
Pubblicazione: (2025)
LLM Performance Predictors are good initializers for Architecture Search
di: Jawahar, Ganesh, et al.
Pubblicazione: (2023)
di: Jawahar, Ganesh, et al.
Pubblicazione: (2023)
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2026)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2026)
GNNs as Predictors of Agentic Workflow Performances
di: Zhang, Yuanshuo, et al.
Pubblicazione: (2025)
di: Zhang, Yuanshuo, et al.
Pubblicazione: (2025)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
di: Karinshak, Elise, et al.
Pubblicazione: (2024)
di: Karinshak, Elise, et al.
Pubblicazione: (2024)
LMEB: Long-horizon Memory Embedding Benchmark
di: Zhao, Xinping, et al.
Pubblicazione: (2026)
di: Zhao, Xinping, et al.
Pubblicazione: (2026)
A Survey of Retentive Network
di: Yang, Haiqi, et al.
Pubblicazione: (2025)
di: Yang, Haiqi, et al.
Pubblicazione: (2025)
Massive Sound Embedding Benchmark (MSEB)
di: Heigold, Georg, et al.
Pubblicazione: (2026)
di: Heigold, Georg, et al.
Pubblicazione: (2026)
Scaling Down Semantic Leakage: Investigating Associative Bias in Smaller Language Models
di: Smilga, Veronika
Pubblicazione: (2025)
di: Smilga, Veronika
Pubblicazione: (2025)
EmbedGrad: Gradient-Based Prompt Optimization in Embedding Space for Large Language Models
di: Hou, Xiaoming, et al.
Pubblicazione: (2025)
di: Hou, Xiaoming, et al.
Pubblicazione: (2025)
TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding
di: Qiang, Minjie, et al.
Pubblicazione: (2026)
di: Qiang, Minjie, et al.
Pubblicazione: (2026)
Benchmarking Cross-Lingual Semantic Alignment in Multilingual Embeddings
di: Gong, Wen G.
Pubblicazione: (2025)
di: Gong, Wen G.
Pubblicazione: (2025)
PL-MTEB: Polish Massive Text Embedding Benchmark
di: Poświata, Rafał, et al.
Pubblicazione: (2024)
di: Poświata, Rafał, et al.
Pubblicazione: (2024)
CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding
di: Huo, Jiahao, et al.
Pubblicazione: (2026)
di: Huo, Jiahao, et al.
Pubblicazione: (2026)
Adjusting Interpretable Dimensions in Embedding Space with Human Judgments
di: Erk, Katrin, et al.
Pubblicazione: (2024)
di: Erk, Katrin, et al.
Pubblicazione: (2024)
Learning Complex Word Embeddings in Classical and Quantum Spaces
di: Harvey, Carys, et al.
Pubblicazione: (2024)
di: Harvey, Carys, et al.
Pubblicazione: (2024)
Lifelong Event Detection with Embedding Space Separation and Compaction
di: Qin, Chengwei, et al.
Pubblicazione: (2024)
di: Qin, Chengwei, et al.
Pubblicazione: (2024)
MIEB: Massive Image Embedding Benchmark
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
One Swallow Does Not Make a Summer: Understanding Semantic Structures in Embedding Spaces
di: Sun, Yandong, et al.
Pubblicazione: (2025)
di: Sun, Yandong, et al.
Pubblicazione: (2025)
German Text Embedding Clustering Benchmark
di: Wehrli, Silvan, et al.
Pubblicazione: (2024)
di: Wehrli, Silvan, et al.
Pubblicazione: (2024)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2024)
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2024)
Fine-Tuning Small Embeddings for Elevated Performance
di: Silwal, Biraj
Pubblicazione: (2024)
di: Silwal, Biraj
Pubblicazione: (2024)
SirLLM: Streaming Infinite Retentive LLM
di: Yao, Yao, et al.
Pubblicazione: (2024)
di: Yao, Yao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
di: Myntti, Amanda, et al.
Pubblicazione: (2025) -
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
di: Kanerva, Jenna, et al.
Pubblicazione: (2025) -
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs
di: Laato, Joonatan, et al.
Pubblicazione: (2025) -
Automatic register identification for the open web using multilingual deep learning
di: Henriksson, Erik, et al.
Pubblicazione: (2024) -
Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs
di: Laato, Joonatan, et al.
Pubblicazione: (2026)