Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.17344 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911605614182400 |
|---|---|
| author | Jiang, Jingzhou Tang, Yixuan Yang, Yi Tam, Kar Yan |
| author_facet | Jiang, Jingzhou Tang, Yixuan Yang, Yi Tam, Kar Yan |
| contents | When task-specific labels are not available, it becomes difficult to select an embedding model for a specific target corpus. Existing labelless measures based on kernel estimators or Gaussian mixes fail in high-dimensional space, resulting in unstable rankings. We propose a flow-based labelless representation embedding evaluation (FLARE), which utilizes normalized streams to estimate information sufficiency directly from log-likelihood and avoid distance-based density estimation. We give a finite sample boundary, indicating that the estimation error depends on the intrinsic dimension of the data manifold rather than the original embedding dimension. On 11 datasets and 8 embedders, FLARE reached Spearman's $ρ$ of 0.90 under the supervised benchmark and remained stable in high-dimensional embeddings ($d \geq 3{,}584$) as the existing labelless baseline collapsed. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_17344 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FLARE: Task-agnostic embedding model evaluation through a normalization process Jiang, Jingzhou Tang, Yixuan Yang, Yi Tam, Kar Yan Machine Learning Computation and Language When task-specific labels are not available, it becomes difficult to select an embedding model for a specific target corpus. Existing labelless measures based on kernel estimators or Gaussian mixes fail in high-dimensional space, resulting in unstable rankings. We propose a flow-based labelless representation embedding evaluation (FLARE), which utilizes normalized streams to estimate information sufficiency directly from log-likelihood and avoid distance-based density estimation. We give a finite sample boundary, indicating that the estimation error depends on the intrinsic dimension of the data manifold rather than the original embedding dimension. On 11 datasets and 8 embedders, FLARE reached Spearman's $ρ$ of 0.90 under the supervised benchmark and remained stable in high-dimensional embeddings ($d \geq 3{,}584$) as the existing labelless baseline collapsed. |
| title | FLARE: Task-agnostic embedding model evaluation through a normalization process |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2604.17344 |