Saved in:
Bibliographic Details
Main Authors: Rojas, Ana, Francisco M. Perez-Canales
Format: Recurso digital
Language:
Published: Zenodo 2025
Subjects:
Online Access:https://doi.org/10.5281/zenodo.17167843
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • <p><strong>FANTASIA V4 – LookUp Table – UniProt September 2025 – Experimental Evidence Code (Layer 0 Only)</strong></p> <p><strong>Release:</strong> UniProt September 2025<br><strong>System:</strong> Protein Information System (PIS v3.0.0)<br><strong>Compatibility:</strong> FANTASIA V4</p> <p> <strong>Overview</strong></p> <p>This PostgreSQL database backup uses the pgvector extension to store high-dimensional protein embeddings.<br>It contains precomputed embeddings and functional annotations from the UniProt September 2025 release, restricted to entries with experimental evidence only.</p> <p>The lookup table was generated with PIS v3.0.0 (Protein Information System), an integrated platform for automated extraction, processing, and management of protein-related data.<br>PIS consolidates information from UniProt, PDB, and GOA, enabling efficient retrieval of sequences, structures, and annotations.</p> <p>This release is designed for direct use with FANTASIA V4, an advanced pipeline for large-scale functional annotation using Protein Language Models (PLMs).<br>Unlike previous releases, this dataset <strong>includes only layer 0 embeddings</strong> from each PLM, providing a compact reference table optimized for fast similarity searches and annotation transfer.</p> <p>At runtime, FANTASIA:</p> <ul> <li> <p>Loads relevant embeddings into memory</p> </li> <li> <p>Performs high-speed nearest-neighbor searches in embedding space</p> </li> <li> <p>Transfers GO terms from experimentally annotated proteins to query sequences</p> </li> </ul> <p> <strong>Dataset Details</strong></p> <ul> <li> <p>Total proteins: 127,546</p> </li> <li> <p>Total sequences: 124,397</p> </li> <li> <p>Total embeddings: 6<em>21,585 (last layer  for 5 models)</em></p> </li> <li> <p>Total GO annotations: 627,932</p> </li> <li> <p>Proteins with experimental annotations: 127,546</p> </li> </ul> <p> <strong>Evidence Codes (GO – Experimental Only)</strong></p> <ul> <li> <p>EXP – Inferred from Experiment</p> </li> <li> <p>IDA – Inferred from Direct Assay</p> </li> <li> <p>IPI – Inferred from Physical Interaction</p> </li> <li> <p>IMP – Inferred from Mutant Phenotype</p> </li> <li> <p>IGI – Inferred from Genetic Interaction</p> </li> <li> <p>IEP – Inferred from Expression Pattern</p> </li> <li> <p>TAS – Traceable Author Statement</p> </li> <li> <p>IC – Inferred by Curator</p> </li> </ul> <p> <strong>Included Embedding Models</strong></p> <ul> <li> <p><strong>ESM-2 (650M, 34 layers: 0–33)</strong><br>Layer: 0</p> </li> <li> <p><strong>ESM3c (Cambrian 600M, 36 layers: 0–35)</strong><br>Layer: 0</p> </li> <li> <p><strong>Ankh3-Large (620M, 49 layers: 0–48)</strong><br>Layer: 0</p> </li> <li> <p><strong>ProtT5-XL-UniRef50 (~1.2B, 25 layers: 0–24)</strong><br>Layer: 0</p> </li> <li> <p><strong>ProstT5 (~1.2B, 25 layers: 0–24)</strong><br>Layer: 0</p> </li> </ul> <p>⚠️ <strong>Missing Proteins</strong></p> <p>A small subset of proteins could not be processed on Finisterrae III (CESGA) due to memory limitations with 40 GB A100 GPUs.</p> <ul> <li> <p>Affected UniProt IDs are listed in: <code>missing_embeddings.csv</code></p> </li> <li> <p>These entries are excluded from the final lookup table</p> </li> </ul>