Ontology Term Embeddings in OAK

This notebook introduces the EmbeddingProviderInterface, which gives direct access to vector embeddings of ontology terms as numpy matrices and pandas DataFrames, along with operations built on them (similarity, nearest neighbours, text search, and best-match comparison of term sets).

Backends:

Adapter

Models

Notes

ols:<ontology>

precomputed LLM embeddings hosted by OLS

several models, 512-d (PCA-reduced); cached locally

llm:<adapter>

any embedding model supported by the llm library

embeds labels (or labels + definitions) on demand; cached locally

any adapter that can compute ancestors (e.g. sqlite:obo:hp)

closure

each term is a multi-hot vector of its reflexive ancestors

The closure model is how “classic” ontology semantic similarity looks when written as vectors. Jaccard similarity of closure vectors is the familiar ancestor-set Jaccard, so classic and learned embeddings can be compared with exactly the same code.

[1]:
import warnings
warnings.filterwarnings("ignore", category=UserWarning, module="eutils")
warnings.filterwarnings("ignore", category=DeprecationWarning)

import matplotlib.pyplot as plt
import seaborn as sns

# reference palette (categorical slots in fixed order; sequential blue ramp)
BLUE, ORANGE, AQUA = "#2a78d6", "#eb6834", "#1baf7a"
INK, INK2, GRID, SURFACE = "#0b0b0b", "#52514e", "#e4e3df", "#fcfcfb"
plt.rcParams.update({
    "figure.facecolor": SURFACE, "axes.facecolor": SURFACE, "savefig.facecolor": SURFACE,
    "axes.edgecolor": GRID, "axes.labelcolor": INK2, "axes.titlecolor": INK,
    "axes.grid": True, "grid.color": GRID, "grid.linewidth": 0.6,
    "axes.spines.top": False, "axes.spines.right": False,
    "xtick.color": INK2, "ytick.color": INK2, "text.color": INK,
    "font.size": 10, "axes.titlesize": 11, "figure.dpi": 110,
})
SEQUENTIAL = sns.blend_palette(["#f0efec", "#86b6ef", "#2a78d6", "#0d366b"], as_cmap=True)

Connecting to OLS

OLS serves embeddings for every class it indexes, for several models:

[2]:
from oaklib import get_adapter

ols = get_adapter("ols:hp")
ols.embedding_models()
[2]:
['harrier-oss-v1-27b_pca512',
 'llama-embed-nemotron-8b_pca512',
 'text-embedding-3-large_pca512',
 'text-embedding-3-small_pca512']

Vectors as numpy and pandas

entity_embeddings returns the CURIEs that have vectors and a matrix with one row per CURIE. Vectors are cached in a local sqlite database (under ~/.data/oaklib/embeddings), so OLS is only asked once per term and model.

[3]:
hand_terms = [
    "HP:0001159",  # Syndactyly
    "HP:0006101",  # Finger syndactyly
    "HP:0001770",  # Toe syndactyly
    "HP:0001161",  # Hand polydactyly
    "HP:0001829",  # Foot polydactyly
    "HP:0001155",  # Abnormality of the hand
    "HP:0001760",  # Abnormal foot morphology
    "HP:0000568",  # Microphthalmia
    "HP:0000478",  # Abnormality of the eye
    "HP:0001250",  # Seizure
]
ids, matrix = ols.entity_embeddings(hand_terms)
matrix.shape
[3]:
(10, 512)
[4]:
ols.embeddings_dataframe(hand_terms, model="text-embedding-3-large_pca512").iloc[:, :6].round(3)
[4]:
0 1 2 3 4 5
id
HP:0001159 0.096 -0.007 -0.201 -0.119 -0.137 0.003
HP:0006101 0.064 -0.042 -0.216 -0.085 -0.159 0.000
HP:0001770 0.091 0.001 -0.189 -0.122 -0.142 -0.009
HP:0001161 0.068 -0.057 -0.183 -0.096 -0.152 -0.024
HP:0001829 0.084 -0.004 -0.178 -0.112 -0.172 -0.007
HP:0001155 0.108 -0.077 -0.241 -0.038 -0.166 -0.016
HP:0001760 0.109 -0.001 -0.228 -0.073 -0.172 -0.020
HP:0000568 0.087 -0.016 -0.254 -0.052 -0.169 0.053
HP:0000478 0.105 -0.047 -0.241 -0.033 -0.142 0.003
HP:0001250 0.064 -0.076 -0.236 -0.020 -0.005 0.030

Pairwise similarity

Scores are cosine similarities computed locally from the vectors. OLS’s own similarity endpoints report (1 + cosine) / 2; OAK converts these back to cosine so that all scores are on the same scale (#915).

Different models disagree substantially, even on the same pair of terms:

[5]:
import pandas as pd

pairs = [
    ("HP:0001159", "HP:0006101"),  # syndactyly vs finger syndactyly (subclass)
    ("HP:0006101", "HP:0001770"),  # finger vs toe syndactyly (siblings)
    ("HP:0001159", "HP:0001161"),  # syndactyly vs hand polydactyly
    ("HP:0001159", "HP:0001250"),  # syndactyly vs seizure (unrelated)
]
labels = dict(ols.labels({t for p in pairs for t in p}))
rows = []
for model in ols.embedding_models():
    for s, o in pairs:
        rows.append({"model": model, "pair": f"{labels[s]} / {labels[o]}",
                     "cosine": ols.embedding_similarity(s, o, model=model)})
pd.DataFrame(rows).pivot(index="pair", columns="model", values="cosine").round(3)
[5]:
model harrier-oss-v1-27b_pca512 llama-embed-nemotron-8b_pca512 text-embedding-3-large_pca512 text-embedding-3-small_pca512
pair
Finger syndactyly / Toe syndactyly 0.853 0.906 0.846 0.854
Syndactyly / Finger syndactyly 0.939 0.875 0.922 0.871
Syndactyly / Hand polydactyly 0.806 0.648 0.815 0.724
Syndactyly / Seizure 0.369 0.366 0.268 0.392

Note that absolute cosine values are not comparable across models: each model has its own “background” similarity for unrelated terms. Rankings within a model are what matter.

Classic semantic similarity as vectors: the closure model

Any adapter that can compute ancestors provides the closure model. Each term is a multi-hot vector over the union of the ancestors of the terms being encoded:

[8]:
from oaklib.utilities.embeddings.closure_embeddings import closure_embeddings

hp = get_adapter("sqlite:obo:hp")
ids, vocab, m = closure_embeddings(hp, hand_terms[:3])
closure_df = pd.DataFrame(m.astype(int), index=[hp.label(i) for i in ids], columns=[hp.label(v) for v in vocab])
print(closure_df.shape)
# show the dimensions (ancestors) that are not shared by all three terms
closure_df.loc[:, closure_df.sum() < 3]
(3, 16)
[8]:
Abnormal foot morphology Toe syndactyly Abnormal toe morphology Abnormality of the lower limb Finger syndactyly
Syndactyly 0 0 0 0 0
Finger syndactyly 0 0 0 0 1
Toe syndactyly 1 1 1 1 0

Jaccard similarity over closure vectors is identical to the ancestor-set Jaccard computed by the SemanticSimilarityInterface:

[9]:
from oaklib.datamodels.vocabulary import IS_A

for s, o in pairs:
    vec = hp.embedding_similarity(s, o, model="closure", metric="jaccard")
    classic = hp.pairwise_similarity(s, o, predicates=[IS_A]).jaccard_similarity
    print(f"{labels[s]:>20} / {labels[o]:<20} vector={vec:.4f} classic={classic:.4f}")
          Syndactyly / Finger syndactyly    vector=0.9167 classic=0.9167
   Finger syndactyly / Toe syndactyly       vector=0.6875 classic=0.6875
          Syndactyly / Hand polydactyly     vector=0.4762 classic=0.4762
          Syndactyly / Seizure              vector=0.1429 classic=0.1429

Comparing the similarity structure of models

Below are all-by-all similarity matrices for the same ten terms: ontology closure (cosine) versus two OLS-hosted LLM embedding models. The closure model only sees the ontology graph; the LLM models only see the text of the terms.

[10]:
import numpy as np

short = {t: ols.label(t) for t in hand_terms}
panels = [
    ("closure (ontology)", hp, "closure"),
    ("text-embedding-3-large", ols, "text-embedding-3-large_pca512"),
    ("harrier-oss-v1-27b", ols, "harrier-oss-v1-27b_pca512"),
]
fig, axes = plt.subplots(1, 3, figsize=(16, 5.2), sharey=True)
for ax, (title, adapter, model) in zip(axes, panels):
    m = adapter.embedding_similarity_matrix(hand_terms, model=model)
    m = m.rename(index=short, columns=short)
    sns.heatmap(m, ax=ax, cmap=SEQUENTIAL, vmin=0, vmax=1, square=True,
                cbar=ax is axes[-1], cbar_kws={"label": "cosine similarity", "shrink": 0.8},
                linewidths=1, linecolor=SURFACE)
    ax.set_title(title)
    ax.grid(False)
    ax.tick_params(length=0)
fig.tight_layout()
../../_images/examples_Embeddings_Embeddings-Tutorial_18_0.png

All three put the limb terms in one block, and the eye terms and seizure apart from it. They differ in the details. The closure model scores the general terms (“Abnormality of the hand”, “Abnormal foot morphology”) as only moderately similar to their own descendants, because those descendants have many extra ancestors. The LLM models place them closer. The two LLM models also have very different “background” similarity between unrelated terms (compare the Seizure row). This is why absolute cosine thresholds do not transfer between models; rankings within a model are more meaningful.

Comparing sets of terms

embedding_termset_similarity compares two sets of terms (e.g. a patient profile and a disease profile) by best-match average, returning the same TermSetPairwiseSimilarity object as classic semantic similarity:

[11]:
patient = ["HP:0006101", "HP:0000568"]           # finger syndactyly, microphthalmia
disease = ["HP:0001159", "HP:0000478", "HP:0001250"]  # syndactyly, eye abnormality, seizure
sim = ols.embedding_termset_similarity(patient, disease, model="text-embedding-3-large_pca512", labels=True)
print("best match average:", round(sim.average_score, 3))
for bm in sim.subject_best_matches.values():
    print(f"  {bm.match_source_label} -> {bm.match_target_label} ({bm.score:.3f})")
best match average: 0.705
  Finger syndactyly -> Syndactyly (0.922)
  Microphthalmia -> Abnormality of the eye (0.703)

LLM embeddings computed on demand

The llm: adapter wraps another adapter and embeds term text with any model supported by the llm library (OpenAI, sentence-transformers via plugins, …). The text template is configurable, so you can compare label-only with label + definition embeddings. This needs an API key (or a local model plugin), so it is not run here:

llm_adapter = get_adapter("llm:sqlite:obo:hp")
llm_adapter.embedding_model_id = "text-embedding-3-small"
llm_adapter.embedding_text_template = "{label}: {definition}"
ids, matrix = llm_adapter.entity_embeddings(hand_terms)

Command line

The same operations are available from runoak:

runoak -i ols:hp embedding-models
runoak -i ols:hp embeddings .desc//p=i HP:0001155 -m text-embedding-3-large_pca512 -o hand.tsv
runoak -i ols:hp nearest-entities HP:0001159 -L 5
runoak -i ols:hp nearest-entities --text "webbed fingers"
runoak -i ols:hp embedding-similarity HP:0001159 HP:0006101 @ HP:0001770 HP:0001250
runoak -i sqlite:obo:hp embedding-similarity -m closure --metric jaccard HP:0001159 @ HP:0001770

See also