Embedding Provider Interface

Provides vector embeddings for ontology entities, as numpy matrices or pandas DataFrames, together with operations derived from them: pairwise and all-by-all similarity, nearest-neighbour search, and set-wise (best match average) comparison.

Implementations:

  • Ontology Lookup Service (OLS) Adapter serves precomputed embeddings for several models (see embedding_models()), cached locally so OLS is only asked once per term and model. The cache follows the OAK --caching policy (default: refresh after 1 month). OLS reports similarity as (1 + cosine) / 2; OAK converts this back to plain cosine.

  • Any adapter that can compute ancestors (e.g. sqlite) provides the closure model, in which each term is a multi-hot vector of its reflexive ancestors. This puts classic ontology-based similarity on the same footing as learned embeddings; e.g. Jaccard over closure vectors is identical to ancestor-set Jaccard.

>>> from oaklib import get_adapter
>>> ols = get_adapter("ols:hp")  
>>> ids, matrix = ols.entity_embeddings(["HP:0001159", "HP:0006101"])  
>>> ols.embedding_similarity("HP:0001159", "HP:0006101", model="text-embedding-3-small_pca512")  
>>> list(ols.nearest_entities_to_text("webbed fingers", limit=5))  

See the Embeddings Examples notebooks for worked examples, and the embedding-models, embeddings, nearest-entities and embedding-similarity commands for command-line access.

class oaklib.interfaces.embedding_provider_interface.EmbeddingProviderInterface(resource: ~oaklib.resource.OntologyResource = None, strict: bool = False, _multilingual: bool = None, autosave: bool = <factory>, exclude_owl_top_and_bottom: bool = <factory>, ontology_metamodel_mapper: ~oaklib.mappers.ontology_metadata_mapper.OntologyMetadataMapper | None = None, _converter: ~curies.api.Converter | None = None, auto_relax_axioms: bool = None, cache_lookups: bool = False, property_cache: ~oaklib.utilities.keyval_cache.KeyValCache = <factory>, _edge_index: ~oaklib.indexes.edge_index.EdgeIndex | None = None, _entailed_edge_index: ~oaklib.indexes.edge_index.EdgeIndex | None = None, _prefix_map: ~typing.Mapping[str, str] | None = None)[source]

Provides vector embeddings for entities, and operations over them.

The core method is entity_embeddings(), which returns a numpy matrix with one row per entity. Similarity, nearest-neighbour search and set-wise (best match) comparison are derived from it, so a backend only needs to supply vectors; backends that can do more natively (e.g. server-side vector search) override the relevant methods.

Each backend offers one or more named models (see embedding_models()). Vectors from different models are not comparable.

Any adapter that also implements OboGraphInterface additionally supports the closure pseudo-model, in which each term is a multi-hot vector of its reflexive ancestors. This makes classic ontology-based similarity directly comparable with learned embeddings:

>>> from oaklib import get_adapter
>>> adapter = get_adapter("tests/input/go-nucleus.db")
>>> ids, m = adapter.entity_embeddings(["GO:0005634", "GO:0005635"], model="closure")
>>> m.shape[0]
2
>>> round(adapter.embedding_similarity("GO:0005634", "GO:0005634", model="closure"), 3)
1.0
default_embedding_model = None

Model used when none is specified.

closure_embedding_predicates = None

Predicates traversed by the closure model; None means is_a only.

embedding_models() → List[str][source]

Names of the models this adapter can provide embeddings for.

Returns:

list of model names

entity_embeddings(curies: Iterable[str], model: str | None = None) → Tuple[List[str], ndarray][source]

Embed a collection of entities as a matrix.

Entities for which no vector is available are dropped; the returned list gives the entity for each row, in input order.

Parameters:
  • curies – entities to embed

  • model – model name; defaults to the adapter’s default model

Returns:

tuple of (curies, matrix with one row per curie)

entity_embedding(curie: str, model: str | None = None) → ndarray | None[source]

Embed a single entity.

Note that for the closure model the vector dimensions depend on the set of entities embedded together, so use entity_embeddings() to compare.

Parameters:
  • curie – entity to embed

  • model – model name

Returns:

vector, or None if the entity has no embedding

embeddings_dataframe(curies: Iterable[str], model: str | None = None) → DataFrame[source]

Embed a collection of entities as a DataFrame indexed by entity.

Parameters:
  • curies – entities to embed

  • model – model name

Returns:

DataFrame with one row per entity and one column per dimension

text_embedding(text: str, model: str | None = None) → ndarray[source]

Embed arbitrary text, using the same space as the entity embeddings.

Parameters:
  • text – text to embed

  • model – model name

Returns:

vector

embedding_similarity_matrix(subjects: Iterable[str], objects: Iterable[str] | None = None, model: str | None = None, metric: str = 'cosine') → DataFrame[source]

All-by-all similarity between two sets of entities.

Entities without embeddings are omitted from the result.

Parameters:
  • subjects – row entities

  • objects – column entities; defaults to subjects

  • model – model name

  • metric – cosine (default) or jaccard

Returns:

DataFrame indexed by subject, with a column per object

embedding_similarity(subject: str, object: str, model: str | None = None, metric: str = 'cosine') → float | None[source]

Similarity between a pair of entities.

Parameters:
  • subject – first entity

  • object – second entity

  • model – model name

  • metric – cosine (default) or jaccard

Returns:

similarity score, or None if either entity has no embedding

embedding_pairwise_similarity(subject: str, object: str, model: str | None = None) → TermPairwiseSimilarity | None[source]

Pairwise similarity as a TermPairwiseSimilarity object.

Only the cosine_similarity slot is populated. This allows embedding-based scores to be used anywhere semantic similarity results are expected.

Parameters:
  • subject – first entity

  • object – second entity

  • model – model name

Returns:

similarity object, or None if either entity has no embedding

embedding_termset_similarity(subjects: List[str], objects: List[str], model: str | None = None, labels: bool = False, metric: str = 'cosine') → TermSetPairwiseSimilarity[source]

Compare two sets of entities using best-match average over vector similarity.

This is the embedding analog of SemanticSimilarityInterface.termset_pairwise_similarity(); the score of each best match is the vector similarity under metric.

Parameters:
  • subjects – first set of entities (e.g. a patient’s phenotypes)

  • objects – second set of entities (e.g. a disease’s phenotypes)

  • model – model name

  • labels – if True, populate labels

  • metric – cosine (default) or jaccard

Returns:

set-wise similarity, with average_score as the best-match average

nearest_entities(curie: str, limit: int = 10, model: str | None = None, candidates: Iterable[str] | None = None) → Iterator[Tuple[str, float]][source]

Find the entities whose embeddings are most similar to that of a given entity.

The default implementation is a brute-force comparison against candidates (by default, all entities in the adapter). Backends with a vector index override this.

Parameters:
  • curie – query entity

  • limit – maximum number of results

  • model – model name

  • candidates – entities to search over

Returns:

iterator of (entity, score) tuples, best first (excluding the query entity)

nearest_entities_to_vector(vector: ndarray, limit: int = 10, model: str | None = None, candidates: Iterable[str] | None = None) → Iterator[Tuple[str, float]][source]

Find the entities whose embeddings are most similar to a given vector.

Parameters:
  • vector – query vector, in the space of model

  • limit – maximum number of results

  • model – model name

  • candidates – entities to search over

Returns:

iterator of (entity, score) tuples, best first

nearest_entities_to_text(text: str, limit: int = 10, model: str | None = None, candidates: Iterable[str] | None = None) → Iterator[Tuple[str, float]][source]

Find the entities whose embeddings are most similar to an embedding of some text.

Parameters:
  • text – query text

  • limit – maximum number of results

  • model – model name

  • candidates – entities to search over

Returns:

iterator of (entity, score) tuples, best first