Embedding Provider Interface
Provides vector embeddings for ontology entities, as numpy matrices or pandas DataFrames, together with operations derived from them: pairwise and all-by-all similarity, nearest-neighbour search, and set-wise (best match average) comparison.
Implementations:
Ontology Lookup Service (OLS) Adapter serves precomputed embeddings for several models (see
embedding_models()), cached locally so OLS is only asked once per term and model. The cache follows the OAK--cachingpolicy (default: refresh after 1 month). OLS reports similarity as(1 + cosine) / 2; OAK converts this back to plain cosine.Any adapter that can compute ancestors (e.g.
sqlite) provides theclosuremodel, in which each term is a multi-hot vector of its reflexive ancestors. This puts classic ontology-based similarity on the same footing as learned embeddings; e.g. Jaccard over closure vectors is identical to ancestor-set Jaccard.
>>> from oaklib import get_adapter
>>> ols = get_adapter("ols:hp")
>>> ids, matrix = ols.entity_embeddings(["HP:0001159", "HP:0006101"])
>>> ols.embedding_similarity("HP:0001159", "HP:0006101", model="text-embedding-3-small_pca512")
>>> list(ols.nearest_entities_to_text("webbed fingers", limit=5))
See the Embeddings Examples notebooks for worked examples, and the
embedding-models, embeddings, nearest-entities and embedding-similarity
commands for command-line access.
- class oaklib.interfaces.embedding_provider_interface.EmbeddingProviderInterface(resource: ~oaklib.resource.OntologyResource = None, strict: bool = False, _multilingual: bool = None, autosave: bool = <factory>, exclude_owl_top_and_bottom: bool = <factory>, ontology_metamodel_mapper: ~oaklib.mappers.ontology_metadata_mapper.OntologyMetadataMapper | None = None, _converter: ~curies.api.Converter | None = None, auto_relax_axioms: bool = None, cache_lookups: bool = False, property_cache: ~oaklib.utilities.keyval_cache.KeyValCache = <factory>, _edge_index: ~oaklib.indexes.edge_index.EdgeIndex | None = None, _entailed_edge_index: ~oaklib.indexes.edge_index.EdgeIndex | None = None, _prefix_map: ~typing.Mapping[str, str] | None = None)[source]
Provides vector embeddings for entities, and operations over them.
The core method is
entity_embeddings(), which returns a numpy matrix with one row per entity. Similarity, nearest-neighbour search and set-wise (best match) comparison are derived from it, so a backend only needs to supply vectors; backends that can do more natively (e.g. server-side vector search) override the relevant methods.Each backend offers one or more named models (see
embedding_models()). Vectors from different models are not comparable.Any adapter that also implements
OboGraphInterfaceadditionally supports theclosurepseudo-model, in which each term is a multi-hot vector of its reflexive ancestors. This makes classic ontology-based similarity directly comparable with learned embeddings:>>> from oaklib import get_adapter >>> adapter = get_adapter("tests/input/go-nucleus.db") >>> ids, m = adapter.entity_embeddings(["GO:0005634", "GO:0005635"], model="closure") >>> m.shape[0] 2 >>> round(adapter.embedding_similarity("GO:0005634", "GO:0005634", model="closure"), 3) 1.0
- default_embedding_model = None
Model used when none is specified.
- closure_embedding_predicates = None
Predicates traversed by the
closuremodel; None means is_a only.
- embedding_models() List[str][source]
Names of the models this adapter can provide embeddings for.
- Returns:
list of model names
- entity_embeddings(curies: Iterable[str], model: str | None = None) Tuple[List[str], ndarray][source]
Embed a collection of entities as a matrix.
Entities for which no vector is available are dropped; the returned list gives the entity for each row, in input order.
- Parameters:
curies – entities to embed
model – model name; defaults to the adapter’s default model
- Returns:
tuple of (curies, matrix with one row per curie)
- entity_embedding(curie: str, model: str | None = None) ndarray | None[source]
Embed a single entity.
Note that for the
closuremodel the vector dimensions depend on the set of entities embedded together, so useentity_embeddings()to compare.- Parameters:
curie – entity to embed
model – model name
- Returns:
vector, or None if the entity has no embedding
- embeddings_dataframe(curies: Iterable[str], model: str | None = None) DataFrame[source]
Embed a collection of entities as a DataFrame indexed by entity.
- Parameters:
curies – entities to embed
model – model name
- Returns:
DataFrame with one row per entity and one column per dimension
- text_embedding(text: str, model: str | None = None) ndarray[source]
Embed arbitrary text, using the same space as the entity embeddings.
- Parameters:
text – text to embed
model – model name
- Returns:
vector
- embedding_similarity_matrix(subjects: Iterable[str], objects: Iterable[str] | None = None, model: str | None = None, metric: str = 'cosine') DataFrame[source]
All-by-all similarity between two sets of entities.
Entities without embeddings are omitted from the result.
- Parameters:
subjects – row entities
objects – column entities; defaults to subjects
model – model name
metric –
cosine(default) orjaccard
- Returns:
DataFrame indexed by subject, with a column per object
- embedding_similarity(subject: str, object: str, model: str | None = None, metric: str = 'cosine') float | None[source]
Similarity between a pair of entities.
- Parameters:
subject – first entity
object – second entity
model – model name
metric –
cosine(default) orjaccard
- Returns:
similarity score, or None if either entity has no embedding
- embedding_pairwise_similarity(subject: str, object: str, model: str | None = None) TermPairwiseSimilarity | None[source]
Pairwise similarity as a
TermPairwiseSimilarityobject.Only the
cosine_similarityslot is populated. This allows embedding-based scores to be used anywhere semantic similarity results are expected.- Parameters:
subject – first entity
object – second entity
model – model name
- Returns:
similarity object, or None if either entity has no embedding
- embedding_termset_similarity(subjects: List[str], objects: List[str], model: str | None = None, labels: bool = False, metric: str = 'cosine') TermSetPairwiseSimilarity[source]
Compare two sets of entities using best-match average over vector similarity.
This is the embedding analog of
SemanticSimilarityInterface.termset_pairwise_similarity(); the score of each best match is the vector similarity undermetric.- Parameters:
subjects – first set of entities (e.g. a patient’s phenotypes)
objects – second set of entities (e.g. a disease’s phenotypes)
model – model name
labels – if True, populate labels
metric –
cosine(default) orjaccard
- Returns:
set-wise similarity, with
average_scoreas the best-match average
- nearest_entities(curie: str, limit: int = 10, model: str | None = None, candidates: Iterable[str] | None = None) Iterator[Tuple[str, float]][source]
Find the entities whose embeddings are most similar to that of a given entity.
The default implementation is a brute-force comparison against
candidates(by default, all entities in the adapter). Backends with a vector index override this.- Parameters:
curie – query entity
limit – maximum number of results
model – model name
candidates – entities to search over
- Returns:
iterator of (entity, score) tuples, best first (excluding the query entity)
- nearest_entities_to_vector(vector: ndarray, limit: int = 10, model: str | None = None, candidates: Iterable[str] | None = None) Iterator[Tuple[str, float]][source]
Find the entities whose embeddings are most similar to a given vector.
- Parameters:
vector – query vector, in the space of
modellimit – maximum number of results
model – model name
candidates – entities to search over
- Returns:
iterator of (entity, score) tuples, best first
- nearest_entities_to_text(text: str, limit: int = 10, model: str | None = None, candidates: Iterable[str] | None = None) Iterator[Tuple[str, float]][source]
Find the entities whose embeddings are most similar to an embedding of some text.
- Parameters:
text – query text
limit – maximum number of results
model – model name
candidates – entities to search over
- Returns:
iterator of (entity, score) tuples, best first