How We Built Note Embeddings
Similar-fragrance search works because we turn every fragrance's scent pyramid into a vector — computed locally, from the notes themselves, not from marketing copy. Here's the actual pipeline.
When you open a fragrance page here and see "similar fragrances," the list is not curated by hand and it is not "people who bought X also bought Y." It is geometry. Every fragrance in the catalog gets a position in a 768-dimensional space, and the similar list is simply its nearest neighbors.
This post explains how those positions get computed — because the details are where the quality comes from.
The scent is the signal
The first decision was what to embed. A fragrance record carries a lot of text: brand blurbs, review snippets, year, concentration. Most of it is marketing, and marketing is noise for smell-alike search — two fragrances described as "bold, magnetic, unforgettable" can smell nothing alike.
So we embed a structured description built from the scent data only. For each fragrance we compose a plain-text profile in a fixed order: accords first — they carry the overall character — then the pyramid, role by role:
Accords: woody, aromatic, citrus
Top notes: Bergamot, Pink Pepper
Heart notes: Lavender, Geranium
Base notes: Vetiver, Cedar, Amber
The order matters. Accords describe the whole; the pyramid describes the arc from spray to drydown. Base notes like Vetiver or Vanilla dominate how a fragrance actually wears, which is the same reason our layering guide tells you to match bases, not openings.
Local model, no API bill
The text goes through nomic-embed-text, an open embedding model we run
locally through Ollama. Each profile becomes a 768-dimensional vector. Running
the model ourselves means the whole catalog — tens of thousands of fragrances
— can be embedded and re-embedded on our own hardware, on a cron, for free.
When new fragrances arrive from the catalog pipeline, a scheduled job picks up
whatever is missing a vector and fills it in.
pgvector does the search
The vectors live in Postgres next to everything else, indexed with pgvector. "Similar to this" is one SQL query: nearest neighbors by vector distance, filtered to fragrances we can actually show you. No separate vector database, no sync problems between two stores — the fragrance row and its vector commit together and get queried together.
Why this beats note overlap
The obvious approach is counting shared notes: two fragrances that both list bergamot and cedar must be similar, right? Sometimes. But note lists are inconsistent across sources — one lists "citruses," another lists bergamot, lemon, and lime separately — and overlap counting treats those as disjoint.
An embedding model has read enough of the world to know that bergamot IS a citrus, that oud and leather live near each other, that an aromatic fougère built on lavender and coumarin is a close cousin of another one even when their note lists barely intersect. The vector space absorbs the messiness of naming, and the neighbors come out smelling right.
That is the whole system: honest inputs, a local model, and one database. The next time the similar-fragrances list surprises you with something your nose agrees with, this is why.