How to read the wheel

Every wheel here is a chord diagram. The notes sit around the rim as short coloured arcs, grouped into 10 families: citrus, fresh, herbal, floral, fruity, sweet, spicy, amber, woody and musky. The ribbons crossing the middle are the pairings. A ribbon between two arcs means those two notes appear in the same perfume, and the thicker it is, the more perfumes they share.

There are 9,645 such pairings between the 140 notes on the accord wheel. Drawn all at once that is a texture rather than a chart, which is why the controls let you keep only the strongest, and why selecting a single note is usually the faster route to an answer. Hold one note and the wheel fades everything that does not touch it.

Families are our grouping, not the industry's. Perfumery has no single agreed taxonomy of notes, so the 10 families here are an editorial call made to give the rim some structure at a glance. The floral family alone carries 37 of the 140 notes, which tells you something about both perfume and the limits of any such grouping.

Nearly half of all perfume contains musk

musk appears in 45% of the 24,059 fragrances in this collection. Not a plurality. Nearly half. After it come bergamot at 36% and sandalwood at 33%, and the pattern holds a long way down: a small number of notes do an enormous amount of the work.

This is why the raw pairing counts are less interesting than they look. The ten heaviest ribbons on the wheel are almost all something plus musk, which tells you only that musk is everywhere. Popularity crowds out everything else, the same way the most common word in any book is "the".

The accords tell a related story from another angle. citrus is the accord most likely to lead a composition, heading the list in 35% of the perfumes it appears in, while others almost never lead at all. Woody appears in more perfumes than any other accord and still leads only a quarter of the time. Some things are written to be the subject and some are written to hold the rest up.

The pairings that beat chance

A more useful question than "which notes appear together most" is "which notes appear together more than they should". If musk is in half of everything, it will share a perfume with almost anything by accident. So compare each pairing against what you would expect if the two notes were thrown together at random, and the accidents fall away.

Do that and something satisfying happens: the list that comes back is a list of real perfumery. saffron and agarwood (oud) lead it, appearing together 5.7 times more often than chance predicts across 485 fragrances. Sage and lavender follow. Then lavender and geranium, which is the spine of the fougère, a structure named in 1882. Then saffron and leather. Then cardamom and nutmeg.

Nobody encoded any of that. It falls out of arithmetic over 24,059 ingredient lists, and it recovers accords that perfumers named and taught each other by hand over a century and a half. 616 pairings clear the bar of appearing together at least 300 times, which is the floor below which this measure starts reporting coincidence as insight.

What this cannot tell you

It cannot tell you what smells good. A thick ribbon means two notes were listed together by whoever catalogued the perfume, which is a record of convention, marketing and habit as much as of chemistry. Nothing here has smelled anything.

Listed notes are also not a formula. A perfume's published note list is a description written for the person buying it, not the brief the perfumer worked from. Quantities are absent, and a note named in the pyramid may be present in a trace or may be most of what you smell. Every chart here treats presence and absence, never dose.

The catalogue has its own shape too. It skews toward what gets released, reviewed and entered into a database, which means recent mainstream releases are represented far better than older or obscure ones. Read it as what these 24,059 listings contain, which is exactly what it is, rather than as what perfume is.

Where the data comes from

24,059 fragrances from a public Fragrantica dataset, published on Kaggle under the CC BY-NC-SA 4.0 licence. The aggregation into pairings, and the grouping of notes into 10 families, are ours. Notes appearing in fewer than a handful of perfumes are dropped, because a pairing seen twice is not a pattern.

The wheels are baked ahead of time rather than computed in the browser: a Python pipeline reads the raw dataset and writes the matrices this page loads, so nothing you do here sends a request anywhere or waits on one.