Skip to content
SciStack
Tool Python Beginner 30 min

seaborn from the ground up: three cell lines in four views

Afterwards you can compare groups and distributions from a DataFrame with seaborn, choose between box, violin, strip, and ECDF plots, and style them for print.

Field
Biology, Chemistry
Prerequisites
none beyond Python basics
Libraries
matplotlib 3.11.2numpy 2.5.3pandas 3.0.6seaborn 0.13.2
Download notebook Save Mark as done

py-seaborn.ipynb, executed with the versions above. The download needs a free account

Run it yourself. In a terminal, this installs exactly the versions above:

pip install numpy==2.5.3 pandas==3.0.6 seaborn==0.13.2 matplotlib==3.11.2 jupyterlab

The problem: is line C larger, or only more variable?

A lab has measured the projected area of 200 cells in each of three cell lines, A, B, and C, from microscope images, and wants to know whether line C is larger than line B. The first figure everyone draws is a bar chart of the means, and seaborn draws it in one call: 123, 154, and 159 µm². The thin line on each bar is seaborn's default error bar, a 95 % confidence interval of the mean found by resampling the cells 1,000 times, which is called the bootstrap. For B it runs from 150 to 159 µm², for C from 150 to 168 µm². The chart invites a verdict: B and C are about the same, C perhaps a little larger.

That is half right, and it hides the important half. C is not larger than B in the middle: it has the same median to within 1 µm², 150.6 against B's 151.5 µm², and more than twice the spread, an interquartile range of 86 µm² against 41 µm². Of C's cells, 23.5 % lie above 200 µm² (B: 8 %) and 19.5 % below 100 µm² (B: 2 %). C's slightly higher mean comes from its long right tail, not from typical cells being larger. A chemist with particle diameters from three synthesis batches meets the same figure. How such areas come out of an image is the subject of Counting cells with scipy.ndimage; here they are generated.

Cell area in µm² of three cell lines, shown four ways on one shared axis: box plot, violin, every cell as a dot, and the cumulative fraction of cells. B and C have the same median, but C spreads about twice as far in both directions.

This is where we end up: the same 600 cells at the width of a two-column journal page, in four views that each answer a different question. Step 6 draws it with four seaborn calls, first as a 2 × 2 grid for the screen, then as this row, inside a block that switches Matplotlib to print settings.

Setup

Matplotlib looks up the default look of every figure, from font size to which frame lines show, in one dictionary, plt.rcParams, under keys such as "font.size". The style block below is a plt.rcParams.update({...}) call: it changes those defaults for every later figure. with plt.rc_context({...}): changes them only for the figures made inside the with block and puts the old values back at its end. Importing seaborn changes none of these settings.

The areas are log-normal: their logarithm is normally distributed, so they are positive and skewed to the right, as cell areas are. sigma is the typical factor of variation: 0.2 means a factor of about 1.2 up or down around the median, 0.4 about 1.5. B and C are drawn with the same median. The data arrive as wide, a pandas DataFrame, a table with named columns, one per cell line, as a spreadsheet has them (more in pandas from the ground up).

import warnings
from pathlib import Path

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

# seaborn 0.13.2 still passes vert= to Matplotlib's boxplot; harmless until Matplotlib 3.13
warnings.filterwarnings("ignore", message="vert: bool was deprecated")

plt.rcParams.update({
    "figure.figsize": (7, 3.6), "figure.dpi": 110,
    "axes.spines.top": False, "axes.spines.right": False,
    "axes.grid": True, "grid.alpha": 0.25,
    "font.size": 11, "lines.linewidth": 1.8,
})
INK, ACCENT, SECOND, MUTED = "#1f2a44", "#c8553d", "#2a7f9e", "#8a8f98"
LINES = {"A": INK, "B": SECOND, "C": ACCENT}   # one color per cell line, in every figure

SEED = 781
rng = np.random.default_rng(SEED)
np.random.seed(SEED)   # seaborn draws the strip plot's jitter from NumPy's global generator

n = 200
wide = pd.DataFrame({
    "A": rng.lognormal(np.log(120), 0.25, n),   # median 120 µm², the reference
    "B": rng.lognormal(np.log(150), 0.20, n),   # shifted up
    "C": rng.lognormal(np.log(150), 0.40, n),   # shifted like B, twice as wide
})
Path("assets").mkdir(exist_ok=True)

print(wide.shape)
print(wide.head(3).round(1))
(200, 3)
       A      B      C
0  161.1  188.1   78.6
1  140.0  173.0  252.4
2  105.7  147.0   51.2

Step 1: Turn the spreadsheet into long format with melt

seaborn wants its data in long format: one row per cell, with one column naming the line and one holding the area. Then any column can be assigned to x, y, or hue by its name. A table with one column per line is wide format, and melt turns it into long:

long = wide.melt(var_name="line", value_name="area")
print(wide.shape, "->", long.shape)
print(long.head(3).round(1))
(200, 3) -> (600, 2)
  line   area
0    A  161.1
1    A  140.0
2    A  105.7

Three columns of 200 became two columns of 600. Now the bar chart. fig, ax = plt.subplots() makes one figure with one set of axes, and every seaborn call in this tutorial draws onto the axes it is handed through ax= (more in Matplotlib from the ground up):

fig, ax = plt.subplots()
sns.barplot(long, x="line", y="area", hue="line", palette=LINES, legend=False,
            seed=SEED, saturation=1, alpha=0.45, err_kws={"color": INK, "linewidth": 1.5}, ax=ax)
ax.set(xlabel="cell line", ylabel="mean cell area / µm²")
plt.show()

means = long.groupby("line")["area"].mean()
for line, mean in means.items():
    print(f"{line}: mean {mean:5.1f} µm²")
Bar chart of the mean cell area in µm² for cell lines A, B, and C, with 95 % bootstrap intervals. The bars for B and C are nearly equal and their intervals overlap.
A: mean 122.7 µm²
B: mean 153.9 µm²
C: mean 158.5 µm²

hue assigns colors from a column. Here it names the same column as x, because we want one color per line, and seaborn 0.13 warns that a palette without hue is deprecated and will stop working in 0.14. legend=False drops the legend that would repeat the x labels, and the Variations show hue doing work of its own. seed=SEED fixes the resampling, so the interval is the same on every run, and groupby("line") computes the mean per line. The bars stand at 123, 154, and 159 µm², and the intervals of B and C overlap from 150 to 159 µm².

Step 2: Draw the histograms with hue, stat, and common_norm

Here hue does something x cannot: one call draws three histograms on one set of axes, told apart only by color.

fig, ax = plt.subplots()
sns.histplot(long, x="area", hue="line", palette=LINES, binwidth=10, element="step",
             fill=False, stat="density", common_norm=False, ax=ax)
ax.set(xlabel="cell area / µm²", ylabel="density / µm⁻²")
sns.move_legend(ax, "upper right", frameon=False, title="cell line")
plt.show()
Histograms of cell area in µm² for lines A, B, and C as step outlines, density on the y axis in per µm². C is lower and broader than B; the three outlines overlap.

With stat="density" the height of a bar times its width, 10 µm², is the fraction of the line's cells in that bar. The y axis is therefore in µm⁻², and each histogram has an area of 1. The alternatives are stat="count", "percent", and "probability".

common_norm decides what "each" means there. The default, True, divides by all cells of all lines together, so each line gets an area equal to its share of the cells. With 200 cells per line that is a third and only scales the curves; it matters when the groups are unequal. The cell below keeps only the first 50 cells of C, draws the histogram both ways, and adds up height times width of the bars seaborn drew for each line:

Show code
uneven = pd.concat([long[long["line"] != "C"], long[long["line"] == "C"].head(50)])
for common_norm in [True, False]:
    fig, ax = plt.subplots()
    sns.histplot(uneven, x="area", hue="line", binwidth=10, stat="density",
                 common_norm=common_norm, ax=ax)
    plt.close(fig)
    # seaborn adds the bars of the last line first
    areas = {line: sum(bar.get_height() * bar.get_width() for bar in bars)
             for line, bars in zip("ABC", ax.containers[::-1])}
    print(f"common_norm={common_norm!s:5}  " + "  ".join(f"{k}: {v:.3f}" for k, v in areas.items()))
common_norm=True   A: 0.444  B: 0.444  C: 0.111
common_norm=False  A: 1.000  B: 1.000  C: 1.000

With the default, the 50 cells of C get a ninth of the area and draw a flatter curve, though their shape is the same. Compare shapes with common_norm=False. Even then, three overlapping outlines are hard to read: C is lower and broader, and that is about all the histogram says.

Step 3: Compare medians and quartiles with a box plot, and show every cell with a strip plot

A box plot draws the middle half of the cells as a box, from the first to the third quartile, and the median as a line across it. The whiskers reach to the farthest cell within 1.5 times the interquartile range beyond the box (whis), and the circles are cells beyond them, not errors. A strip plot on top shows every cell, each shifted sideways by a small random amount, the jitter, so that dots do not pile up.

fig, ax = plt.subplots()
sns.boxplot(long, x="line", y="area", hue="line", palette=LINES, legend=False,
            fill=False, ax=ax)
sns.stripplot(long, x="line", y="area", hue="line", palette=LINES, legend=False,
              size=2.5, alpha=0.5, jitter=0.25, ax=ax)
ax.set(xlabel="cell line", ylabel="cell area / µm²")
plt.show()

quartiles = long.groupby("line")["area"].quantile([0.25, 0.5, 0.75]).unstack()
quartiles["IQR"] = quartiles[0.75] - quartiles[0.25]
print(quartiles.round(1))
Box plots of cell area in µm² for lines A, B, and C with every cell drawn as a dot. B and C have their median at the same height, and C's box is twice as tall as B's.
       0.25    0.5   0.75   IQR
line                           
A     102.6  119.7  140.0  37.4
B     131.4  151.5  172.0  40.6
C     111.7  150.6  197.9  86.2

The printed quartiles are the box edges, and the 0.5 column holds the medians. B's box is A's moved up by about 30 µm². C's median sits at B's, 150.6 against 151.5 µm², and its box is twice as tall, with an interquartile range of 86 µm² against 41 µm².

Step 4: Show the shape with a violin plot

A box reduces each line to five numbers; a violin shows the whole shape. Each cell contributes a small bell-shaped bump centered at its area, and the bumps are added up into one smooth curve, a kernel density estimate. The width of each bump is the bandwidth, which seaborn chooses from the data and bw_adjust scales (2 is smoother, 0.5 bumpier). The violin is that curve drawn on both sides of a vertical line.

fig, ax = plt.subplots()
sns.violinplot(long, x="line", y="area", hue="line", palette=LINES, legend=False,
               inner="quart", alpha=0.35, ax=ax)
sns.stripplot(long, x="line", y="area", hue="line", palette=LINES, legend=False,
              size=2.5, alpha=0.5, jitter=0.25, ax=ax)
ax.axhline(0, color=MUTED, lw=1)
ax.set(xlabel="cell line", ylabel="cell area / µm²")
plt.show()
Violins of cell area in µm² for lines A, B, and C, with quartile lines inside and every cell as a dot. C has the longest tail upward, and its violin reaches below its smallest cell, down to the zero line.

The longer dashes inside each violin mark the median and the shorter dashes the quartiles of Step 3. All three lines are skewed to the right, and C's tail reaches past 400 µm². At the bottom, C's violin runs below its smallest cell, down to the zero line; the pitfalls say why, and how far past zero it goes.

Step 5: Separate shift from spread with the ECDF

The ECDF, the empirical cumulative distribution function, gives for each area the fraction of the line's cells at or below it. Each cell lifts its line's curve by 1/200 at its own area, so there are no bins and no bandwidth to choose.

fig, ax = plt.subplots()
sns.ecdfplot(long, x="area", hue="line", palette=LINES, legend=False, ax=ax)
for x in [100, 200]:
    ax.axvline(x, color=MUTED, lw=1, ls="--")
for line, f, dx in [("A", 0.6, -8), ("B", 0.3, 8), ("C", 0.85, 8)]:   # where each curve is alone
    ax.text(wide[line].quantile(f) + dx, f, line, color=LINES[line], va="center",
            ha="right" if dx < 0 else "left")
ax.set(xlabel="cell area / µm²", ylabel="fraction of cells", ylim=(0, 1.03))
plt.show()

for line, area in wide.items():
    print(f"{line}: below 100 µm² {100 * (area < 100).mean():4.1f} %   above 200 µm² {100 * (area > 200).mean():4.1f} %"
          f"   median {area.median():5.1f} µm²")
Cumulative fraction of cells against cell area in µm² for lines A, B, and C, with dashed guides at 100 and 200 µm². B is A moved right; C crosses B at a fraction of 0.5 and rises more slowly.
A: below 100 µm² 22.0 %   above 200 µm²  1.5 %   median 119.7 µm²
B: below 100 µm²  2.0 %   above 200 µm²  8.0 %   median 151.5 µm²
C: below 100 µm² 19.5 %   above 200 µm² 23.5 %   median 150.6 µm²

Read one number first: B's curve passes 0.5 at its median, 151.5 µm². B's curve is A's moved right by about 30 µm². C's curve crosses B's at 0.5, the shared median, and rises more slowly, so C's cells spread further on both sides: 19.5 % of them lie below 100 µm² against 2 % of B's, and 23.5 % above 200 µm² against 8 %. Every number in the second paragraph of this page can be read off this one figure.

Step 6: Put the four views side by side at journal size

The figure for print is one row of four panels, 17.8 cm wide for a two-column journal page, with 8 pt text. That row is too small to read on a phone, so this web page first draws the same views as a 2 × 2 grid, 7 by 6 inches (figsize is in inches). Both come from one function that draws the four views onto any four Axes. The violins get cut=0, which ends them at the smallest and largest cell, where Step 4's ran past the data; the pitfalls say why. The three categorical calls share one dict, the ECDF gets y="area" so that sharey=True can give all four panels one area axis, and layout="constrained" makes Matplotlib fit the labels inside the figure.

CAT = dict(x="line", y="area", hue="line", palette=LINES, legend=False)

def four_views(axs, dot):
    box, violin, strip, ecdf = axs
    sns.boxplot(long, **CAT, fill=False, fliersize=dot, ax=box)
    sns.violinplot(long, **CAT, inner="quart", cut=0, alpha=0.35, ax=violin)
    sns.stripplot(long, **CAT, size=dot, alpha=0.5, jitter=0.25, ax=strip)
    sns.ecdfplot(long, y="area", hue="line", palette=LINES, legend=False, ax=ecdf)
    for line, f, dy in [("A", 0.6, -8), ("B", 0.3, 8)]:   # as in Step 5, turned
        ecdf.text(f, wide[line].quantile(f) + dy, line, color=LINES[line], ha="center",
                  va="bottom" if dy > 0 else "top")
    ecdf.text(0.72, wide["C"].quantile(0.85), "C", color=LINES["C"], ha="right", va="center")
    for ax in axs:
        left = ax.get_subplotspec().is_first_col()   # one area label per row
        ax.set(xlabel="cell line", ylabel="cell area / µm²" if left else "")
    ecdf.set(xlabel="fraction of cells", xlim=(0, 1.03), xticks=[0, 0.5, 1])

fig, axs = plt.subplots(2, 2, figsize=(7, 6), sharey=True, layout="constrained")
four_views(axs.ravel(), dot=2.5)
plt.show()
Four panels in a 2 by 2 grid sharing one cell-area axis in µm²: box plots, violins cut at the data, every cell as a dot, and the ECDF turned on its side with the fraction of cells on the x axis. B and C share a median; C spreads twice as far.

The turned ECDF has the fraction on the x axis: from an area on the y axis, read across to the curve and down to the fraction of cells below it.

For print, the settings are written out in PRINT and used with rc_context, so they hold only inside the block and the page figures keep the site style. Other rcParams keys can join it for another journal (where these numbers come from). The same function fills the row:

PRINT = {"font.size": 8, "lines.linewidth": 1.0, "axes.linewidth": 0.6,
         "xtick.major.width": 0.6, "ytick.major.width": 0.6}

with plt.rc_context(PRINT):
    fig, axs = plt.subplots(1, 4, figsize=(17.8 / 2.54, 6.5 / 2.54), sharey=True,
                            layout="constrained")
    four_views(axs, dot=1.5)
    fig.savefig("assets/four-views.pdf", metadata={"CreationDate": None})
    plt.close(fig)   # shown at the top of the page, at print size

w, h = fig.get_size_inches() * 2.54
print(f"assets/four-views.pdf: {w:.1f} x {h:.1f} cm")
assets/four-views.pdf: 17.8 x 6.5 cm

savefig keeps the size you asked for unless you pass bbox_inches="tight", which fits the file around what is drawn and changes its size. For a journal leave it out: the PDF keeps its 17.8 × 6.5 cm, and 8 pt text prints at 8 pt. The row is the figure at the top of this page. The box plot gives medians, quartiles, and outliers at a glance; the violin, the shape; the strip, every cell and how many there are; the ECDF, shift against spread and the fraction above any area, without bins or smoothing.

Pitfalls

A violin that claims cells of negative area. In Step 4, C's violin reaches below zero, although no cell is smaller than 44 µm². The cause is the bumps of Step 4: those of the smallest cells reach past them, and seaborn draws the curve out to cut=2 bandwidths beyond the extreme cells. So the lowest point of the violin is the smallest cell minus two bandwidths, and half the gap between them is the bandwidth seaborn chose. The cell below reads the lowest point of C's outline, which Matplotlib keeps as a list of points after drawing, with the default and with cut=0:

Show code
smallest = wide["C"].min()
bottom = {}
for cut in [2, 0]:
    fig, ax = plt.subplots()
    sns.violinplot(long, x="line", y="area", hue="line", legend=False, cut=cut, ax=ax)
    plt.close(fig)
    bottom[cut] = ax.collections[2].get_paths()[0].vertices[:, 1].min()   # the third violin is C
print(f"smallest cell of C:          {smallest:5.1f} µm²")
print(f"C's violin with cut=2 ends at {bottom[2]:5.1f} µm², {smallest - bottom[2]:.1f} µm² below it:"
      f" two bandwidths of {(smallest - bottom[2]) / 2:.1f} µm²")
print(f"C's violin with cut=0 ends at {bottom[0]:5.1f} µm²")
smallest cell of C:           44.0 µm²
C's violin with cut=2 ends at  -1.6 µm², 45.6 µm² below it: two bandwidths of 22.8 µm²
C's violin with cut=0 ends at  44.0 µm²

C's violin ends at -1.6 µm², 45.6 µm² below its smallest cell: two bandwidths of 22.8 µm². The fix is cut=0, which ends each violin at the extreme cells, with the strip on top so that the reader sees where the cells are.

sns.set_theme() throws away your style. It looks like a harmless first line of a plotting script, and it overwrites Matplotlib's settings with seaborn's own. The cell runs it inside plt.rc_context(), which with no dict changes nothing on entry and only puts every setting back at the end, so the later cells are untouched:

with plt.rc_context():
    before = dict(plt.rcParams)
    sns.set_theme()
    after = dict(plt.rcParams)
changed = [k for k in before if before[k] != after[k]]
print(f"{len(changed)} of {len(before)} settings changed, among them")
for k in ["font.size", "axes.facecolor", "axes.spines.top", "axes.spines.right"]:
    print(f"  {k:18s} {before[k]!s:>6} -> {after[k]!s}")
37 of 339 settings changed, among them
  font.size            11.0 -> 12.0
  axes.facecolor      white -> #EAEAF2
  axes.spines.top     False -> True
  axes.spines.right   False -> True

The font grows, the background turns gray, and the top and right frame lines come back, 37 settings changed by one call. Do not call it, or call it before your own rcParams.update, or pass your settings to it as rc=.

Cell lines stored as numbers. With line coded as 1, 2, and 3, hue treats it as a quantity: the colors run on a sequential palette from pale pink to dark purple, as for three doses of a drug, and line 3 looks like more of line 1. Store the lines as strings, or as pd.Categorical, pandas' column type for labels from a fixed set, or pass the palette as a dict keyed by the values, which also fixes the colors.

Variations

  • A second factor. If your table has another column, say treatment with treated and untreated cells, keep x="line" and give hue that column: the box of each line splits in two, side by side.
  • One panel per line. sns.displot(long, x="area", col="line", stat="density", binwidth=10) draws each line's histogram in its own panel with shared axes, which removes the overlap that made Step 2 hard to read.
  • A log axis. log_scale=True in histplot and ecdfplot. On a log axis a log-normal distribution is symmetric about its median, so C's extra spread shows evenly on both sides of the median it shares with B.
  • Is the gap real? The largest vertical gap between two ECDFs is the Kolmogorov-Smirnov statistic, and scipy.stats.ks_2samp(wide["B"], wide["C"]) turns it into a test. scipy.stats from the ground up covers comparing two samples.

Cheat sheet

long = wide.melt(var_name="line", value_name="area")        # one row per observation
kw = dict(x="line", y="area", hue="line", palette=LINES, legend=False)   # palette without hue: deprecated
sns.barplot(long, **kw, seed=SEED, ax=ax)                    # mean, bootstrap 95 % CI; seed it
sns.histplot(long, x="area", hue="line", stat="density", common_norm=False, ax=ax)
sns.boxplot(long, **kw, fill=False, ax=ax)                 # whiskers: 1.5 IQR (whis)
sns.violinplot(long, **kw, cut=0, inner="quart", ax=ax)    # cut=0: no tail past the data
sns.stripplot(long, **kw, jitter=0.25, ax=ax)              # jitter uses np.random.seed
sns.ecdfplot(long, x="area", hue="line", ax=ax)            # no bins, no bandwidth
sns.move_legend(ax, "upper right", frameon=False)
with plt.rc_context(PRINT): fig, axs = plt.subplots(1, 4, sharey=True)   # print settings here only

Further reading

Was this tutorial helpful? Sign in to tell the author with one click.

Found a mistake, or something unclear? Report a problem (with a free account).

Cite this tutorial

SciStack (2026). seaborn from the ground up: three cell lines in four views. https://scistack.dev/t/py-seaborn/ (accessed 2026-10-08).

@online{scistack-py-seaborn,
  author  = {{SciStack}},
  title   = {seaborn from the ground up: three cell lines in four views},
  date    = {2026-10-08},
  url     = {https://scistack.dev/t/py-seaborn/},
  urldate = {2026-10-08},
  note    = {numpy 2.5.3, pandas 3.0.6, seaborn 0.13.2, matplotlib 3.11.2}
}

Tags

barplotboxplotecdfplothistplotmatplotlibpandasseabornstripplotviolinplot

Comments

No comments yet.

Sign in to comment, with a free account.