CRAFT – Art & Science of Visualization

A platform-independent thinking discipline for building visualizations.

A visualization isn’t produced by picking a chart type. It is crafted — shaped, refined, and reduced through a sequence of deliberate decisions until what remains actually communicates something.

CRAFT names that sequence:

  • C – Context
  • R – Representation
  • A – Aesthetics
  • F – Framing
  • T – Telling

None of the five letters names a tool, a library, or a chart type. That is deliberate — CRAFT is the thinking that happens before and around any tool, not a feature of one.

Before anything is drawn, the data has to be understood on its own terms. This is the foundational stage: nothing gets plotted until the nature of the variable is clear.

The first question is always: is the variable numeric or categorical? The second is: how many variables are involved at once – one, or several? These two questions decide the entire shape of what follows.

Univariate – Numeric

  • Histogram
  • Scatter plot
  • Box plot
  • KDE (density plot)

Univariate – Categorical

  • Bar plot
  • Pie chart

Multivariate – Numeric vs Numeric

  • Scatter plot
  • Line chart

Multivariate – Numeric vs Categorical

  • Box plot
  • Violin plot

Multivariate – Categorical vs Categorical

  • Grouped bar plot
  • Stacked bar plot
  • Heatmap

Context establishes ground truth about the data before any visual decision is made.

Note: The above charts are examples for these scenarios. There are additional charts beyond these examples. The included charts are for reference only.

Once the base understanding is in place, it has to be built into something visual. Representation is the working stage – where the chosen chart form is developed, refined, and adjusted until it actually carries the data correctly.

Representation asks: does this form hold the data honestly? Does the scale distort the comparison? Does the grouping obscure the pattern? A correct chart choice can still be executed poorly – this is the stage where the work continues past selection into refinement.

Aesthetics governs how data values become visual properties: color, size, shape, texture, position. Its purpose is differentiation and emphasis — making the point that matters impossible to miss, and nothing more.

Perceptual accuracy differs across these properties: position is read more precisely than length, which is read more precisely than area or angle. Aesthetic choices should follow this hierarchy, not just visual preference.

Framing surrounds the visual with everything the viewer needs to read it correctly — titles, axis labels, legends, annotations, source notes.

The job here is to give the viewer the frame of reference required to interpret the chart without guessing. A title that only names the topic does less work than one that states the insight.

This is the layer that turns a correct chart into a meaningful one. A visual can be well-scoped, well-built, well-labeled — and still say nothing. Telling is the discipline of asking one more question: what should the viewer walk away believing?

A title that states the insight instead of the topic, a chart ordered to lead the eye toward the conclusion, a single annotation that turns a data point into a takeaway — this is Telling. It is not decoration on top of the chart; it is the reason the chart was made in the first place.


Every visualization ultimately reduces to two variable types:

TypeInterested inKey measures
Numeric / CountDistribution of valuesCentral tendency (mean, median, mode) and Dispersion (range, variance, IQR, skew)
CategoricalDistribution across groupsCount (frequency) and Proportion (share of whole)

Every plot below — including multivariate and trellis/grid layouts — is just this same question (“what is the distribution?”) asked for one variable, two variables, or many variables laid out together. When you look at any plot, first identify which of the two types each axis/encoding represents, then ask “what does this tell me about central tendency/dispersion (numeric) or count/proportion (categorical)?”


Shows: Frequency/count distribution via bins. Interpretation:

  • Peak(s) = mode(s) — unimodal vs. bimodal/multimodal signals distinct subgroups.
  • Symmetry vs. skew — right-skew (long tail right) means mean > median; left-skew is the reverse.
  • Bar height spread = dispersion; bin width choice can hide or exaggerate structure, so check sensitivity to bin count.
  • Empty bins/gaps can indicate missing data ranges or natural boundaries (e.g., minimum wage floors).

Shows: Smoothed probability density curve. Interpretation:

  • Same reads as histogram (modality, skew) but smoothing removes bin-width artifacts — better for comparing shapes across groups (overlay multiple KDEs).
  • Bandwidth is the KDE’s “bin width” — too wide oversmooths and hides bimodality; too narrow looks noisy.
  • Area under the curve always sums to 1 — height alone isn’t “count,” only relative density.

Shows: Five-number summary + outliers. Interpretation:

  • Median line = central tendency; box (IQR) = middle 50% dispersion.
  • Whisker length asymmetry = skew direction.
  • Points beyond whiskers = outliers — check if they’re data errors or genuine extremes.
  • Best for quick side-by-side comparison of central tendency/spread across many groups; loses information about multimodality (a box plot can look identical for a bimodal and a normal distribution).

Shows: Empirical cumulative distribution function. Interpretation:

  • Y-axis at any X = proportion of data ≤ that X — direct way to answer “what % is below/above this value.”
  • Steepness = density (steep section = many points concentrated there, same info as a histogram peak).
  • Flat stretches = data gaps/sparse regions.
  • Useful for comparing distributions without binning bias, and for reading off percentiles/medians directly (where curve crosses 0.5).

Shows: Count or proportion per category. Interpretation:

  • Bar height ranking = which categories dominate (count) or which share is largest (proportion).
  • Compare bar heights, not areas or angles — this is why bar charts are more accurate than pie charts for count/proportion comparison.
  • Watch for a truncated y-axis (not starting at 0), which exaggerates differences.

Shows: Proportion as part of a whole. Interpretation:

  • Only meaningful when parts sum to 100% of a single whole.
  • Human eyes judge angle/area poorly — use only when there are few categories (≤5) and one or two clearly dominate.
  • Avoid for comparing similarly-sized slices or across multiple pies (a bar chart reads better in both cases).

Shows: Relationship/correlation between two numeric variables. Interpretation:

  • Overall direction of the cloud = sign of correlation (positive/negative/none).
  • Tightness of the cloud = strength of relationship; a wide scatter = weak/no relationship.
  • Look for non-linear patterns (curves, U-shapes) — correlation coefficients only capture linear relationships.
  • Clusters or outlier points may indicate subgroups or data issues.

Shows: Ordered sequence or trend (X is ordered/time). Interpretation:

  • Slope = rate of change; look for the overall trend (central tendency of the trajectory) plus local dispersion (volatility, noise around the trend).
  • Turning points = where trend direction changes — worth investigating what happened there.
  • Multiple lines = comparing trends’ central tendency (are they trending the same way) and dispersion (do they diverge/converge over time).

Shows: Joint distribution of two numeric variables. Interpretation:

  • This is the 2-variable version of a histogram — color intensity = concentration of points in that (x,y) region.
  • Look for the “hot spot” = joint mode (most common combination of both variables).
  • Elongated hot regions along a diagonal = correlation; a circular/symmetric hot region = weak/no correlation.
  • Useful when a scatter plot would be too overplotted (too many points) to read density visually.

Shows: Distribution of numeric across categories. Interpretation:

  • Compare median (central tendency) and box/whisker length (dispersion) side by side across category groups.
  • Overlapping boxes suggest categories aren’t meaningfully different on this numeric variable; well-separated boxes suggest a real group effect.
  • Watch for differing outlier counts per group — some categories may just have more data.

Shows: Aggregated numeric per category + subgroup. Interpretation:

  • Bar height = the chosen aggregate (mean, sum, median) of the numeric variable per category — you’re only seeing central tendency, not spread (no dispersion info, unlike a box plot).
  • Compare heights within a category cluster (subgroup differences) and across clusters (category differences).
  • Be cautious: if bars show “mean,” a wide-spread or skewed group can look deceptively similar to a tight one — pair with a box plot if dispersion matters.

Shows: Frequency/proportion across two categorical variables. Interpretation:

  • Total bar height = count/proportion for the primary category (same as a simple bar chart).
  • Segment sizes within each bar = the breakdown by the second category — easiest to compare the bottom segment or the total; harder to compare middle segments precisely since they don’t share a baseline.
  • Switch to a 100%-stacked version when you care about proportion breakdown regardless of group size differences.

Shows: Cross-tabulation frequency as color intensity. Interpretation:

  • Each cell = count/proportion for a specific combination of the two categories.
  • Look for rows/columns that are uniformly light or dark (a category with a consistent relationship to all others) vs. a distinct hot cell (a strong, specific association between two particular category levels).
  • Row-normalizing or column-normalizing the counts (to proportions) changes what “hot” means — check whether the color scale is raw count or proportion.

Shows: Pairwise scatter of all numeric variables. Interpretation:

  • Each off-diagonal cell is just plot #8 (Scatter Plot) repeated for every variable pair — read each the same way (direction, tightness, non-linearity).
  • Diagonal cells are typically each variable’s own distribution (histogram/KDE) — read as plot #1/#2.
  • Scan for which variable pairs show the strongest relationships (tightest, most linear clouds) versus which are unrelated (circular clouds).

Shows: All numeric variables per observation across axes. Interpretation:

  • Each vertical axis is one numeric variable’s distribution (central tendency = where most lines cross that axis, dispersion = spread of crossing points).
  • Each line is one observation — bundles of near-parallel lines between two axes = correlated variables; heavy line-crossing (an “X” pattern) = negatively correlated or unrelated variables.
  • Color-coding lines by a categorical variable reveals whether categories occupy distinct “corridors” across variables (i.e., whether the categories are numerically separable).

Shows: All categorical variables as flows across axes. Interpretation:

  • This is the categorical analog of #16 — instead of lines per observation, you get flows (ribbons) whose width = count/proportion moving from one category level to another.
  • Wide ribbons = common combinations (high joint frequency); thin/absent ribbons = rare or impossible combinations.
  • Good for spotting funnel-like drop-offs or dominant paths through multiple categorical stages.

Shows: Any base plot replicated across 1–2 categorical variables. Interpretation:

  • This is not a new plot type — it’s a layout. Each panel is one of the 18 other plots, so interpret each panel using its own rules above (histogram panel → modality/skew; scatter panel → correlation; bar panel → count/proportion).
  • The added value is comparison across panels: look for panels that look different from the rest (a subgroup behaving unusually) versus a consistent pattern repeating across all panels (the categorical split doesn’t matter for this relationship).
  • Keep axis scales consistent across panels — mismatched scales make cross-panel comparison misleading.

This framework applied in R and Python platforms. you can refer them.

Scroll to Top