Plots with ROOTs

Data visualization with gnuplot | Live coding in terminal

0 — Start here

Follow one question at a time: inspect the data, make a first plot, improve its readability, then choose different charts for different questions. Both editions use the same five published CSV datasets, include crew and work through the same eighteen figures. No local AWK preparation is needed.

Begin with the independent Meet gnuplot chapter below; it needs no CSV files. Chapters 1–11 then apply the tool to counts, percentages and composition. Chapters 12–17 are chart extensions: choose them to answer a question, not because every option must be learned on the first day.

The workflow follows the ROOT principles: Robust, Open, Ongoing, Time-tested. Keep the data, plotting instructions and resulting figures understandable. Change one setting at a time and compare the result.

Use gnuplot 6.0 or newer with SVG and PNG output, and curl for downloads. Retrieve the data before an offline lesson. Once the downloads are available, all plotting can run without further network access.

Work in a terminal

Use Bash on macOS or Linux, WSL or Git Bash. Git Bash provides the shell commands; install gnuplot separately and add it to PATH. Use a plain-text editor and a browser to view SVG or PNG figures. Students do not need Emacs, LaTeX or Inkscape to run these examples.

During live coding, create each .gp file in the editor, save it, then run its printed gnuplot command. The supplied terminal/ directory is a worked solution for comparison or catch-up, not a prerequisite for writing the plots.

The listing titles distinguish three places to work:

Listing label Where to enter the code
Bash / shell At your normal terminal prompt; $ here represents the shell prompt.
At gnuplot> Inside the running gnuplot program, after starting gnuplot -d.
Editor In the named file; save it before running it from the shell.

Prompts such as $ and gnuplot> are orientation marks, not commands to copy. Omit them when entering code. Indented lines without a prompt continue a command or enter its data; finish those lines before starting the next command. Saved .gp files never contain prompts.

Check the programs

Run listing 1. command -v checks whether a program is available through PATH. If it says MISSING, install the program or fix its path, then open a new terminal and repeat the check.

Listing 1 — Bash / shell — Check the required programs
for program in bash gnuplot curl head cat mkdir; do
  if command -v "$program" >/dev/null 2>&1; then
    printf 'OK       %s\n' "$program"
  else
    printf 'MISSING  %s - install it or add it to PATH\n' "$program"
  fi
done

Check the gnuplot build

Check gnuplot's version and both output terminals with listing 2. Continue only when the checks succeed.

Listing 2 — Bash / shell — Check the gnuplot version and image terminals
set -e
gnuplot --version
gnuplot -d -e 'if (GPVAL_VERSION < 6.0) { exit error "Need gnuplot 6.0 or newer" }; set terminal svg; print "OK: version and SVG output"'
gnuplot -d -e 'set terminal pngcairo; print "OK: PNG output"'

Enter the working directory

Run listing 3 once from the course folder. Keep all later commands in terminal/. Quote paths containing spaces and use forward slashes, for example cd "/c/Users/Your Name/course/terminal" in Git Bash.

Listing 3 — Bash / shell — Create the working directories and set numeric formatting
mkdir -p terminal
cd terminal
mkdir -p data figures plots
export LC_NUMERIC=C

LC_NUMERIC=C selects decimal-point numeric formatting for programs started by this shell. It does not change the system language or the CSV separator. export passes the setting to child processes. If LC_ALL is set, it overrides LC_NUMERIC; run unset LC_ALL first in that case.

Practice this topic in Challenge 1: A working plotting environment.

Before Titanic — Meet gnuplot, from a first command to a simple chart

This chapter needs only gnuplot. No Titanic files, shared style or network connection is used. Start with a function; then use one small synthetic table. The short continuations below change one thing at a time. Predict the change, run it and look at the plot before proceeding.

First, draw a function

First run listing 4 at your Bash / shell prompt.

Listing 4 — Bash / shell — Start gnuplot from the shell
$ gnuplot -d

The prompt now changes to gnuplot>: you are inside gnuplot, not Bash. Enter 5 there, then 6. Do not type gnuplot -d again at this prompt. Gnuplot calls its output device a terminal: dumb draws with text characters directly in Bash, Git Bash or another terminal, with no additional window. This is a rough preview, not an SVG-quality image.

Listing 5 — At gnuplot> — Select a text preview in the terminal, without opening a window
gnuplot> set terminal dumb size 79,24

The SVGs printed in this guide are illustrations. Their caption filenames identify those illustrations; typing plot does not save them. Use the explicit export step below when you want an SVG file of your own.

Listing 6 — At gnuplot> — intro-first — A function, with no input file
gnuplot> reset
gnuplot> plot sin(x)
first.svg
Figure 1: The first function plot. (first.svg)

reset starts with known plotting settings. plot samples the built-in function; x is its argument. Sine uses radians by default. Gnuplot chooses the initial ranges automatically.

Optional — Smoother function curves in SVG

Gnuplot draws a function by joining calculated points. If the sine curve looks angular when you zoom in, add set samples 500 after reset and before plot sin(x) in 6, then rerun the complete example. In an ongoing terminal session, change the sample count, then repeat the whole SVG export sequence: reopen the output filename, replot, and close it with unset output. show samples checks the active setting; a later reset restores the default.

More samples refine the curve's geometry, not the display's pixels. Once the segments are smaller than screen pixels, 10000 samples may look no better than 500 while producing a much larger SVG. Reopen the newly exported SVG in a browser and zoom in to distinguish straight segments from pixelated edges. This setting does not improve the character-grid resolution of dumb or add points to ordinary tabular-data plots.

Save the current plot as SVG

After 6, enter 7 in the same gnuplot session. The figures/ directory was created in the setup chapter; keep working from terminal/. set terminal svg selects the file format, set output names the file, and replot draws the current plot into it. unset output closes the file, but SVG remains the selected output format. We keep it selected for the rest of this chapter.

Listing 7 — At gnuplot> — Save the current plot as figures/first.svg and keep SVG selected
gnuplot> set terminal svg size 900,400
gnuplot> set output "figures/first.svg"
gnuplot> replot
gnuplot> unset output

Open figures/first.svg in a browser for the full graphic. Reuse this export step after any later plot, changing the filename to match that example (for instance figures/two-series.svg). Reusing a filename overwrites that file. There is no need to select the SVG terminal again unless you deliberately change the output format.

One SVG document per file: do not call replot repeatedly while the same SVG output file stays open. Each call appends another complete SVG, causing an XML error such as “junk after document element”. After changing a setting such as set samples 500, repeat all of listing 7. Its set output reopens and replaces the file; unset output finishes it. This also repairs a file containing several concatenated SVG documents. Then reload the file in your browser.

For each following plotting listing, including the short continuations: enter set output "figures/first.svg" before its commands, and unset output after its last plot or replot. Reload the browser to see the change. Choose a different filename to keep a previous version. For example, 8 now runs between those two output commands. The same rule applies to the later synthetic-data examples. With SVG selected and no output file open, a bare plot or replot writes SVG markup to the shell; it does not update the saved file.

One change, then look again

Listings 8 through 13 are continuations in the same interactive session, after 6. Enter each pair of commands in order. set changes a setting; replot redraws the previous plot with that setting. It does not create a plot from nothing.

In Org, these short fragments are deliberately not executable: separate :session none blocks do not remember the previous plot. Instead, add each new set line just before plot sin(x) in a copy of 6 and rerun that complete block. Give the copy a new block name and output path. No persistent Babel session is needed.

Listing 8 — At gnuplot> — Continuation — Add just a title
gnuplot> set title "A sine wave"
gnuplot> replot
Listing 9 — At gnuplot> — Continuation — Name the horizontal axis
gnuplot> set xlabel "Angle (radians)"
gnuplot> replot
Listing 10 — At gnuplot> — Continuation — Name the vertical axis
gnuplot> set ylabel "sin(x)"
gnuplot> replot
Listing 11 — At gnuplot> — Continuation — Show one full cycle
gnuplot> set xrange [0:2*pi]
gnuplot> replot
Listing 12 — At gnuplot> — Continuation — Leave space above and below the wave
gnuplot> set yrange [-1.2:1.2]
gnuplot> replot
Listing 13 — At gnuplot> — Continuation — Add reference lines
gnuplot> set grid
gnuplot> replot

Listing 14 is the complete, independently runnable checkpoint after the six changes. It does not depend on 6.

Listing 14 — At gnuplot> — intro-function-settings — Complete function checkpoint
gnuplot> reset
gnuplot> set title "A sine wave"
gnuplot> set xlabel "Angle (radians)"
gnuplot> set ylabel "sin(x)"
gnuplot> set xrange [0:2*pi]
gnuplot> set yrange [-1.2:1.2]
gnuplot> set grid
gnuplot> plot sin(x)
function-settings.svg
Figure 2: The same function after the six individual changes. (function-settings.svg)

Practice this step in Intro A: Change the visible range.

One synthetic table, three numeric columns

The following six rows are synthetic, not observations. Column 1 is a step number; columns 2 and 3 are two invented measurements in arbitrary units. All later data examples in this chapter reuse this one table.

Listing 15 is another complete checkpoint. The lines between $Measurements << EOD and EOD define a gnuplot data block: a little table stored in memory, without a separate file. Enter the whole definition before plotting it. Spaces separate its numeric columns.

Listing 15 — At gnuplot> — intro-data — The complete synthetic table and its first plot
gnuplot> reset
gnuplot> $Measurements << EOD
         1 2 3
         2 4 4
         3 3 5
         4 6 5
         5 5 7
         6 8 6
         EOD
gnuplot> set xlabel "Step"
gnuplot> set ylabel "Measurement (arbitrary units)"
gnuplot> plot $Measurements using 1:2 with points title "Series A"
points.svg
Figure 3: Six synthetic points; step is x and column 2 is y. (points.svg)

using 1:2 means x from column 1, y from column 2. It is not a row range. For example, row 3 places a point at x = 3, y = 3. title "Series A" names the legend entry, not the whole chart.

Change only the drawing style

Run 16, then 17 in the same prompt after 15. In Org, replace only the final plot line in 15 and rerun the whole block. These are replacement lines, not independent examples: the data block and axis labels must still be present.

Listing 16 — At gnuplot> — Continuation — Replace points with connecting lines
gnuplot> plot $Measurements using 1:2 with lines title "Series A"
lines.svg
Figure 4: Lines connect the same six rows in their input order. (lines.svg)
Listing 17 — At gnuplot> — Continuation — Keep the line and show the measured positions
gnuplot> plot $Measurements using 1:2 with linespoints title "Series A"
linespoints.svg
Figure 5: Linespoints makes both the trend and the six positions visible. (linespoints.svg)

Practice this step in Intro B: Change the drawing style.

Add one series, keep the axes

Column 3 supplies a second series on the same axes, called “Series B”. The comma separates two series in one plot command; a backslash continues the command on the next line. Listing 18 draws both series: replace the final plot line in 15, or enter it in that session.

Listing 18 — At gnuplot> — Continuation — Add the second synthetic series
gnuplot> plot $Measurements using 1:2 with linespoints title "Series A", \
              $Measurements using 1:3 with linespoints title "Series B"
two-series.svg
Figure 6: A second y column, with its own legend entry. (two-series.svg)

At step 6, Series A is 8 and Series B is 6. Both series share the same x values. No percentages or grouping have been introduced yet.

Practice this step in Intro C: Add a second series.

From the prompt to a saved script

For a reproducible checkpoint, copy the complete body of 15 into plots/intro.gp in your editor and save it (not just a continuation). Optionally replace its final plot line with 18. Use listing 19 inside gnuplot to return to Bash.

Listing 19 — At gnuplot> — Leave gnuplot and return to the shell
gnuplot> exit

You are now back at the shell prompt. From terminal/, listing 20 runs this script as a text preview in Bash; listing 21 saves an SVG instead. Choose the output you need. Neither opens a plot window.

Listing 20 — Bash / shell — Preview the saved synthetic example directly in the terminal
gnuplot -d -e 'set terminal dumb size 79,24' plots/intro.gp

The two output settings in 21 are file-export plumbing; they do not change the data. Gnuplot closes the output file when it exits.

Listing 21 — Bash / shell — Run the saved synthetic example and create figures/intro.svg
gnuplot -d -e 'set terminal svg; set output "figures/intro.svg"' plots/intro.gp

Open figures/intro.svg in a browser. Save, rerun and reload after an edit. The later Titanic chapters use the same editor/save/run cycle, but their scripts already contain set terminal and set output commands. Those scripts save the named file directly; the command-line preview setting above would be overridden. See 31 for a text preview of the first Titanic script.

1 — Download and inspect the data

We now apply the plotting commands we already know to a real dataset. The introduction supplied the tool; from here on we ask questions about people, groups and survival. The new step in this chapter is retrieving and interpreting published columns, not calculating a new dataset.

Use Titanic Passengers and Crew: Demographics and Survival Data, version 1.0.1. This version-specific record fixes which data the lesson uses. All counting and age grouping have already been done; no conversion or preprocessing is needed.

Published CSV Rows Contents and use
titanic.csv 2,207 Individual records; inspect the underlying data
titanic_by_class.csv 4 Class/crew counts; bars, rates, pie and donut
titanic_by_sex.csv 4 Counts by class/crew and sex; comparisons, heatmap, dumbbell and spider
titanic_by_age_and_class.csv 100 Three-year age/class cells; large heatmap and four panels
titanic_by_age.csv 25 Three-year age groups; sparklines

The person-level columns are class, sex, age, survived and embarked. There are 2,207 records: 1,317 passengers and 890 crew, with 711 survivors. The categories 1st, 2nd, 3rd and Crew appear in that order. Crew is a separate group, not a fourth passenger class.

NA marks missing values, not zero. The two unknown ages are excluded from age summaries, leaving 2,205 people. The dataset description documents its provenance, transformations and limitations; these are descriptive records, not evidence of causal effects.

Retrieve the five files

From terminal/, run listing 22 to save the published CSVs in data/. The failure guard stops the shell if a download fails; resolve the error and rerun it before plotting.

Listing 22 — Bash / shell — Download all five Titanic CSVs from Zenodo
for file in titanic.csv titanic_by_class.csv titanic_by_sex.csv \
            titanic_by_age_and_class.csv titanic_by_age.csv; do
  curl --fail --location \
    "https://zenodo.org/api/records/22983324/files/$file/content" \
    --output "data/$file" || exit 1
done

--output writes each response to its named CSV file. The loop downloads each file once; it does not calculate or regroup the data.

--fail reports HTTP errors and --location follows redirects. Where used, --silent --show-error hides the progress meter but retains errors.

Inspect individual records

Run listing 23. head displays the requested lines without truncating or modifying the downloaded file.

Listing 23 — Bash / shell — Inspect the person-level CSV
head -n 5 data/titanic.csv

A short preview of the records:

class,sex,age,survived,embarked
1st,female,2,0,S
1st,female,13,1,S
1st,female,16,1,C
1st,female,16,1,S

survived is 1 for survival and 0 otherwise. Ages can be fractional; the two missing ages remain NA. Names and identifiers are not present.

Inspect the class-and-crew counts

Run listing 24 to inspect the same ready-made counts.

Listing 24 — Bash / shell — Inspect counts and outcomes by class/crew
cat data/titanic_by_class.csv
Class,Aboard,Survived,Died
1st,324,201,123
2nd,284,118,166
3rd,709,181,528
Crew,890,211,679

Columns 2, 3 and 4 are aboard, survived and died. All ages and both sexes are included. Every row satisfies survived + died = aboard; for Crew, 211 + 679 = 890. First class has 324 people and 201 survivors.

Inspect the class-and-sex counts

The file is already downloaded; listing 25 displays it.

Listing 25 — Bash / shell — Inspect the published counts by class/crew and sex
cat data/titanic_by_sex.csv

Columns 2–4 are women aboard, survived and died; columns 5–7 are the corresponding men's counts. These categories include children. A women's survival rate uses $3/$2; a men's uses $6/$5.

The first row is 1st,144,139,5,180,62,118: 324 people and 201 survivors. The Crew row is Crew,23,20,3,867,191,676: 890 people and 211 survivors. Adding the women's and men's counts reproduces the class table.

Inspect the age-and-class summary

Preview the downloaded summary with listing 26.

Listing 26 — Bash / shell — Inspect the published three-year age/class cells
head -n 9 data/titanic_by_age_and_class.csv

Its columns are age_midpoint, class_code, survival_pct, n and survivors. The 100 rows cross 25 three-year age bands with four groups. Class codes 1–4 mean first, second, third and crew. Labels 1, 4, …, 73 represent completed-year groups 0–2, 3–5, …, 72–74.

Rows cycle through class codes 1, 2, 3, 4 within each age band. Keep that order for the four-panel plot. Empty cells are retained with n=0 and survival_pct=NA; these are not zero-percent survival observations.

Inspect the age-only summary

Listing 27 — Bash / shell — Inspect the 25 published three-year age groups
cat data/titanic_by_age.csv

The columns are Age, People and Survived %. All 25 groups run in age order, from 0–2 through 72–74, with no combined 60+ group. The first row is 0-2,34,61.8 and the last is 72-74,2,0.0. Counts sum to 2,205; both sexes, all passenger classes and crew are pooled. Percentages have one decimal place; NaN would mark an empty age-only group.

Keep headers separate from observations

Each script uses set datafile separator "," for CSV and set datafile columnheaders for its header. Put those settings after reset. With headers enabled, row index 0 is already the first observation: do not add every ::1, which would skip it. The variable data is just a string holding a relative filename such as data/titanic_by_class.csv.

Practice this topic in Challenge 2: Match a question to a dataset.

2 — A first plot

Start with everyone aboard by class and crew. Each table row is one category and its aboard count supplies the bar height. Keep gnuplot's default appearance for this first experiment; later topics make each design choice explicit.

In your editor, create plots/by_class_aboard.gp inside terminal/. Enter listing 28 and save it.

Listing 28 — Editor — plots/by_class_aboard.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,300 font 'sans,13'
set output 'figures/by_class_aboard.svg'
set style data histograms
set yrange [0:*]
plot data using 2:xtic(1) title "Aboard"
unset output

After saving, run listing 29 from Bash in terminal/.

Listing 29 — Bash / shell — Run plots/by_class_aboard.gp from Bash
gnuplot -d plots/by_class_aboard.gp
by_class_aboard.svg
Figure 7: People aboard by passenger class and crew. (by_class_aboard.svg)

Practice this topic in Challenge 3: Ask the first plot a different question.

Open figures/by_class_aboard.svg in a browser. set output names that file. Use a plain-text editor, save with the exact .gp extension—not =.gp.txt=— then rerun and reload the browser after each change.

Bash receives the gnuplot -d ... command; the editor contains gnuplot statements. The -d flag skips personal startup files. reset clears settings left by earlier plots. Relative paths are resolved from the working directory, not from the script: stay in terminal/.

You can also work at the gnuplot prompt. Run listing 30 now, then use load after saving a script. Bash's $ and gnuplot's gnuplot> prompts are not characters to copy into commands.

Listing 30 — Bash / shell — Start the interactive gnuplot prompt
$ gnuplot -d

For a text preview without a separate window, enter listing 31 at that prompt. load first runs the script and saves its SVG; switching to dumb afterwards and calling replot also draws the chart in the terminal. The saved SVG remains unchanged. Text output has limited resolution and cannot faithfully reproduce every later chart, especially multiplots and colour heatmaps. Use their saved SVG or PNG files for the full result.

Listing 31 — At gnuplot> — Save figures/by_class_aboard.svg and also preview it as text
gnuplot> load 'plots/by_class_aboard.gp'
gnuplot> unset output
gnuplot> set terminal dumb size 79,24
gnuplot> replot

3 — One series, minimal code

New here: categorical positions and boxes, rather than the numeric x coordinates and points in the introduction. using 3:xtic(1) takes values from column 3 and category labels from column 1. This version shows only survivors. notitle suppresses an unhelpful data-source legend. The zero baseline is explicit; descriptive wording is the next refinement.

Category labels versus numerical coordinates. xtic(1) supplies labels, not numerical x positions: the classes are spaced equally by row. For age data, use the actual age as x, for example using 1:2 when column 1 holds age and column 2 holds the measured value. Ages 1, 4 and 10 must have gaps of 3 and 6 years; plotting against row number would make both gaps equal.

In your editor, create plots/by_class_survivors.gp inside terminal/. Enter listing 32, then save the file before running it.

Listing 32 — Editor — plots/by_class_survivors.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_survivors.svg'
set boxwidth .5 relative
set yrange [0:*]
plot data using 3:xtic(1) with boxes notitle
unset output

Run listing 33 from Bash in terminal/.

Listing 33 — Bash / shell — Run plots/by_class_survivors.gp from Bash
gnuplot -d plots/by_class_survivors.gp
by_class_survivors.svg
Figure 8: Survivors by class, with gnuplot's defaults. (by_class_survivors.svg)

Practice this topic in Challenge 4: Give bars an honest baseline.

4 — Two outcomes, readable labels

New here: histograms groups bars by category. Both outcomes share axes and a legend. The comma and series title work just as they did for the two synthetic series in the introduction. We repeat data for clarity; later '' means reuse the preceding data source. Colours are still gnuplot's defaults.

In your editor, create plots/by_class_counts.gp inside terminal/. Enter listing 34, then save the file before running it.

Listing 34 — Editor — plots/by_class_counts.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcomes.svg'
set style data histograms
set style fill solid 1.0
set title "Titanic: survivors and deaths by class and crew"
set ylabel "People aboard"
set yrange [0:*]
set key top left
plot data using 3:xtic(1) title "Survived", \
     data using 4         title "Died"
unset output

Run listing 35 from Bash in terminal/.

Listing 35 — Bash / shell — Run plots/by_class_counts.gp from Bash
gnuplot -d plots/by_class_counts.gp
by_class_outcomes.svg
Figure 9: Survivors and deaths by class, side by side. (by_class_outcomes.svg)

Practice this topic in Challenge 5: Check two series against their total.

5 — Make the appearance deliberate

Every visual choice is explicit: left and bottom axes only (set border 3), outward ticks, light horizontal guides and two colours chosen on purpose. Use a strong colour for survivors and a quiet one for deaths. In a report, the caption can carry the title rather than repeating it inside the figure.

In your editor, create plots/by_class_styled.gp inside terminal/. Enter listing 36, then save the file before running it.

Listing 36 — Editor — plots/by_class_styled.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcomes_styled.svg'
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
set border 3 lw 1.9
set tics out nomirror scale 0.9
set grid ytics lc rgb "#d5d8dc" lw .5
set key top left
set linetype 1 lc rgb "#14507d"
set linetype 2 lc rgb "#aab0b8"
set style fill solid 0.9 border -2
set style data histograms
set style histogram clustered gap 1
set boxwidth 0.9

set ylabel "People aboard"
set xlabel "Passenger class or crew"
set yrange [0:*]
plot data using 3:xtic(1) title "Survived", \
     data using 4         title "Died"
unset output

Run listing 37 from Bash in terminal/.

Listing 37 — Bash / shell — Run plots/by_class_styled.gp from Bash
gnuplot -d plots/by_class_styled.gp
by_class_outcomes_styled.svg
Figure 10: The same figure, with every visual choice written out. (by_class_outcomes_styled.svg)

Most of the setup describes a visual style, not this particular question. Repeating those settings in every plot invites inconsistent edits.

Practice this topic in Challenge 6: Make one visual change at a time.

Define the shared style

A shared style contains settings, not a plot command. It travels with the project and can be inspected by another reader. A personal ~/.gnuplot would apply invisibly on one machine without necessarily travelling with the lesson.

Before using styled plots, create style.gp directly inside terminal/, not inside plots/. Save listing 38 there. Each later script uses load 'style.gp' to read those settings.

Listing 38 — Editor — style.gp — Shared plotting settings
# Shared appearance for the terminal lesson.
reset
if (GPVAL_VERSION < 6.0) { exit error "Need gnuplot 6.0 or newer" }
set encoding utf8
set datafile separator ","
set datafile columnheaders
set border 3 lw 1.9
set tics out nomirror scale 0.9
set grid ytics lc rgb "#d5d8dc" lw .5
set key top left
set linetype 1 lc rgb "#14507d"
set linetype 2 lc rgb "#aab0b8"
set style fill solid 0.9 border -2
set style data histograms
set style histogram clustered gap 1
set boxwidth 0.9

6 — Reuse the shared style

This recreates the deliberately styled grouped plot. Its appearance should match the previous figure apart from the x-axis label: the style has moved, not the question. Each plot now contains primarily what is specific to its data and comparison.

In your editor, create plots/by_class_shared.gp inside terminal/. Enter listing 39, then save the file before running it.

Listing 39 — Editor — plots/by_class_shared.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcomes_shared_style.svg'
load 'style.gp'
set ylabel "People aboard"
set yrange [0:*]
plot data using 3:xtic(1) title "Survived", \
     ''   using 4         title "Died"
unset output

Run listing 40 from Bash in terminal/.

Listing 40 — Bash / shell — Run plots/by_class_shared.gp from Bash
gnuplot -d plots/by_class_shared.gp
by_class_outcomes_shared_style.svg
Figure 11: The same figure, drawn with the shared style. (by_class_outcomes_shared_style.svg)

Practice this topic in Challenge 7: Change a shared setting once.

7 — Stacked counts including crew

rowstacked puts survivors and deaths on top of each other. The total height of each bar is the size of the group, so this chart answers both “how large was the group?” and “how did its outcomes divide?”

In your editor, create plots/by_class_stacked.gp inside terminal/. Enter listing 41, then save the file before running it.

Listing 41 — Editor — plots/by_class_stacked.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcomes_stacked.svg'
load 'style.gp'
set style histogram rowstacked
set ylabel "People aboard"
set yrange [0:*]
plot data using 3:xtic(1) title "Survived", \
     ''   using 4         title "Died"
unset output

Run listing 42 from Bash in terminal/.

Listing 42 — Bash / shell — Run plots/by_class_stacked.gp from Bash
gnuplot -d plots/by_class_stacked.gp
by_class_outcomes_stacked.svg
Figure 12: People by class and crew, survivors and dead stacked. (by_class_outcomes_stacked.svg)

The bar totals are 324, 284 and 709 passengers, followed by 890 crew. Crew is the largest group, with 679 deaths; third class has 528 deaths. Both groups are larger than either first or second class.

Practice this topic in Challenge 8: Choose between stacked and grouped counts.

8 — One hundred percent stacks

The same stacks are now scaled to 100 per cent. A computed column is an expression in parentheses; $3 means “column 3 of this row”. Survived divided by aboard, multiplied by 100, is the survival percentage. The survivors' and deaths' percentages add to 100 within each group.

Keep the expression in parentheses: ($3/$2*100) is a calculation, whereas an unparenthesised using entry is interpreted as a column specification.

In your editor, create plots/by_class_shares.gp inside terminal/. Enter listing 43, then save the file before running it.

Listing 43 — Editor — plots/by_class_shares.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcome_percentages.svg'
load 'style.gp'
set style histogram rowstacked
set ylabel "Share of each class"
set yrange [0:100]
set format y "%.0f%%"
set key below
plot data using ($3/$2*100):xtic(1) title "Survived", \
     ''   using ($4/$2*100)         title "Died"
unset output

Run listing 44 from Bash in terminal/.

Listing 44 — Bash / shell — Run plots/by_class_shares.gp from Bash
gnuplot -d plots/by_class_shares.gp
by_class_outcome_percentages.svg
Figure 13: Share of survivors and dead in each class. (by_class_outcome_percentages.svg)

Practice this topic in Challenge 9: Distinguish a share from a count.

9 — Rates with exact labels

New here: one rate per group, followed by an optional text layer. The denominator is still the aboard count of that same class or Crew group, not the total of 2,207. First draw the bars without formatting strings.

First pass — only the rates

Use listing 45 after the input setup from the preceding class-count example: data must refer to titanic_by_class. In Org, copy that example's complete block (including :var data=titanic_by_class), give it a new name and file path, and replace its body with this listing. In a terminal script, retain the CSV separator, column-header, data, terminal and output setup, but replace the plotting body. This is a replacement body, not an independently runnable source block.

Listing 45 — Editor — First pass — Survival rates without the label layer
set ylabel "Survived (%)"
set yrange [0:100]
set boxwidth 0.6
plot data using 0:($3/$2*100):xtic(1) with boxes notitle

0 is the row index (0, 1, 2, 3), used for the bar positions. The calculated y value is the same percentage introduced in the preceding chapter. Check the first-class bar: 201/324*100 is about 62.0%.

Report extension — exact labels and shared appearance

The complete version below retains the report styling. Its only new plot layer is with labels: x and y place the text, and sprintf supplies it. The function survival_percent names the repeated calculation; it does not create or regroup data. %.1f formats one decimal place, and %% prints a literal percent sign. Compare bar heights before and after adding the label layer: none should move. Challenge 10 practices the precision.

In your editor, create plots/by_class_survival_rate.gp inside terminal/. Enter listing 46, then save the file before running it.

Listing 46 — Editor — plots/by_class_survival_rate.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_survival_rate.svg'
load 'style.gp'
survival_percent(survived, aboard) = 100.0 * survived / aboard
set ylabel "Survived"
set yrange [0:100]
set format y "%.0f%%"
set boxwidth 0.6
set offsets 0.5, 0.5, 0, 0

# Both layers use the same x position and the same survival percentage.
plot data using 0:(survival_percent($3,$2)):xtic(1) with boxes notitle, \
     data using 0:(survival_percent($3,$2)):(sprintf("%.1f%%", survival_percent($3,$2))) \
          with labels offset 0,0.5 notitle
unset output

Run listing 47 from Bash in terminal/.

Listing 47 — Bash / shell — Run plots/by_class_survival_rate.gp from Bash
gnuplot -d plots/by_class_survival_rate.gp
by_class_survival_rate.svg
Figure 14: Survival rate by class. (by_class_survival_rate.svg)

The percentages are 62.0 for first class, 41.5 for second, 25.5 for third and 23.7 for Crew. Rates compare survival within groups; they do not show how many people were aboard.

Class Survived %
1st 62.0
2nd 41.5
3rd 25.5
Crew 23.7

Practice this topic in Challenge 10: Decide how much precision to show.

10 — Class and sex

Compare women's and men's survival rates within each group. Women use columns 3/2; men use 6/5. Each denominator must be that group's own aboard count, not the total number of survivors or the total aboard.

The colours now encode sex rather than survival outcome. A linetype number only selects a style: change its colour and the legend together when the plot's question changes. These categories include children.

In your editor, create plots/by_class_and_sex.gp inside terminal/. Enter listing 48, then save the file before running it.

Listing 48 — Editor — plots/by_class_and_sex.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_and_sex_survival_rates.svg'
load 'style.gp'
set linetype 1 lc rgb "#14507d"   # women
set linetype 2 lc rgb "#c47a2c"   # men
set ylabel "Survived"
set yrange [0:100]
set format y "%.0f%%"
set key top right
plot sex using ($3/$2*100):xtic(1) title "Women", \
     sex using ($6/$5*100)         title "Men"
unset output

Run listing 49 from Bash in terminal/.

Listing 49 — Bash / shell — Run plots/by_class_and_sex.gp from Bash
gnuplot -d plots/by_class_and_sex.gp
by_class_and_sex_survival_rates.svg
Figure 15: Survival rate by class, women and men. (by_class_and_sex_survival_rates.svg)

Practice this topic in Challenge 11: Compare within a group.

11 — A pie with circles

A pie shows one whole divided into parts. Here the whole is everyone aboard, including crew. Use it for one composition; bars are clearer for comparing survival rates between groups.

For a simple pie, with circles is sufficient in gnuplot 6: supply the centre, radius, start angle and end angle for each slice, plus its colour. The default set style circle wedge connects the arc to the centre. No polar mode is needed. The donut chapter introduces with sectors for actual ring segments; using gnuplot 6 does not mean replacing every older style.

New here: stats data using 2 nooutput reads the aboard column and stores its sum as STATS_sum. Save that as total (2,207). The named function angle_for_count converts a count into its share of 360 degrees. For example, first class occupies 324/2207*360, about 52.85 degrees. The running variable angle_end carries one slice's end to the next slice's start. Start with slices only: labels are an extension in Challenge 12.

In your editor, create plots/by_class_pie.gp inside terminal/. Enter listing 50, then save the file before running it.

Listing 50 — Editor — plots/by_class_pie.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_aboard_pie.svg'
# reset leaves Cartesian coordinates and unrestricted ranges for stats.
stats data using 2 nooutput
total = STATS_sum
angle_for_count(people) = people * 360.0 / total

set title "Titanic: people aboard by class and crew"
set angles degrees
unset border
unset tics
unset grid
unset key
set size ratio -1
set xrange [-1.05:1.05]
set yrange [-1.05:1.05]
set style fill solid 1.0 border lc rgb "white"
# circles: x : y : radius : start angle : end angle : color
angle_end = 0
plot data using (0):(0):(1):(angle_end): \
     (angle_end = angle_end + angle_for_count($2)):($0 + 1) \
     with circles linecolor variable notitle
unset output

Run listing 51 from Bash in terminal/.

Listing 51 — Bash / shell — Run plots/by_class_pie.gp from Bash
gnuplot -d plots/by_class_pie.gp
by_class_aboard_pie.svg
Figure 16: People aboard by passenger class and crew. (by_class_aboard_pie.svg)

This example starts with reset, so stats runs before restrictive ranges or polar mode are enabled. For a pasted fragment in an existing session, run unset polar and set autoscale before stats.

Read the last line of the calculation from left to right: the fourth circle column reads the old angle_end; the fifth adds this row's angular width and returns the new end. This single update is necessary to join successive slices. Before repeating the plot, reset angle_end to zero; plain replot alone would reuse the accumulated value. The sixth column selects a distinct default colour via linecolor variable. ($0 + 1) turns row indices 0–3 into colour indices 1–4. The explicit ranges leave room around the unit circle, avoid the empty-y-range warning and prevent clipping artefacts at the circle's boundary. The equal axis scale (set size ratio -1) keeps the pie circular, not elliptical.

Optional report palette

The default colours distinguish the construction's four slices. To recover the muted report palette, insert listing 52 before the plot and rerun the complete example. This changes appearance, not shares. The label challenge uses this darker palette so white text remains readable.

Listing 52 — Editor — Extension — The four-colour report palette
set linetype 1 linecolor rgb "#14507d"
set linetype 2 linecolor rgb "#4f86b5"
set linetype 3 linecolor rgb "#6f9bc4"
set linetype 4 linecolor rgb "#476178"

Angles start at three o'clock and run anticlockwise in table order: first class, second class, third class, then Crew. This deliberately unlabelled plot is a construction step, not a finished communication graphic. Without the table order, a reader cannot identify the slices from colour alone.

Complete it in Challenge 12: Give the slices names and percentages.

12 — Heatmaps and age panels

Class and sex

Each cell is a rectangle coloured by its survival percentage. With only eight cells a table would also work; the heatmap becomes more useful as the grid grows. Here with boxxyerror draws the cells.

The five using values are x, y, horizontal half-width, vertical half-height and colour value. using (0):0 places women at x=0 and each data row on its own y coordinate. Pseudo-column 0 is the row index.

In your editor, create plots/by_class_and_sex_heatmap.gp inside terminal/. Enter listing 53, then save the file before running it.

Listing 53 — Editor — plots/by_class_and_sex_heatmap.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_and_sex_survival_heatmap.svg'
load 'style.gp'
unset grid; unset key
set xrange [-0.5:1.5]
set yrange [3.5:-0.5]
set xtics ("Women" 0, "Men" 1) scale 0
set ytics scale 0
set palette defined (0 "#ffffff", 100 "#4f86b5")
set cbrange [0:100]
set format cb "%.0f%%"
set cblabel "Survived"
set style fill solid 1.0 border lc rgb "white"

plot sex using (0):0:(0.5):(0.5):($3/$2*100):ytic(1) with boxxyerror lc palette, \
     ''  using (1):0:(0.5):(0.5):($6/$5*100)         with boxxyerror lc palette, \
     ''  using (0):0:(sprintf("%.0f%%", $3/$2*100))  with labels, \
     ''  using (1):0:(sprintf("%.0f%%", $6/$5*100))  with labels
unset output

Run listing 54 from Bash in terminal/.

Listing 54 — Bash / shell — Run plots/by_class_and_sex_heatmap.gp from Bash
gnuplot -d plots/by_class_and_sex_heatmap.gp
by_class_and_sex_survival_heatmap.svg
Figure 17: Survival rate by class and sex, as a heatmap. (by_class_and_sex_survival_heatmap.svg)

Reversing the y range puts first class at the top. The palette represents 0–100%; its darkest colour remains light enough for the dark labels.

Practice this topic in Challenge 13: Read the same rates in two encodings.

Age and class — the large heatmap

Now show 25 three-year age bands across all four groups, including Crew. Use titanic_by_age_and_class.csv, whose columns are age_midpoint, class_code, survival_pct, n and survivors. These 100 cells describe 2,205 people with known age; the two unknown ages are excluded. Percentages and group counts are already calculated.

Use the downloaded data/titanic_by_age_and_class.csv. The script's CSV separator and column-header settings handle the file format; do not additionally skip the first observation.

In your editor, create plots/by_age_and_class_heatmap.gp inside terminal/. Enter listing 55, then save the file before running it.

Listing 55 — Editor — plots/by_age_and_class_heatmap.gp
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_age_and_class.csv'
set terminal pngcairo size 2400,650 enhanced font "Sans,12"
set output 'figures/by_age_and_class_survival_heatmap.png'
set datafile missing "NA"
set title "Titanic survival by age and class\nPercent survived; n = group count"
set xlabel "Age band (completed years)"
set xrange [-0.5:74.5]
set yrange [4.5:0.5]
set xtics ("0-2" 1, "3-5" 4, "6-8" 7, "9-11" 10, "12-14" 13, \
           "15-17" 16, "18-20" 19, "21-23" 22, "24-26" 25, "27-29" 28, \
           "30-32" 31, "33-35" 34, "36-38" 37, "39-41" 40, "42-44" 43, \
           "45-47" 46, "48-50" 49, "51-53" 52, "54-56" 55, "57-59" 58, \
           "60-62" 61, "63-65" 64, "66-68" 67, "69-71" 70, "72-74" 73)
set ytics ("1st" 1, "2nd" 2, "3rd" 3, "Crew" 4)
set tics out nomirror
set key off
set palette defined (0 "#440154", 25 "#3b528b", 50 "#21918c", \
                     75 "#5ec962", 100 "#fde725")
set cbrange [0:100]
set cblabel "Survival (%)"
set style fill solid 1.0 border lc rgb "white"

# Column headers are already enabled; each cell spans three years and one class.
plot data using 1:2:($1-1.5):($1+1.5):($2-0.5):($2+0.5):3 \
        with boxxyerror linecolor palette, \
     data using 1:2:(sprintf("%.1f%%\nn=%d", $3, $4)):($3 >= 55 ? 0x111111 : 0xffffff) \
        with labels textcolor rgb variable font "Sans,12"
unset output

Run listing 56 from Bash in terminal/.

Listing 56 — Bash / shell — Run plots/by_age_and_class_heatmap.gp from Bash
gnuplot -d plots/by_age_and_class_heatmap.gp
by_age_and_class_survival_heatmap.png
Figure 18: Survival by three-year age band and class, including crew. Blank cells have no records. (by_age_and_class_survival_heatmap.png)

This form of boxxyerror uses seven values: x/y centre, left/right bounds, lower/upper bounds and colour value. Each cell spans three years and one class category.

The second layer adds the percentage and group count, switching text colour for contrast. NA leaves the 12 empty groups blank, unlike an observed 0%. Small n matters: 100% with n=1 means one survivor out of one person. Open the PNG at full size to inspect all 25 columns.

Practice this topic in Challenge 14: Show when a percentage has little support.

Age trends in four panels

Use the same age/class summary for a four-panel plot, including Crew. Identical axes make the age patterns easier to compare. Points represent three-year groups, not individual ages, and the labels report group counts.

In your editor, create plots/by_age_and_class_panels.gp inside terminal/. Enter listing 57, then save the file before running it.

Listing 57 — Editor — plots/by_age_and_class_panels.gp
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
header_rows = 0 # columnheaders already removes the CSV header
data = "data/titanic_by_age_and_class.csv"
set terminal pngcairo size 2200,1100 enhanced font "Sans,11"
set output "figures/by_age_and_class_survival_panels.png"
set datafile missing "NA"
set xrange [-0.5:74.5]
set yrange [-8:115]
set xtics 1,3,73
set ytics 0,20,100
set xlabel "Age band centre (whole years)"
set ylabel "Survival (%)"
set grid ytics
set tics out nomirror
set key off
set style line 1 lc rgb "#0072B2" lw 2 pt 7
set style line 2 lc rgb "#D55E00" lw 2 pt 5
set style line 3 lc rgb "#009E73" lw 2 pt 9
set style line 4 lc rgb "#CC79A7" lw 2 pt 13
classes = "1st 2nd 3rd Crew"

set multiplot layout 2,2 rowsfirst title "Titanic survival by class\nLabels show group counts (n)" font "Sans,16"
do for [c=1:4] {
    set title word(classes,c)
    plot data every 4::(header_rows+c-1) using 1:3 with linespoints linestyle c pointsize 1.1, \
         data every 4::(header_rows+c-1) using 1:3:(sprintf("n=%d", $4)) \
             with labels offset char 0,1 textcolor rgb "#333333" font "Sans,10"
}
unset multiplot
unset output

Run listing 58 from Bash in terminal/.

Listing 58 — Bash / shell — Run plots/by_age_and_class_panels.gp from Bash
gnuplot -d plots/by_age_and_class_panels.gp
by_age_and_class_survival_panels.png
Figure 19: Survival by age in four class panels; labels give each group's count. (by_age_and_class_survival_panels.png)

set multiplot layout 2,2 arranges four panels. The loop selects class codes 1–4, and every 4::(header_rows+c-1) takes every fourth record for class c. Keep the published ordering: each age band contains all four class codes, even where a group is empty.

word(classes,c) chooses the title and linestyle c its colour and marker. unset multiplot finishes the layout. The lines guide the eye; they are not a fitted model. Missing groups are not zero-percent survival, and extreme percentages from small groups deserve caution.

set datafile columnheaders removes the CSV header before row selection, so header_rows is 0. Using 1 as well would shift the selected class.

Practice this topic in Challenge 15: Do lines imply more than the data show?.

13 — A dumbbell chart

A dumbbell connects the men's survival rate to the women's rate for each class or crew group. The line's length is their gap.

New here: horizontal displacement between two already familiar rates. As in the rate example, survival_percent names the repeated arithmetic. Column 3 divided by 2 is the women's rate; column 6 divided by 5 is the men's rate. Subtracting them gives percentage points, not percent change. The colours and larger points below are the report layer, not new data.

with vectors takes x:y:dx:dy: a starting point and a displacement. Start at the men's rate, use the difference between the rates as dx, and keep dy at zero. nohead removes the arrowhead to leave a connecting line.

In your editor, create plots/by_class_sex_gap.gp inside terminal/. Enter listing 59, then save the file before running it.

Listing 59 — Editor — plots/by_class_sex_gap.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 820,400 font 'sans,14'
set output 'figures/by_class_and_sex_survival_gap.svg'
load 'style.gp'
survival_percent(survived, aboard) = 100.0 * survived / aboard
unset grid
set grid xtics lc rgb "#d5d8dc" lw 0.8
set xrange [0:100]
set yrange [3.6:-0.6]
set format x "%.0f%%"
set xlabel "Survived"
set ytics scale 0
set key above
set linetype 1 lc rgb "#14507d"   # women
set linetype 2 lc rgb "#c47a2c"   # men

plot sex using (survival_percent($6,$5)):0:(survival_percent($3,$2) - survival_percent($6,$5)):(0):ytic(1) \
         with vectors nohead lw 4 lc rgb "#d5d8dc" notitle, \
     ''  using (survival_percent($3,$2)):0 with points pt 7 ps 2 lc 1 title "Women", \
     ''  using (survival_percent($6,$5)):0 with points pt 7 ps 2 lc 2 title "Men"
unset output

Run listing 60 from Bash in terminal/.

Listing 60 — Bash / shell — Run plots/by_class_sex_gap.gp from Bash
gnuplot -d plots/by_class_sex_gap.gp
by_class_and_sex_survival_gap.svg
Figure 20: The gap between women's and men's survival rates. (by_class_and_sex_survival_gap.svg)

Practice this topic in Challenge 16: Add an honest name for the gap.

14 — Five measures on a spider plot

A spider or radar plot shows several measures of one thing at once: one axis per measure, one polygon per row. Each class or crew group is a polygon across five percentages derived from the class-and-sex counts. This course uses gnuplot 6.0 or newer, including for spider plots.

Each plot clause adds an axis. All axes run from 0 to 100, and set paxis ... label supplies each spoke label and its rotation.

In your editor, create plots/by_class_profile.gp inside terminal/. Enter listing 61, then save the file before running it.

Listing 61 — Editor — plots/by_class_profile.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 820,640 font 'sans,14'
set output 'figures/by_class_profile_spider.svg'
load 'style.gp'

  set title "Titanic survival by class and sex" offset 0,1 font ",18"
  unset border; unset tics; unset grid
  set spiderplot

  # One axis per plot clause below: five measures, all in per cent
  set for [p=1:5] paxis p range [0:100]
  set paxis 1 tics 0,25,100 format "%.0f%%"
  set for [p=2:5] paxis p tics format ""

  # Spoke labels, each at right angles to its spoke. Spoke p points at
  # 90 - (p-1)*72 degrees; the label is turned 90 degrees less than that,
  # and turned over by 180 where it would otherwise stand upside down.
  set paxis 1 label "All survived"   rotate by   0
  set paxis 2 label "Women survived" rotate by -72
  set paxis 3 label "Men survived"   rotate by  36
  set paxis 4 label "Women aboard"   rotate by -36
  set paxis 5 label "Men aboard"     rotate by  72

  set grid spiderplot lc rgb "#d0d0d0" dt 2 lw 1
  set style spiderplot fs transparent solid 0.18 border lw 2
  set key outside right center

  # One polygon per row of the table, i.e. per class
  set linetype 1 lc rgb "#14507d" lw 3
  set linetype 2 lc rgb "#c47a2c" lw 3
  set linetype 3 lc rgb "#3f8f5a" lw 3
  set linetype 4 lc rgb "#8064a2" lw 3

  # Columns: 2 women aboard, 3 women survived, 5 men aboard, 6 men survived
  # key(1) names each polygon after column 1 (1st, 2nd, 3rd, Crew)
  plot \
       sex using (($3+$6)/($2+$5)*100):key(1) with spiderplot notitle, \
       sex using ($3/$2*100)                  with spiderplot notitle, \
       sex using ($6/$5*100)                  with spiderplot notitle, \
       sex using ($2/($2+$5)*100)             with spiderplot notitle, \
       sex using ($5/($2+$5)*100)             with spiderplot notitle
unset output

Run listing 62 from Bash in terminal/.

Listing 62 — Bash / shell — Run plots/by_class_profile.gp from Bash
gnuplot -d plots/by_class_profile.gp
by_class_profile_spider.svg
Figure 21: Five measures for each class. (by_class_profile_spider.svg)

The last two axes are group composition—the shares of women and men aboard—not survival rates. They sum to 100 and therefore mirror each other. Crew is overwhelmingly male and has the lowest overall survival rate, 23.7%. This describes the group; it does not establish why people survived.

Practice this topic in Challenge 17: Test whether polygon shape is evidence.

15 — Class size and fate in a donut

The inner ring divides everyone aboard into three passenger classes and crew. The outer ring splits each group into survivors and deaths. It combines the two messages of the stacked counts: group size and outcome.

This is an optional report example, not the next minimal plotting command. New here are ring boundaries and layered angular intervals; the calculations still use the same counts and whole as the pie. Read the inner-ring pass first, then the two outcome passes, and only then the label passes. The stacked chart remains a simpler way to communicate the same quantities.

Draw real ring segments with with sectors: class segments run from hole_radius to split_radius; outcome segments run from split_radius to outer_radius. The centre stays empty without a white covering disc. Each pass resets pos so both rings share the same group boundaries.

Here the sectors columns are start angle:inner radius:angular width:radial width, optionally followed by a colour. Where circles takes start and end angles, sectors takes a start angle and angular width. In polar mode, labels take angle:radius:text directly. See the gnuplot 6 manual, Sectors.

In your editor, create plots/by_class_donut.gp inside terminal/. Enter listing 63, then save the file before running it.

Listing 63 — Editor — plots/by_class_donut.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 600,600 font 'sans,11'
set output 'figures/by_class_and_outcome_donut.svg'
load 'style.gp'
unset polar
set autoscale
stats data using 2 nooutput
total = STATS_sum
set angles degrees
ang(n) = n * 360.0 / total

set polar
set theta right ccw
unset raxis
unset border; unset tics; unset grid
set key below horizontal center
set size ratio -1
set xrange [-1.35:1.35]
set yrange [-1.35:1.35]
set rrange [0:1.35]
set style fill solid 1.0 border lc rgb "white"
set linetype 3 lc rgb "#2b3036"
set linetype 4 lc rgb "#4a5058"
set linetype 5 lc rgb "#6c737c"
set linetype 6 lc rgb "#515969"

hole_radius = 0.40
split_radius = 0.72
outer_radius = 1.0
label_radius = (hole_radius + split_radius)/2
rate_radius = 1.16

# sectors: start angle : inner radius : angular width : radial width
# Each outcome pass advances pos by the WHOLE group, not just that outcome.
plot pos = 0, \
     data using (pos):(split_radius): \
          (span = ang($3), pos = pos + ang($2), span): \
          (outer_radius - split_radius) \
          with sectors lc rgb "#14507d" title "Survived", \
     pos = 0, \
     data using (pos + ang($3)):(split_radius): \
          (span = ang($4), pos = pos + ang($2), span): \
          (outer_radius - split_radius) \
          with sectors lc rgb "#dfe3e8" title "Died", \
     pos = 0, \
     data using (pos):(hole_radius): \
          (span = ang($2), pos = pos + span, span): \
          (split_radius - hole_radius):($0 + 3) \
          with sectors lc variable notitle, \
     pos = 0, \
     data using (mid = pos + ang($2)/2, pos = pos + ang($2), mid): \
                (label_radius):(strcol(1)) \
          with labels tc rgb "white" center notitle, \
     pos = 0, \
     data using (mid = pos + ang($2)/2, pos = pos + ang($2), mid): \
                (rate_radius):(sprintf("%.0f%%", 100.0 * $3 / $2)) \
          with labels tc rgb "#2b3036" center notitle
unset output

Run listing 64 from Bash in terminal/.

Listing 64 — Bash / shell — Run plots/by_class_donut.gp from Bash
gnuplot -d plots/by_class_donut.gp
by_class_and_outcome_donut.svg
Figure 22: Class size and survival in one chart. (by_class_and_outcome_donut.svg)

The first pass draws survivors in blue; the second starts after those survivors and fills the remaining group angle in light grey. Both passes advance pos by the whole group's angle, but return only the outcome's angular width. These widths use counts divided by the total aboard, so survivors and deaths together align with the corresponding inner segment. set polar lets the labels use angle:radius:text directly.

The plot uses square output dimensions rather than the wide default. Class names sit inside the inner ring and survival rates outside. Crew is the largest inner wedge, 40.3%, with a mostly grey outer rim. First class is 14.7% of those aboard and has a mostly blue rim.

Change hole_radius to adjust the hole. Keep 0 < hole_radius < split_radius < outer_radius=; the radial widths are differences between these boundaries, not absolute outer radii. label_radius keeps class names halfway across the inner ring.

Practice this topic in Challenge 18: Separate decoration from information.

16 — Tiny age-count bars

A sparkline is a chart about the size of a word: no axes, labels or legend, just the data's shape inside a sentence. It answers “what does the trend look like?” without interrupting the reading.

Use titanic_by_age.csv with its 25 three-year groups from 0–2 to 72–74. Both sexes, all classes and crew are included; two unknown ages are excluded. These bars show how many people were in each age group, not how many survived.

In your editor, create plots/by_age_counts.gp inside terminal/. Enter listing 65, then save the file before running it.

Listing 65 — Editor — plots/by_age_counts.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_age.csv'
set terminal svg size 120,24
set output 'figures/by_age_people_sparkline.svg'
load 'style.gp'
unset border; unset tics; unset key; unset grid
set margins 0, 0, 0, 0
set yrange [0:*]
set offsets 0.5, 0.5, 0, 0
set boxwidth 0.7
set style fill solid 1.0 noborder
plot data using 0:2 with boxes lc 1
unset output

Run listing 66 from Bash in terminal/.

Listing 66 — Bash / shell — Run plots/by_age_counts.gp from Bash
gnuplot -d plots/by_age_counts.gp
by_age_people_sparkline.svg
Figure 23: Age-group counts including crew (by_age_people_sparkline.svg)

The SVG is only 120 by 24 pixels. Removing axes, tics, legend and grid, and setting all margins to zero, leaves room for the data. set offsets preserves a little horizontal space so the end bars are not cut off. Pseudo-column 0 places the rows at positions 0, 1, 2, … because the age labels in the first column are text.

Practice this topic in Challenge 19: Make a tiny chart understandable.

17 — An inline survival trend

The same age groups now show the survival percentage as a line—not the number of survivors. Keep set yrange [0:100] so the vertical scale remains comparable between survival sparklines. Without it, small differences could be stretched to the full height.

In your editor, create plots/by_age_survival_rate.gp inside terminal/. Enter listing 67, then save the file before running it.

Listing 67 — Editor — plots/by_age_survival_rate.gp
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_age.csv'
set terminal svg size 120,24
set output 'figures/by_age_survival_sparkline.svg'
load 'style.gp'
unset border; unset tics; unset key; unset grid
set margins 0, 0, 0, 0
set yrange [0:100]
set offsets 0.3, 0.3, 0, 0
plot data using 0:3 with lines lw 1.5 lc 1
unset output

Run listing 68 from Bash in terminal/.

Listing 68 — Bash / shell — Run plots/by_age_survival_rate.gp from Bash
gnuplot -d plots/by_age_survival_rate.gp
by_age_survival_sparkline.svg
Figure 24: Age-group survival trend (by_age_survival_sparkline.svg)

The supplied browser gallery displays the tiny SVGs at sparkline size. When placing them in another document, use a height close to its text line rather than stretching them to a full-width chart.

The highest published rate is 67.9% for ages 3–5, compared with 61.8% for ages 0–2. The three oldest groups have no recorded survivors, but their counts are small. The trend is not monotonic.

Practice this topic in Challenge 20: See what an automatic scale hides.

18 — Run the complete set again

After working through every plot, rerun the complete sequence. Check that all eighteen figures are present, including the large age/class heatmap and four-panel age comparison. Revisit the data description and the interpretation whenever changing the dataset version.

A reproducible rerun should not depend on invisible settings from an earlier interactive session. Keep the published inputs, explicit plotting commands and shared style together.

After saving all eighteen plot files and style.gp, create all.gp directly inside terminal/. Listing 69 contains one load command per plot; it does not download data, calculate summaries or create missing scripts.

Listing 69 — Editor — all.gp — Run all eighteen saved plots
# Run from terminal/: gnuplot -d all.gp
load 'plots/by_class_aboard.gp'
load 'plots/by_class_survivors.gp'
load 'plots/by_class_counts.gp'
load 'plots/by_class_styled.gp'
load 'plots/by_class_shared.gp'
load 'plots/by_class_stacked.gp'
load 'plots/by_class_shares.gp'
load 'plots/by_class_survival_rate.gp'
load 'plots/by_class_and_sex.gp'
load 'plots/by_class_pie.gp'
load 'plots/by_class_and_sex_heatmap.gp'
load 'plots/by_age_and_class_heatmap.gp'
load 'plots/by_age_and_class_panels.gp'
load 'plots/by_class_sex_gap.gp'
load 'plots/by_class_profile.gp'
load 'plots/by_class_donut.gp'
load 'plots/by_age_counts.gp'
load 'plots/by_age_survival_rate.gp'

Run listing 70 only after every script and all five CSVs exist.

Listing 70 — Bash / shell — Regenerate all eighteen plots
gnuplot -d all.gp

Open the SVGs and PNGs in figures/, or reload the supplied index.html gallery. If a script is missing, return to its topic and save it first.

For individual reruns from Bash, use listing 71.

Listing 71 — Bash / shell — Examples of running saved plot files
gnuplot -d plots/by_class_survivors.gp
gnuplot -d plots/by_class_counts.gp
gnuplot -d plots/by_class_styled.gp
gnuplot -d plots/by_class_survival_rate.gp

At the gnuplot prompt, use listing 72 instead. load reads the saved file again; replot repeats the most recent plot command with the current settings. Save edits before loading a file.

Listing 72 — At gnuplot> — Reload a saved file at the gnuplot prompt
gnuplot> load 'plots/by_class_survival_rate.gp'
gnuplot> # Edit the file in your editor, then load it again:
gnuplot> load 'plots/by_class_survival_rate.gp'
gnuplot> exit

Practice this topic in Challenge 21: Prove that an edit reaches the output.

19 — What the figures say

Of the 2,207 people in the published class table, 711 survived—about one in three. Passenger survival rates were 62.0% in first class, 41.5% in second and 25.5% in third; Crew's rate was 23.7%.

Crew also has the largest death count, 679, followed by third class with 528. Together these groups account for most deaths. Compare the stacked counts with the rate bars: a large number of deaths and a low survival rate are related but different descriptions.

Within every class and crew group, women survived at a higher rate than men. Among passengers, women's rates were approximately 97%, 89% and 49% in first, second and third class. Men's rates were approximately 34%, 13% and 15%. Among crew, 20 of 23 women survived (87%), compared with 191 of 867 men (22%). Read the small female crew count alongside its rate.

The age heatmap, four panels and sparkline show age patterns at different levels of detail. Only the two unknown ages are excluded from age-based plots; the class and sex plots include all records. Empty groups are not zero survival.

These numbers describe who survived, not why. The selected columns do not record deck location, access to lifeboats, evacuation timing or individual decisions. Do not infer causal effects of age, class or sex from these plots. Consult the dataset provenance before treating the supplied records as a definitive historical reconstruction.

The reported values belong to the pinned release. A new dataset version requires reviewing this interpretation as well as rerunning the code.

Practice this topic in Challenge 22: Write a claim the figure supports.

20 — Troubleshooting

Work through the same questions when a plot surprises you: is the input available, is it read correctly, are the selected columns appropriate, and have you rerun and reopened the right output?

Symptom Check
command not found Install the required program or correct PATH, then restart the shell or Emacs.
Only one series appears Check the comma separating plot elements and the selected columns.
Unexpected percentage Check the numerator and that group's own denominator.
Missing first observation Do not skip a row twice after handling the header.
Missing output figure Check the output path and directory, and that the plot ran successfully.
Old figure remains visible Rerun the code, then refresh the image display or browser.
Unexpected colours or axes Start from reset and the explicit shared settings; check plot-specific overrides.
Spider syntax fails Check that gnuplot is version 6.0 or newer.
An empty heatmap cell Inspect n and NA before interpreting it as zero survival.

Stay in terminal/; confirm that all five files are under data/ and the figures/ directory exists. Use comma separation and column headers after reset. The scripts write SVG or PNG rather than opening a plot window. Keep Bash commands and gnuplot statements in their respective contexts.

If a browser still shows an older result, reload after running the correct script. The only network step in the lesson is downloading the CSVs.

Practice this topic in Challenge 23: Catch a plausible but wrong plot.

Further reading