Plots with ROOTs
Data visualization with gnuplot | Live coding in terminal
0 — Start here
Follow one question at a time: inspect the data, make a first plot, improve its readability, then choose different charts for different questions. Both editions use the same five published CSV datasets, include crew and work through the same eighteen figures. No local AWK preparation is needed.
Begin with the independent Meet gnuplot chapter below; it needs no CSV files. Chapters 1–11 then apply the tool to counts, percentages and composition. Chapters 12–17 are chart extensions: choose them to answer a question, not because every option must be learned on the first day.
The workflow follows the ROOT principles: Robust, Open, Ongoing, Time-tested. Keep the data, plotting instructions and resulting figures understandable. Change one setting at a time and compare the result.
Use gnuplot 6.0 or newer with SVG and PNG output, and curl for downloads. Retrieve the data before an offline lesson. Once the downloads are available, all plotting can run without further network access.
Work in a terminal
Use Bash on macOS or Linux, WSL or Git Bash. Git Bash provides the shell
commands; install gnuplot separately and add it to PATH. Use a plain-text
editor and a browser to view SVG or PNG figures. Students do not need
Emacs, LaTeX or Inkscape to run these examples.
During live coding, create each .gp file in the editor, save it, then run
its printed gnuplot command. The supplied terminal/ directory is a worked
solution for comparison or catch-up, not a prerequisite for writing the plots.
The listing titles distinguish three places to work:
| Listing label | Where to enter the code |
|---|---|
| Bash / shell | At your normal terminal prompt; $ here represents the shell prompt. |
| At gnuplot> | Inside the running gnuplot program, after starting gnuplot -d. |
| Editor | In the named file; save it before running it from the shell. |
Prompts such as $ and gnuplot> are orientation marks, not commands to
copy. Omit them when entering code. Indented lines without a prompt
continue a command or enter its data; finish those lines before starting
the next command. Saved .gp files never contain prompts.
Check the programs
Run listing 1. command -v checks whether a program is
available through PATH. If it says MISSING, install the program or fix
its path, then open a new terminal and repeat the check.
for program in bash gnuplot curl head cat mkdir; do
if command -v "$program" >/dev/null 2>&1; then
printf 'OK %s\n' "$program"
else
printf 'MISSING %s - install it or add it to PATH\n' "$program"
fi
done
Check the gnuplot build
Check gnuplot's version and both output terminals with listing 2. Continue only when the checks succeed.
set -e
gnuplot --version
gnuplot -d -e 'if (GPVAL_VERSION < 6.0) { exit error "Need gnuplot 6.0 or newer" }; set terminal svg; print "OK: version and SVG output"'
gnuplot -d -e 'set terminal pngcairo; print "OK: PNG output"'
Enter the working directory
Run listing 3 once from the course folder. Keep all later
commands in terminal/. Quote paths containing spaces and use forward
slashes, for example cd "/c/Users/Your Name/course/terminal" in Git Bash.
mkdir -p terminal cd terminal mkdir -p data figures plots export LC_NUMERIC=C
LC_NUMERIC=C selects decimal-point numeric formatting for programs
started by this shell. It does not change the system language or the CSV
separator. export passes the setting to child processes. If LC_ALL is set,
it overrides LC_NUMERIC; run unset LC_ALL first in that case.
Practice this topic in Challenge 1: A working plotting environment.
Before Titanic — Meet gnuplot, from a first command to a simple chart
This chapter needs only gnuplot. No Titanic files, shared style or network connection is used. Start with a function; then use one small synthetic table. The short continuations below change one thing at a time. Predict the change, run it and look at the plot before proceeding.
First, draw a function
First run listing 4 at your Bash / shell prompt.
$ gnuplot -d
The prompt now changes to gnuplot>: you are inside gnuplot, not Bash.
Enter 5 there, then 6. Do not type
gnuplot -d again at this prompt. Gnuplot calls its output device a
terminal: dumb draws
with text characters directly in Bash, Git Bash or another terminal, with
no additional window. This is a rough preview, not an SVG-quality image.
gnuplot> set terminal dumb size 79,24
The SVGs printed in this guide are illustrations. Their caption filenames
identify those illustrations; typing plot does not save them. Use the
explicit export step below when you want an SVG file of your own.
gnuplot> reset gnuplot> plot sin(x)
first.svg)
reset starts with known plotting settings. plot samples the built-in
function; x is its argument. Sine uses radians by default. Gnuplot chooses
the initial ranges automatically.
Optional — Smoother function curves in SVG
Gnuplot draws a function by joining calculated points. If the sine curve looks angular when you zoom in, add
set samples 500afterresetand beforeplot sin(x)in 6, then rerun the complete example. In an ongoing terminal session, change the sample count, then repeat the whole SVG export sequence: reopen the output filename,replot, and close it withunset output.show sampleschecks the active setting; a laterresetrestores the default.More samples refine the curve's geometry, not the display's pixels. Once the segments are smaller than screen pixels,
10000samples may look no better than500while producing a much larger SVG. Reopen the newly exported SVG in a browser and zoom in to distinguish straight segments from pixelated edges. This setting does not improve the character-grid resolution ofdumbor add points to ordinary tabular-data plots.
Save the current plot as SVG
After 6, enter 7 in the same gnuplot
session. The figures/ directory was created in the setup chapter; keep
working from terminal/. set terminal svg selects the file format,
set output names the file, and replot draws the current plot into it.
unset output closes the file, but SVG remains the selected output
format. We keep it selected for the rest of this chapter.
gnuplot> set terminal svg size 900,400 gnuplot> set output "figures/first.svg" gnuplot> replot gnuplot> unset output
Open figures/first.svg in a browser for the full graphic. Reuse this
export step after any later plot, changing the filename to match that
example (for instance figures/two-series.svg). Reusing a filename
overwrites that file. There is no need to select the SVG terminal again
unless you deliberately change the output format.
One SVG document per file: do not call replot repeatedly while the
same SVG output file stays open. Each call appends another complete SVG,
causing an XML error such as “junk after document element”. After changing
a setting such as set samples 500, repeat all of listing
7. Its set output reopens and replaces the file;
unset output finishes it. This also repairs a file containing several
concatenated SVG documents. Then reload the file in your browser.
For each following plotting listing, including the short continuations:
enter set output "figures/first.svg" before its commands, and unset output
after its last plot or replot. Reload the browser to see the change.
Choose a different filename to keep a previous version. For example,
8 now runs between those two output commands. The same rule
applies to the later synthetic-data examples. With SVG selected and no
output file open, a bare plot or replot writes SVG markup to the shell;
it does not update the saved file.
One change, then look again
Listings 8 through 13 are continuations in the same
interactive session, after 6. Enter each pair of commands
in order. set changes a setting; replot redraws the previous plot with
that setting. It does not create a plot from nothing.
In Org, these short fragments are deliberately not executable: separate
:session none blocks do not remember the previous plot. Instead, add
each new set line just before plot sin(x) in a copy of 6
and rerun that complete block. Give the copy a new block name and output
path. No persistent Babel session is needed.
gnuplot> set title "A sine wave" gnuplot> replot
gnuplot> set xlabel "Angle (radians)" gnuplot> replot
gnuplot> set ylabel "sin(x)" gnuplot> replot
gnuplot> set xrange [0:2*pi] gnuplot> replot
gnuplot> set yrange [-1.2:1.2] gnuplot> replot
gnuplot> set grid gnuplot> replot
Listing 14 is the complete, independently runnable checkpoint after the six changes. It does not depend on 6.
gnuplot> reset gnuplot> set title "A sine wave" gnuplot> set xlabel "Angle (radians)" gnuplot> set ylabel "sin(x)" gnuplot> set xrange [0:2*pi] gnuplot> set yrange [-1.2:1.2] gnuplot> set grid gnuplot> plot sin(x)
function-settings.svg)Practice this step in Intro A: Change the visible range.
One synthetic table, three numeric columns
The following six rows are synthetic, not observations. Column 1 is a step number; columns 2 and 3 are two invented measurements in arbitrary units. All later data examples in this chapter reuse this one table.
Listing 15 is another complete checkpoint. The lines between
$Measurements << EOD and EOD define a gnuplot data block: a little
table stored in memory, without a separate file. Enter the whole definition
before plotting it. Spaces separate its numeric columns.
gnuplot> reset
gnuplot> $Measurements << EOD
1 2 3
2 4 4
3 3 5
4 6 5
5 5 7
6 8 6
EOD
gnuplot> set xlabel "Step"
gnuplot> set ylabel "Measurement (arbitrary units)"
gnuplot> plot $Measurements using 1:2 with points title "Series A"
points.svg)
using 1:2 means x from column 1, y from column 2. It is not a row
range. For example, row 3 places a point at x = 3, y = 3.
title "Series A" names the legend entry, not the whole chart.
Change only the drawing style
Run 16, then 17 in the same prompt after 15. In Org, replace only the final plot line in 15 and rerun the whole block. These are replacement lines, not independent examples: the data block and axis labels must still be present.
gnuplot> plot $Measurements using 1:2 with lines title "Series A"
lines.svg)gnuplot> plot $Measurements using 1:2 with linespoints title "Series A"
linespoints.svg)Practice this step in Intro B: Change the drawing style.
Add one series, keep the axes
Column 3 supplies a second series on the same axes, called “Series B”.
The comma separates two series in one plot command; a backslash continues
the command on the next line. Listing 18 draws both series:
replace the final plot line in 15, or enter it in that session.
gnuplot> plot $Measurements using 1:2 with linespoints title "Series A", \
$Measurements using 1:3 with linespoints title "Series B"
two-series.svg)At step 6, Series A is 8 and Series B is 6. Both series share the same x values. No percentages or grouping have been introduced yet.
Practice this step in Intro C: Add a second series.
From the prompt to a saved script
For a reproducible checkpoint, copy the complete body of 15
into plots/intro.gp in your editor and save it (not just a continuation).
Optionally replace its final plot line with 18. Use
listing 19 inside gnuplot to return to Bash.
gnuplot> exit
You are now back at the shell prompt. From terminal/,
listing 20
runs this script as a text preview in Bash; listing 21 saves
an SVG instead. Choose the output you need. Neither opens a plot window.
gnuplot -d -e 'set terminal dumb size 79,24' plots/intro.gp
The two output settings in 21 are file-export plumbing; they do not change the data. Gnuplot closes the output file when it exits.
gnuplot -d -e 'set terminal svg; set output "figures/intro.svg"' plots/intro.gp
Open figures/intro.svg in a browser. Save, rerun and reload after an edit.
The later Titanic chapters use the same editor/save/run cycle, but their
scripts already contain set terminal and set output commands. Those
scripts save the named file directly; the command-line preview setting
above would be overridden. See 31 for a text
preview of the first Titanic script.
1 — Download and inspect the data
We now apply the plotting commands we already know to a real dataset. The introduction supplied the tool; from here on we ask questions about people, groups and survival. The new step in this chapter is retrieving and interpreting published columns, not calculating a new dataset.
Use Titanic Passengers and Crew: Demographics and Survival Data, version 1.0.1. This version-specific record fixes which data the lesson uses. All counting and age grouping have already been done; no conversion or preprocessing is needed.
| Published CSV | Rows | Contents and use |
|---|---|---|
titanic.csv |
2,207 | Individual records; inspect the underlying data |
titanic_by_class.csv |
4 | Class/crew counts; bars, rates, pie and donut |
titanic_by_sex.csv |
4 | Counts by class/crew and sex; comparisons, heatmap, dumbbell and spider |
titanic_by_age_and_class.csv |
100 | Three-year age/class cells; large heatmap and four panels |
titanic_by_age.csv |
25 | Three-year age groups; sparklines |
The person-level columns are class, sex, age, survived and embarked.
There are 2,207 records: 1,317 passengers and 890 crew, with 711 survivors.
The categories 1st, 2nd, 3rd and Crew appear in that order.
Crew is a separate group, not a fourth passenger class.
NA marks missing values, not zero. The two unknown ages are excluded from
age summaries, leaving 2,205 people. The dataset description documents its
provenance, transformations and limitations; these are descriptive records,
not evidence of causal effects.
Retrieve the five files
From terminal/, run listing 22 to save the published CSVs
in data/. The failure guard stops the shell if a download fails; resolve
the error and rerun it before plotting.
for file in titanic.csv titanic_by_class.csv titanic_by_sex.csv \
titanic_by_age_and_class.csv titanic_by_age.csv; do
curl --fail --location \
"https://zenodo.org/api/records/22983324/files/$file/content" \
--output "data/$file" || exit 1
done
--output writes each response to its named CSV file. The loop downloads
each file once; it does not calculate or regroup the data.
--fail reports HTTP errors and --location follows redirects.
Where used, --silent --show-error hides the progress meter but retains errors.
Inspect individual records
Run listing 23. head displays the requested lines without
truncating or modifying the downloaded file.
head -n 5 data/titanic.csv
A short preview of the records:
class,sex,age,survived,embarked 1st,female,2,0,S 1st,female,13,1,S 1st,female,16,1,C 1st,female,16,1,S
survived is 1 for survival and 0 otherwise. Ages can be fractional;
the two missing ages remain NA. Names and identifiers are not present.
Inspect the class-and-crew counts
Run listing 24 to inspect the same ready-made counts.
cat data/titanic_by_class.csv
Class,Aboard,Survived,Died 1st,324,201,123 2nd,284,118,166 3rd,709,181,528 Crew,890,211,679
Columns 2, 3 and 4 are aboard, survived and died. All ages and both sexes
are included. Every row satisfies survived + died = aboard; for Crew,
211 + 679 = 890. First class has 324 people and 201 survivors.
Inspect the class-and-sex counts
The file is already downloaded; listing 25 displays it.
cat data/titanic_by_sex.csv
Columns 2–4 are women aboard, survived and died; columns 5–7 are the
corresponding men's counts. These categories include children.
A women's survival rate uses $3/$2; a men's uses $6/$5.
The first row is 1st,144,139,5,180,62,118: 324 people and 201 survivors.
The Crew row is Crew,23,20,3,867,191,676: 890 people and 211 survivors.
Adding the women's and men's counts reproduces the class table.
Inspect the age-and-class summary
Preview the downloaded summary with listing 26.
head -n 9 data/titanic_by_age_and_class.csv
Its columns are age_midpoint, class_code, survival_pct, n and
survivors. The 100 rows cross 25 three-year age bands with four groups.
Class codes 1–4 mean first, second, third and crew. Labels 1, 4, …, 73
represent completed-year groups 0–2, 3–5, …, 72–74.
Rows cycle through class codes 1, 2, 3, 4 within each age band. Keep that
order for the four-panel plot. Empty cells are retained with n=0 and
survival_pct=NA; these are not zero-percent survival observations.
Inspect the age-only summary
cat data/titanic_by_age.csv
The columns are Age, People and Survived %. All 25 groups run in age
order, from 0–2 through 72–74, with no combined 60+ group. The first row is
0-2,34,61.8 and the last is 72-74,2,0.0. Counts sum to 2,205; both sexes,
all passenger classes and crew are pooled. Percentages have one decimal
place; NaN would mark an empty age-only group.
Keep headers separate from observations
Each script uses set datafile separator "," for CSV and
set datafile columnheaders for its header. Put those settings after
reset. With headers enabled, row index 0 is already the first observation:
do not add every ::1, which would skip it. The variable data is just a
string holding a relative filename such as data/titanic_by_class.csv.
Practice this topic in Challenge 2: Match a question to a dataset.
2 — A first plot
Start with everyone aboard by class and crew. Each table row is one category and its aboard count supplies the bar height. Keep gnuplot's default appearance for this first experiment; later topics make each design choice explicit.
In your editor, create plots/by_class_aboard.gp inside terminal/.
Enter listing 28 and save it.
# Generated from plots-with-roots.dtx; run from the terminal directory. reset set encoding utf8 set datafile separator "," set datafile columnheaders data = 'data/titanic_by_class.csv' set terminal svg size 900,300 font 'sans,13' set output 'figures/by_class_aboard.svg' set style data histograms set yrange [0:*] plot data using 2:xtic(1) title "Aboard" unset output
After saving, run listing 29 from Bash in terminal/.
gnuplot -d plots/by_class_aboard.gp
by_class_aboard.svg)Practice this topic in Challenge 3: Ask the first plot a different question.
Open figures/by_class_aboard.svg in a browser. set output names that file.
Use a plain-text editor, save with the exact .gp extension—not =.gp.txt=—
then rerun and reload the browser after each change.
Bash receives the gnuplot -d ... command; the editor contains gnuplot
statements. The -d flag skips personal startup files. reset clears
settings left by earlier plots. Relative paths are resolved from the
working directory, not from the script: stay in terminal/.
You can also work at the gnuplot prompt. Run listing 30 now,
then use load after saving a script. Bash's $ and gnuplot's gnuplot>
prompts are not characters to copy into commands.
$ gnuplot -d
For a text preview without a separate window, enter listing
31 at that prompt. load first runs the script
and saves its SVG; switching to dumb afterwards and calling replot
also draws the chart in the terminal. The saved SVG remains unchanged.
Text output has limited resolution and cannot faithfully reproduce every
later chart, especially multiplots and colour heatmaps. Use their saved
SVG or PNG files for the full result.
gnuplot> load 'plots/by_class_aboard.gp' gnuplot> unset output gnuplot> set terminal dumb size 79,24 gnuplot> replot
3 — One series, minimal code
New here: categorical positions and boxes, rather than the numeric x
coordinates and points in the introduction. using 3:xtic(1) takes
values from column 3 and category labels from column 1. This version shows
only survivors. notitle suppresses an unhelpful data-source legend.
The zero baseline is explicit; descriptive wording is the next refinement.
Category labels versus numerical coordinates. xtic(1) supplies labels,
not numerical x positions: the classes are spaced equally by row. For age
data, use the actual age as x, for example using 1:2 when column 1 holds
age and column 2 holds the measured value. Ages 1, 4 and 10 must have gaps
of 3 and 6 years; plotting against row number would make both gaps equal.
In your editor, create plots/by_class_survivors.gp inside terminal/.
Enter listing 32, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory. reset set encoding utf8 set datafile separator "," set datafile columnheaders data = 'data/titanic_by_class.csv' set terminal svg size 900,400 font 'sans,13' set output 'figures/by_class_survivors.svg' set boxwidth .5 relative set yrange [0:*] plot data using 3:xtic(1) with boxes notitle unset output
Run listing 33 from Bash in terminal/.
gnuplot -d plots/by_class_survivors.gp
by_class_survivors.svg)Practice this topic in Challenge 4: Give bars an honest baseline.
4 — Two outcomes, readable labels
New here: histograms groups bars by category. Both outcomes share axes
and a legend. The comma and series title work just as they did for the
two synthetic series in the introduction. We repeat data for
clarity; later '' means reuse the preceding data source. Colours are
still gnuplot's defaults.
In your editor, create plots/by_class_counts.gp inside terminal/.
Enter listing 34, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcomes.svg'
set style data histograms
set style fill solid 1.0
set title "Titanic: survivors and deaths by class and crew"
set ylabel "People aboard"
set yrange [0:*]
set key top left
plot data using 3:xtic(1) title "Survived", \
data using 4 title "Died"
unset output
Run listing 35 from Bash in terminal/.
gnuplot -d plots/by_class_counts.gp
by_class_outcomes.svg)Practice this topic in Challenge 5: Check two series against their total.
5 — Make the appearance deliberate
Every visual choice is explicit: left and bottom axes only (set border 3),
outward ticks, light horizontal guides and two colours chosen on purpose.
Use a strong colour for survivors and a quiet one for deaths. In a report,
the caption can carry the title rather than repeating it inside the figure.
In your editor, create plots/by_class_styled.gp inside terminal/.
Enter listing 36, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcomes_styled.svg'
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
set border 3 lw 1.9
set tics out nomirror scale 0.9
set grid ytics lc rgb "#d5d8dc" lw .5
set key top left
set linetype 1 lc rgb "#14507d"
set linetype 2 lc rgb "#aab0b8"
set style fill solid 0.9 border -2
set style data histograms
set style histogram clustered gap 1
set boxwidth 0.9
set ylabel "People aboard"
set xlabel "Passenger class or crew"
set yrange [0:*]
plot data using 3:xtic(1) title "Survived", \
data using 4 title "Died"
unset output
Run listing 37 from Bash in terminal/.
gnuplot -d plots/by_class_styled.gp
by_class_outcomes_styled.svg)Most of the setup describes a visual style, not this particular question. Repeating those settings in every plot invites inconsistent edits.
Practice this topic in Challenge 6: Make one visual change at a time.
Define the shared style
A shared style contains settings, not a plot command. It travels with
the project and can be inspected by another reader. A personal ~/.gnuplot
would apply invisibly on one machine without necessarily travelling with
the lesson.
Before using styled plots, create style.gp directly inside terminal/,
not inside plots/. Save listing 38 there. Each later script
uses load 'style.gp' to read those settings.
# Shared appearance for the terminal lesson.
reset
if (GPVAL_VERSION < 6.0) { exit error "Need gnuplot 6.0 or newer" }
set encoding utf8
set datafile separator ","
set datafile columnheaders
set border 3 lw 1.9
set tics out nomirror scale 0.9
set grid ytics lc rgb "#d5d8dc" lw .5
set key top left
set linetype 1 lc rgb "#14507d"
set linetype 2 lc rgb "#aab0b8"
set style fill solid 0.9 border -2
set style data histograms
set style histogram clustered gap 1
set boxwidth 0.9
7 — Stacked counts including crew
rowstacked puts survivors and deaths on top of each other. The total
height of each bar is the size of the group, so this chart answers both
“how large was the group?” and “how did its outcomes divide?”
In your editor, create plots/by_class_stacked.gp inside terminal/.
Enter listing 41, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_outcomes_stacked.svg'
load 'style.gp'
set style histogram rowstacked
set ylabel "People aboard"
set yrange [0:*]
plot data using 3:xtic(1) title "Survived", \
'' using 4 title "Died"
unset output
Run listing 42 from Bash in terminal/.
gnuplot -d plots/by_class_stacked.gp
by_class_outcomes_stacked.svg)The bar totals are 324, 284 and 709 passengers, followed by 890 crew. Crew is the largest group, with 679 deaths; third class has 528 deaths. Both groups are larger than either first or second class.
Practice this topic in Challenge 8: Choose between stacked and grouped counts.
9 — Rates with exact labels
New here: one rate per group, followed by an optional text layer. The denominator is still the aboard count of that same class or Crew group, not the total of 2,207. First draw the bars without formatting strings.
First pass — only the rates
Use listing 45 after the input setup from the preceding
class-count example: data must refer to titanic_by_class. In Org, copy
that example's complete block (including :var data=titanic_by_class),
give it a new name and file path, and replace its body with this listing.
In a terminal script, retain the CSV separator, column-header, data,
terminal and output setup, but replace the plotting body. This is a
replacement body, not an independently runnable source block.
set ylabel "Survived (%)" set yrange [0:100] set boxwidth 0.6 plot data using 0:($3/$2*100):xtic(1) with boxes notitle
0 is the row index (0, 1, 2, 3), used for the bar positions. The
calculated y value is the same percentage introduced in the preceding
chapter. Check the first-class bar: 201/324*100 is about 62.0%.
Report extension — exact labels and shared appearance
The complete version below retains the report styling. Its only new plot
layer is with labels: x and y place the text, and sprintf supplies it.
The function survival_percent names the repeated calculation; it does
not create or regroup data. %.1f formats one decimal place, and %%
prints a literal percent sign. Compare bar heights before and after adding
the label layer: none should move. Challenge 10 practices the precision.
In your editor, create plots/by_class_survival_rate.gp inside terminal/.
Enter listing 46, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_survival_rate.svg'
load 'style.gp'
survival_percent(survived, aboard) = 100.0 * survived / aboard
set ylabel "Survived"
set yrange [0:100]
set format y "%.0f%%"
set boxwidth 0.6
set offsets 0.5, 0.5, 0, 0
# Both layers use the same x position and the same survival percentage.
plot data using 0:(survival_percent($3,$2)):xtic(1) with boxes notitle, \
data using 0:(survival_percent($3,$2)):(sprintf("%.1f%%", survival_percent($3,$2))) \
with labels offset 0,0.5 notitle
unset output
Run listing 47 from Bash in terminal/.
gnuplot -d plots/by_class_survival_rate.gp
by_class_survival_rate.svg)The percentages are 62.0 for first class, 41.5 for second, 25.5 for third and 23.7 for Crew. Rates compare survival within groups; they do not show how many people were aboard.
| Class | Survived % |
|---|---|
| 1st | 62.0 |
| 2nd | 41.5 |
| 3rd | 25.5 |
| Crew | 23.7 |
Practice this topic in Challenge 10: Decide how much precision to show.
10 — Class and sex
Compare women's and men's survival rates within each group. Women use columns 3/2; men use 6/5. Each denominator must be that group's own aboard count, not the total number of survivors or the total aboard.
The colours now encode sex rather than survival outcome. A linetype number only selects a style: change its colour and the legend together when the plot's question changes. These categories include children.
In your editor, create plots/by_class_and_sex.gp inside terminal/.
Enter listing 48, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_and_sex_survival_rates.svg'
load 'style.gp'
set linetype 1 lc rgb "#14507d" # women
set linetype 2 lc rgb "#c47a2c" # men
set ylabel "Survived"
set yrange [0:100]
set format y "%.0f%%"
set key top right
plot sex using ($3/$2*100):xtic(1) title "Women", \
sex using ($6/$5*100) title "Men"
unset output
Run listing 49 from Bash in terminal/.
gnuplot -d plots/by_class_and_sex.gp
by_class_and_sex_survival_rates.svg)Practice this topic in Challenge 11: Compare within a group.
11 — A pie with circles
A pie shows one whole divided into parts. Here the whole is everyone aboard, including crew. Use it for one composition; bars are clearer for comparing survival rates between groups.
For a simple pie, with circles is sufficient in gnuplot 6: supply the
centre, radius, start angle and end angle for each slice, plus its colour.
The default set style circle wedge connects the arc to the centre.
No polar mode is needed. The donut chapter introduces with sectors for
actual ring segments; using gnuplot 6 does not mean replacing every older style.
New here: stats data using 2 nooutput reads the aboard column and stores
its sum as STATS_sum. Save that as total (2,207). The named function
angle_for_count converts a count into its share of 360 degrees.
For example, first class occupies 324/2207*360, about 52.85 degrees.
The running variable angle_end carries one slice's end to the next
slice's start. Start with slices only: labels are an extension in Challenge 12.
In your editor, create plots/by_class_pie.gp inside terminal/.
Enter listing 50, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_aboard_pie.svg'
# reset leaves Cartesian coordinates and unrestricted ranges for stats.
stats data using 2 nooutput
total = STATS_sum
angle_for_count(people) = people * 360.0 / total
set title "Titanic: people aboard by class and crew"
set angles degrees
unset border
unset tics
unset grid
unset key
set size ratio -1
set xrange [-1.05:1.05]
set yrange [-1.05:1.05]
set style fill solid 1.0 border lc rgb "white"
# circles: x : y : radius : start angle : end angle : color
angle_end = 0
plot data using (0):(0):(1):(angle_end): \
(angle_end = angle_end + angle_for_count($2)):($0 + 1) \
with circles linecolor variable notitle
unset output
Run listing 51 from Bash in terminal/.
gnuplot -d plots/by_class_pie.gp
by_class_aboard_pie.svg)
This example starts with reset, so stats runs before restrictive
ranges or polar mode are enabled. For a pasted fragment in an existing
session, run unset polar and set autoscale before stats.
Read the last line of the calculation from left to right: the fourth
circle column reads the old angle_end; the fifth adds this row's angular
width and returns the new end. This single update is necessary to join
successive slices. Before repeating the plot, reset angle_end to zero;
plain replot alone would reuse the accumulated value.
The sixth column selects a distinct default colour via linecolor variable.
($0 + 1) turns row indices 0–3 into colour indices 1–4.
The explicit ranges leave room around the unit circle, avoid the empty-y-range
warning and prevent clipping artefacts at the circle's boundary. The equal
axis scale (set size ratio -1) keeps the pie circular, not elliptical.
Optional report palette
The default colours distinguish the construction's four slices. To recover the muted report palette, insert listing 52 before the plot and rerun the complete example. This changes appearance, not shares. The label challenge uses this darker palette so white text remains readable.
set linetype 1 linecolor rgb "#14507d" set linetype 2 linecolor rgb "#4f86b5" set linetype 3 linecolor rgb "#6f9bc4" set linetype 4 linecolor rgb "#476178"
Angles start at three o'clock and run anticlockwise in table order: first class, second class, third class, then Crew. This deliberately unlabelled plot is a construction step, not a finished communication graphic. Without the table order, a reader cannot identify the slices from colour alone.
Complete it in Challenge 12: Give the slices names and percentages.
12 — Heatmaps and age panels
Class and sex
Each cell is a rectangle coloured by its survival percentage. With only
eight cells a table would also work; the heatmap becomes more useful as
the grid grows. Here with boxxyerror draws the cells.
The five using values are x, y, horizontal half-width, vertical half-height
and colour value. using (0):0 places women at x=0 and each data row on its
own y coordinate. Pseudo-column 0 is the row index.
In your editor, create plots/by_class_and_sex_heatmap.gp inside terminal/.
Enter listing 53, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 900,400 font 'sans,13'
set output 'figures/by_class_and_sex_survival_heatmap.svg'
load 'style.gp'
unset grid; unset key
set xrange [-0.5:1.5]
set yrange [3.5:-0.5]
set xtics ("Women" 0, "Men" 1) scale 0
set ytics scale 0
set palette defined (0 "#ffffff", 100 "#4f86b5")
set cbrange [0:100]
set format cb "%.0f%%"
set cblabel "Survived"
set style fill solid 1.0 border lc rgb "white"
plot sex using (0):0:(0.5):(0.5):($3/$2*100):ytic(1) with boxxyerror lc palette, \
'' using (1):0:(0.5):(0.5):($6/$5*100) with boxxyerror lc palette, \
'' using (0):0:(sprintf("%.0f%%", $3/$2*100)) with labels, \
'' using (1):0:(sprintf("%.0f%%", $6/$5*100)) with labels
unset output
Run listing 54 from Bash in terminal/.
gnuplot -d plots/by_class_and_sex_heatmap.gp
by_class_and_sex_survival_heatmap.svg)Reversing the y range puts first class at the top. The palette represents 0–100%; its darkest colour remains light enough for the dark labels.
Practice this topic in Challenge 13: Read the same rates in two encodings.
Age and class — the large heatmap
Now show 25 three-year age bands across all four groups, including Crew.
Use titanic_by_age_and_class.csv, whose columns are age_midpoint,
class_code, survival_pct, n and survivors. These 100 cells describe
2,205 people with known age; the two unknown ages are excluded.
Percentages and group counts are already calculated.
Use the downloaded data/titanic_by_age_and_class.csv.
The script's CSV separator and column-header settings handle the file
format; do not additionally skip the first observation.
In your editor, create plots/by_age_and_class_heatmap.gp inside terminal/.
Enter listing 55, then save the file before running it.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_age_and_class.csv'
set terminal pngcairo size 2400,650 enhanced font "Sans,12"
set output 'figures/by_age_and_class_survival_heatmap.png'
set datafile missing "NA"
set title "Titanic survival by age and class\nPercent survived; n = group count"
set xlabel "Age band (completed years)"
set xrange [-0.5:74.5]
set yrange [4.5:0.5]
set xtics ("0-2" 1, "3-5" 4, "6-8" 7, "9-11" 10, "12-14" 13, \
"15-17" 16, "18-20" 19, "21-23" 22, "24-26" 25, "27-29" 28, \
"30-32" 31, "33-35" 34, "36-38" 37, "39-41" 40, "42-44" 43, \
"45-47" 46, "48-50" 49, "51-53" 52, "54-56" 55, "57-59" 58, \
"60-62" 61, "63-65" 64, "66-68" 67, "69-71" 70, "72-74" 73)
set ytics ("1st" 1, "2nd" 2, "3rd" 3, "Crew" 4)
set tics out nomirror
set key off
set palette defined (0 "#440154", 25 "#3b528b", 50 "#21918c", \
75 "#5ec962", 100 "#fde725")
set cbrange [0:100]
set cblabel "Survival (%)"
set style fill solid 1.0 border lc rgb "white"
# Column headers are already enabled; each cell spans three years and one class.
plot data using 1:2:($1-1.5):($1+1.5):($2-0.5):($2+0.5):3 \
with boxxyerror linecolor palette, \
data using 1:2:(sprintf("%.1f%%\nn=%d", $3, $4)):($3 >= 55 ? 0x111111 : 0xffffff) \
with labels textcolor rgb variable font "Sans,12"
unset output
Run listing 56 from Bash in terminal/.
gnuplot -d plots/by_age_and_class_heatmap.gp
by_age_and_class_survival_heatmap.png)
This form of boxxyerror uses seven values: x/y centre, left/right bounds,
lower/upper bounds and colour value. Each cell spans three years and one
class category.
The second layer adds the percentage and group count, switching text colour
for contrast. NA leaves the 12 empty groups blank, unlike an observed 0%.
Small n matters: 100% with n=1 means one survivor out of one person.
Open the PNG at full size to inspect all 25 columns.
Practice this topic in Challenge 14: Show when a percentage has little support.
Age trends in four panels
Use the same age/class summary for a four-panel plot, including Crew. Identical axes make the age patterns easier to compare. Points represent three-year groups, not individual ages, and the labels report group counts.
In your editor, create plots/by_age_and_class_panels.gp inside terminal/.
Enter listing 57, then save the file before running it.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
header_rows = 0 # columnheaders already removes the CSV header
data = "data/titanic_by_age_and_class.csv"
set terminal pngcairo size 2200,1100 enhanced font "Sans,11"
set output "figures/by_age_and_class_survival_panels.png"
set datafile missing "NA"
set xrange [-0.5:74.5]
set yrange [-8:115]
set xtics 1,3,73
set ytics 0,20,100
set xlabel "Age band centre (whole years)"
set ylabel "Survival (%)"
set grid ytics
set tics out nomirror
set key off
set style line 1 lc rgb "#0072B2" lw 2 pt 7
set style line 2 lc rgb "#D55E00" lw 2 pt 5
set style line 3 lc rgb "#009E73" lw 2 pt 9
set style line 4 lc rgb "#CC79A7" lw 2 pt 13
classes = "1st 2nd 3rd Crew"
set multiplot layout 2,2 rowsfirst title "Titanic survival by class\nLabels show group counts (n)" font "Sans,16"
do for [c=1:4] {
set title word(classes,c)
plot data every 4::(header_rows+c-1) using 1:3 with linespoints linestyle c pointsize 1.1, \
data every 4::(header_rows+c-1) using 1:3:(sprintf("n=%d", $4)) \
with labels offset char 0,1 textcolor rgb "#333333" font "Sans,10"
}
unset multiplot
unset output
Run listing 58 from Bash in terminal/.
gnuplot -d plots/by_age_and_class_panels.gp
by_age_and_class_survival_panels.png)
set multiplot layout 2,2 arranges four panels. The loop selects class
codes 1–4, and every 4::(header_rows+c-1) takes every fourth record for
class c. Keep the published ordering: each age band contains all four
class codes, even where a group is empty.
word(classes,c) chooses the title and linestyle c its colour and marker.
unset multiplot finishes the layout. The lines guide the eye; they are
not a fitted model. Missing groups are not zero-percent survival, and
extreme percentages from small groups deserve caution.
set datafile columnheaders removes the CSV header before row selection,
so header_rows is 0. Using 1 as well would shift the selected class.
Practice this topic in Challenge 15: Do lines imply more than the data show?.
13 — A dumbbell chart
A dumbbell connects the men's survival rate to the women's rate for each class or crew group. The line's length is their gap.
New here: horizontal displacement between two already familiar rates.
As in the rate example, survival_percent names the repeated arithmetic.
Column 3 divided by 2 is the women's rate; column 6 divided by 5 is the
men's rate. Subtracting them gives percentage points, not percent change.
The colours and larger points below are the report layer, not new data.
with vectors takes x:y:dx:dy: a starting point and a displacement.
Start at the men's rate, use the difference between the rates as dx, and
keep dy at zero. nohead removes the arrowhead to leave a connecting line.
In your editor, create plots/by_class_sex_gap.gp inside terminal/.
Enter listing 59, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 820,400 font 'sans,14'
set output 'figures/by_class_and_sex_survival_gap.svg'
load 'style.gp'
survival_percent(survived, aboard) = 100.0 * survived / aboard
unset grid
set grid xtics lc rgb "#d5d8dc" lw 0.8
set xrange [0:100]
set yrange [3.6:-0.6]
set format x "%.0f%%"
set xlabel "Survived"
set ytics scale 0
set key above
set linetype 1 lc rgb "#14507d" # women
set linetype 2 lc rgb "#c47a2c" # men
plot sex using (survival_percent($6,$5)):0:(survival_percent($3,$2) - survival_percent($6,$5)):(0):ytic(1) \
with vectors nohead lw 4 lc rgb "#d5d8dc" notitle, \
'' using (survival_percent($3,$2)):0 with points pt 7 ps 2 lc 1 title "Women", \
'' using (survival_percent($6,$5)):0 with points pt 7 ps 2 lc 2 title "Men"
unset output
Run listing 60 from Bash in terminal/.
gnuplot -d plots/by_class_sex_gap.gp
by_class_and_sex_survival_gap.svg)Practice this topic in Challenge 16: Add an honest name for the gap.
14 — Five measures on a spider plot
A spider or radar plot shows several measures of one thing at once: one axis per measure, one polygon per row. Each class or crew group is a polygon across five percentages derived from the class-and-sex counts. This course uses gnuplot 6.0 or newer, including for spider plots.
Each plot clause adds an axis. All axes run from 0 to 100, and
set paxis ... label supplies each spoke label and its rotation.
In your editor, create plots/by_class_profile.gp inside terminal/.
Enter listing 61, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
sex = 'data/titanic_by_sex.csv'
set terminal svg size 820,640 font 'sans,14'
set output 'figures/by_class_profile_spider.svg'
load 'style.gp'
set title "Titanic survival by class and sex" offset 0,1 font ",18"
unset border; unset tics; unset grid
set spiderplot
# One axis per plot clause below: five measures, all in per cent
set for [p=1:5] paxis p range [0:100]
set paxis 1 tics 0,25,100 format "%.0f%%"
set for [p=2:5] paxis p tics format ""
# Spoke labels, each at right angles to its spoke. Spoke p points at
# 90 - (p-1)*72 degrees; the label is turned 90 degrees less than that,
# and turned over by 180 where it would otherwise stand upside down.
set paxis 1 label "All survived" rotate by 0
set paxis 2 label "Women survived" rotate by -72
set paxis 3 label "Men survived" rotate by 36
set paxis 4 label "Women aboard" rotate by -36
set paxis 5 label "Men aboard" rotate by 72
set grid spiderplot lc rgb "#d0d0d0" dt 2 lw 1
set style spiderplot fs transparent solid 0.18 border lw 2
set key outside right center
# One polygon per row of the table, i.e. per class
set linetype 1 lc rgb "#14507d" lw 3
set linetype 2 lc rgb "#c47a2c" lw 3
set linetype 3 lc rgb "#3f8f5a" lw 3
set linetype 4 lc rgb "#8064a2" lw 3
# Columns: 2 women aboard, 3 women survived, 5 men aboard, 6 men survived
# key(1) names each polygon after column 1 (1st, 2nd, 3rd, Crew)
plot \
sex using (($3+$6)/($2+$5)*100):key(1) with spiderplot notitle, \
sex using ($3/$2*100) with spiderplot notitle, \
sex using ($6/$5*100) with spiderplot notitle, \
sex using ($2/($2+$5)*100) with spiderplot notitle, \
sex using ($5/($2+$5)*100) with spiderplot notitle
unset output
Run listing 62 from Bash in terminal/.
gnuplot -d plots/by_class_profile.gp
by_class_profile_spider.svg)The last two axes are group composition—the shares of women and men aboard—not survival rates. They sum to 100 and therefore mirror each other. Crew is overwhelmingly male and has the lowest overall survival rate, 23.7%. This describes the group; it does not establish why people survived.
Practice this topic in Challenge 17: Test whether polygon shape is evidence.
15 — Class size and fate in a donut
The inner ring divides everyone aboard into three passenger classes and crew. The outer ring splits each group into survivors and deaths. It combines the two messages of the stacked counts: group size and outcome.
This is an optional report example, not the next minimal plotting command. New here are ring boundaries and layered angular intervals; the calculations still use the same counts and whole as the pie. Read the inner-ring pass first, then the two outcome passes, and only then the label passes. The stacked chart remains a simpler way to communicate the same quantities.
Draw real ring segments with with sectors: class segments run from
hole_radius to split_radius; outcome segments run from split_radius
to outer_radius. The centre stays empty without a white covering disc.
Each pass resets pos so both rings share the same group boundaries.
Here the sectors columns are start angle:inner radius:angular width:radial width,
optionally followed by a colour. Where circles takes start and end
angles, sectors takes a start angle and angular width. In polar mode, labels take
angle:radius:text directly. See the gnuplot 6 manual, Sectors.
In your editor, create plots/by_class_donut.gp inside terminal/.
Enter listing 63, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory.
reset
set encoding utf8
set datafile separator ","
set datafile columnheaders
data = 'data/titanic_by_class.csv'
set terminal svg size 600,600 font 'sans,11'
set output 'figures/by_class_and_outcome_donut.svg'
load 'style.gp'
unset polar
set autoscale
stats data using 2 nooutput
total = STATS_sum
set angles degrees
ang(n) = n * 360.0 / total
set polar
set theta right ccw
unset raxis
unset border; unset tics; unset grid
set key below horizontal center
set size ratio -1
set xrange [-1.35:1.35]
set yrange [-1.35:1.35]
set rrange [0:1.35]
set style fill solid 1.0 border lc rgb "white"
set linetype 3 lc rgb "#2b3036"
set linetype 4 lc rgb "#4a5058"
set linetype 5 lc rgb "#6c737c"
set linetype 6 lc rgb "#515969"
hole_radius = 0.40
split_radius = 0.72
outer_radius = 1.0
label_radius = (hole_radius + split_radius)/2
rate_radius = 1.16
# sectors: start angle : inner radius : angular width : radial width
# Each outcome pass advances pos by the WHOLE group, not just that outcome.
plot pos = 0, \
data using (pos):(split_radius): \
(span = ang($3), pos = pos + ang($2), span): \
(outer_radius - split_radius) \
with sectors lc rgb "#14507d" title "Survived", \
pos = 0, \
data using (pos + ang($3)):(split_radius): \
(span = ang($4), pos = pos + ang($2), span): \
(outer_radius - split_radius) \
with sectors lc rgb "#dfe3e8" title "Died", \
pos = 0, \
data using (pos):(hole_radius): \
(span = ang($2), pos = pos + span, span): \
(split_radius - hole_radius):($0 + 3) \
with sectors lc variable notitle, \
pos = 0, \
data using (mid = pos + ang($2)/2, pos = pos + ang($2), mid): \
(label_radius):(strcol(1)) \
with labels tc rgb "white" center notitle, \
pos = 0, \
data using (mid = pos + ang($2)/2, pos = pos + ang($2), mid): \
(rate_radius):(sprintf("%.0f%%", 100.0 * $3 / $2)) \
with labels tc rgb "#2b3036" center notitle
unset output
Run listing 64 from Bash in terminal/.
gnuplot -d plots/by_class_donut.gp
by_class_and_outcome_donut.svg)
The first pass draws survivors in blue; the second starts after those
survivors and fills the remaining group angle in light grey. Both passes
advance pos by the whole group's angle, but return only the outcome's
angular width. These widths use counts divided by the total aboard, so
survivors and deaths together align with the corresponding inner segment.
set polar lets the labels use angle:radius:text directly.
The plot uses square output dimensions rather than the wide default. Class names sit inside the inner ring and survival rates outside. Crew is the largest inner wedge, 40.3%, with a mostly grey outer rim. First class is 14.7% of those aboard and has a mostly blue rim.
Change hole_radius to adjust the hole. Keep
0 < hole_radius < split_radius < outer_radius=; the radial widths are
differences between these boundaries, not absolute outer radii.
label_radius keeps class names halfway across the inner ring.
Practice this topic in Challenge 18: Separate decoration from information.
16 — Tiny age-count bars
A sparkline is a chart about the size of a word: no axes, labels or legend, just the data's shape inside a sentence. It answers “what does the trend look like?” without interrupting the reading.
Use titanic_by_age.csv with its 25 three-year groups from 0–2 to 72–74.
Both sexes, all classes and crew are included; two unknown ages are excluded.
These bars show how many people were in each age group, not how many survived.
In your editor, create plots/by_age_counts.gp inside terminal/.
Enter listing 65, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory. reset set encoding utf8 set datafile separator "," set datafile columnheaders data = 'data/titanic_by_age.csv' set terminal svg size 120,24 set output 'figures/by_age_people_sparkline.svg' load 'style.gp' unset border; unset tics; unset key; unset grid set margins 0, 0, 0, 0 set yrange [0:*] set offsets 0.5, 0.5, 0, 0 set boxwidth 0.7 set style fill solid 1.0 noborder plot data using 0:2 with boxes lc 1 unset output
Run listing 66 from Bash in terminal/.
gnuplot -d plots/by_age_counts.gp
by_age_people_sparkline.svg)
The SVG is only 120 by 24 pixels. Removing axes, tics, legend and grid,
and setting all margins to zero, leaves room for the data.
set offsets preserves a little horizontal space so the end bars are not
cut off. Pseudo-column 0 places the rows at positions 0, 1, 2, … because
the age labels in the first column are text.
Practice this topic in Challenge 19: Make a tiny chart understandable.
17 — An inline survival trend
The same age groups now show the survival percentage as a line—not the
number of survivors. Keep set yrange [0:100] so the vertical scale remains
comparable between survival sparklines. Without it, small differences could
be stretched to the full height.
In your editor, create plots/by_age_survival_rate.gp inside terminal/.
Enter listing 67, then save the file before running it.
# Generated from plots-with-roots.dtx; run from the terminal directory. reset set encoding utf8 set datafile separator "," set datafile columnheaders data = 'data/titanic_by_age.csv' set terminal svg size 120,24 set output 'figures/by_age_survival_sparkline.svg' load 'style.gp' unset border; unset tics; unset key; unset grid set margins 0, 0, 0, 0 set yrange [0:100] set offsets 0.3, 0.3, 0, 0 plot data using 0:3 with lines lw 1.5 lc 1 unset output
Run listing 68 from Bash in terminal/.
gnuplot -d plots/by_age_survival_rate.gp
by_age_survival_sparkline.svg)The supplied browser gallery displays the tiny SVGs at sparkline size. When placing them in another document, use a height close to its text line rather than stretching them to a full-width chart.
The highest published rate is 67.9% for ages 3–5, compared with 61.8% for ages 0–2. The three oldest groups have no recorded survivors, but their counts are small. The trend is not monotonic.
Practice this topic in Challenge 20: See what an automatic scale hides.
18 — Run the complete set again
After working through every plot, rerun the complete sequence. Check that all eighteen figures are present, including the large age/class heatmap and four-panel age comparison. Revisit the data description and the interpretation whenever changing the dataset version.
A reproducible rerun should not depend on invisible settings from an earlier interactive session. Keep the published inputs, explicit plotting commands and shared style together.
After saving all eighteen plot files and style.gp, create all.gp
directly inside terminal/. Listing 69 contains one load
command per plot; it does not download data, calculate summaries or create
missing scripts.
# Run from terminal/: gnuplot -d all.gp load 'plots/by_class_aboard.gp' load 'plots/by_class_survivors.gp' load 'plots/by_class_counts.gp' load 'plots/by_class_styled.gp' load 'plots/by_class_shared.gp' load 'plots/by_class_stacked.gp' load 'plots/by_class_shares.gp' load 'plots/by_class_survival_rate.gp' load 'plots/by_class_and_sex.gp' load 'plots/by_class_pie.gp' load 'plots/by_class_and_sex_heatmap.gp' load 'plots/by_age_and_class_heatmap.gp' load 'plots/by_age_and_class_panels.gp' load 'plots/by_class_sex_gap.gp' load 'plots/by_class_profile.gp' load 'plots/by_class_donut.gp' load 'plots/by_age_counts.gp' load 'plots/by_age_survival_rate.gp'
Run listing 70 only after every script and all five CSVs exist.
gnuplot -d all.gp
Open the SVGs and PNGs in figures/, or reload the supplied index.html
gallery. If a script is missing, return to its topic and save it first.
For individual reruns from Bash, use listing 71.
gnuplot -d plots/by_class_survivors.gp gnuplot -d plots/by_class_counts.gp gnuplot -d plots/by_class_styled.gp gnuplot -d plots/by_class_survival_rate.gp
At the gnuplot prompt, use listing 72 instead.
load reads the saved file again; replot repeats the most recent plot
command with the current settings. Save edits before loading a file.
gnuplot> load 'plots/by_class_survival_rate.gp' gnuplot> # Edit the file in your editor, then load it again: gnuplot> load 'plots/by_class_survival_rate.gp' gnuplot> exit
Practice this topic in Challenge 21: Prove that an edit reaches the output.
19 — What the figures say
Of the 2,207 people in the published class table, 711 survived—about one in three. Passenger survival rates were 62.0% in first class, 41.5% in second and 25.5% in third; Crew's rate was 23.7%.
Crew also has the largest death count, 679, followed by third class with 528. Together these groups account for most deaths. Compare the stacked counts with the rate bars: a large number of deaths and a low survival rate are related but different descriptions.
Within every class and crew group, women survived at a higher rate than men. Among passengers, women's rates were approximately 97%, 89% and 49% in first, second and third class. Men's rates were approximately 34%, 13% and 15%. Among crew, 20 of 23 women survived (87%), compared with 191 of 867 men (22%). Read the small female crew count alongside its rate.
The age heatmap, four panels and sparkline show age patterns at different levels of detail. Only the two unknown ages are excluded from age-based plots; the class and sex plots include all records. Empty groups are not zero survival.
These numbers describe who survived, not why. The selected columns do not record deck location, access to lifeboats, evacuation timing or individual decisions. Do not infer causal effects of age, class or sex from these plots. Consult the dataset provenance before treating the supplied records as a definitive historical reconstruction.
The reported values belong to the pinned release. A new dataset version requires reviewing this interpretation as well as rerunning the code.
Practice this topic in Challenge 22: Write a claim the figure supports.
20 — Troubleshooting
Work through the same questions when a plot surprises you: is the input available, is it read correctly, are the selected columns appropriate, and have you rerun and reopened the right output?
| Symptom | Check |
|---|---|
command not found |
Install the required program or correct PATH, then restart the shell or Emacs. |
| Only one series appears | Check the comma separating plot elements and the selected columns. |
| Unexpected percentage | Check the numerator and that group's own denominator. |
| Missing first observation | Do not skip a row twice after handling the header. |
| Missing output figure | Check the output path and directory, and that the plot ran successfully. |
| Old figure remains visible | Rerun the code, then refresh the image display or browser. |
| Unexpected colours or axes | Start from reset and the explicit shared settings; check plot-specific overrides. |
| Spider syntax fails | Check that gnuplot is version 6.0 or newer. |
| An empty heatmap cell | Inspect n and NA before interpreting it as zero survival. |
Stay in terminal/; confirm that all five files are under data/ and
the figures/ directory exists. Use comma separation and column headers
after reset. The scripts write SVG or PNG rather than opening a plot
window. Keep Bash commands and gnuplot statements in their respective
contexts.
If a browser still shows an older result, reload after running the correct script. The only network step in the lesson is downloading the CSVs.
Practice this topic in Challenge 23: Catch a plausible but wrong plot.
Further reading
- Philipp K. Janert, Gnuplot in Action, Second Edition (Manning, 2016). Optional book on plotting and understanding data.
- Org Babel gnuplot tutorial. Worked examples of plotting functions and named tables in Org; some setup and output advice is historical.
- Plotting tables with org-plot. Examples and options for
#+PLOT:directives above Org tables.