Titanic Passengers and Crew
Demographics and Survival Data
Purpose and outputs
This document contains 2,207 person-level records, their provenance and all
calculations needed to produce the CSV datasets below. Tangle the embedded
base CSV, then evaluate the AWK blocks. Org saves their CSV output using
:results output file and :file. No external data download or generated
export script is needed.
This document is extracted from dataset-titanic.dtx. Edit the embedded
records or calculations in that DTX source, not the generated files.
For an independent offline rebuild, run make datasets.
Requirements: Emacs with AWK and Bash support enabled in Org Babel, Bash, AWK,
head, cat, tee and either sha256sum or shasum. All paths are relative
to the directory containing this document. Tangling creates data/ if needed.
Output in data/ |
Rows | Contents |
|---|---|---|
| titanic.csv | 2,207 | Person-level class, sex, age, survival and embarkation |
| titanic_by_class.csv | 4 | Outcome counts by passenger class and crew membership |
| titanic_by_sex.csv | 4 | Outcome counts by class/crew and sex |
| titanic_by_age_and_class.csv | 100 | Counts and survival percentages by three-year age band and class/crew |
| titanic_by_age.csv | 25 | Counts and survival percentages by three-year age band |
| SHA256SUMS | 5 | Checksums of the five CSV files |
All datasets are comma-separated .csv files. They use UTF-8,
one header row, LF line endings and decimal points. Row counts exclude headers.
NA marks unknown ages or empty age/class cells; the age-only table uses
NaN for an empty group's percentage. Missing values are never zero-filled.
Provenance and limitations
Dataset collection: doi:10.5281/zenodo.22967792. This is the all-versions (concept) DOI for Titanic Passengers and Crew: Demographics and Survival Data. Use it to find the dataset collection; use a version-specific DOI when citing the exact files used in an analysis.
The direct source was
Dennis Sun's Principles of Data Science Titanic CSV,
accessed 25 September 2026. Selected fields were manually transcribed from
the browser-readable CSV, then exported and sorted. Four crew categories
were merged into Crew; gender became sex; names and identifiers were
omitted. Identical five-field rows remain because they may represent
different people. The historical extraction is not reproduced by this file;
only transformations from the embedded snapshot onward are reproducible.
Treat this as a teaching snapshot, not a definitive historical reconstruction. Ages and transcriptions may contain errors. Crew is not a fourth passenger class. Pooling sexes or crew roles can hide differences, and small cells need their counts shown alongside percentages. These data do not establish causes.
Person-level records
| Column | Meaning | Values |
|---|---|---|
| class | Passenger class or crew membership | 1st, 2nd, 3rd, Crew |
| sex | Sex as recorded in the source | female, male |
| age | Reported age in years at the sinking | Numeric, including fractional infant ages; NA for unknown age |
| survived | Recorded survival outcome | 1 = survived; 0 = did not survive |
| embarked | Embarkation location | B = Belfast; C = Cherbourg; Q = Queenstown; S = Southampton |
The complete records are embedded in the Org source and tangled unchanged as data/titanic.csv. The following section is omitted from HTML to avoid printing 2,207 lines; the downloadable CSV retains them all.
What the snapshot preserves
Two ages are NA; other fields have no missing values. Fractional infant
ages are retained. Records are ordered by class, sex, numeric age, outcome
and embarkation, with missing ages last within class/sex groups.
The build does not impute ages, deduplicate people or reconstruct identifiers.
Listing 1 shows ten lines without truncating the files
used for calculations. Set rows to change only this preview.
head -n "$rows" data/titanic.csv
class,sex,age,survived,embarked 1st,female,2,0,S 1st,female,13,1,S 1st,female,16,1,C 1st,female,16,1,S 1st,female,17,1,C 1st,female,17,1,S 1st,female,18,1,C 1st,female,18,1,C 1st,female,18,1,C
Generate the datasets
In Emacs, first use C-c C-v t to tangle the embedded data/titanic.csv.
Then evaluate the following blocks in order with C-c C-c. The AWK blocks
read their CSV input through :in-file. FS = OFS = "," sets comma-separated
input and output. :results output file with :file data/filename.csv tells
Org to save the printed CSV as that file. The preview blocks display it.
Only the previews and checksum commands use Bash; no export script is needed.
Use a decimal-point numeric locale. If necessary, set Emacs's process
environment with M-: (setenv "LC_ALL" nil) and
M-: (setenv "LC_NUMERIC" "C") before evaluating the AWK blocks.
Alternatively, after tangling, use C-c C-v b to evaluate all executable
blocks in document order, including previews and checks. No Python, Makefile
or network access is required.
Validate the snapshot
Listing 2 rejects malformed rows, unexpected
categories, ages outside the documented range and changed cohort totals.
Tangling preserves every record and its order. The five embedded fields
contain neither quoted text nor commas, so FS = "," is sufficient here;
it is not a general parser for arbitrary quoted CSV files.
BEGIN { FS = "," }
NR == 1 {
if ($0 != "class,sex,age,survived,embarked") bad = 1
next
}
{
if (NF != 5 || $1 !~ /^(1st|2nd|3rd|Crew)$/ ||
$2 !~ /^(female|male)$/ || $4 !~ /^[01]$/ ||
$5 !~ /^[BCQS]$/) bad = 1
if ($3 == "NA") missing++
else if ($3 !~ /^[0-9]+([.][0-9]+)?$/ || $3 < 0 || $3 >= 75) bad = 1
n++; survived += $4; crew += ($1 == "Crew")
}
END {
if (bad || n != 2207 || survived != 711 || crew != 890 || missing != 2) {
print "Invalid Titanic snapshot" > "/dev/stderr"; exit 1
}
}
Outcomes by class and crew
Listing 3 counts everyone, including the two unknown ages. For each group, deaths are aboard minus survivors. The explicit loop fixes the display order; associative-array iteration would not guarantee it.
BEGIN { FS = OFS = ","; print "Class", "Aboard", "Survived", "Died" }
NR > 1 { n[$1]++; s[$1] += $4 }
END {
split("1st 2nd 3rd Crew", label, " ")
for (c = 1; c <= 4; c++) {
k = label[c]; print k, n[k]+0, s[k]+0, n[k]-s[k]
}
}
Listing 4 displays the generated class-and-crew counts.
cat data/titanic_by_class.csv
Class,Aboard,Survived,Died 1st,324,201,123 2nd,284,118,166 3rd,709,181,528 Crew,890,211,679
Outcomes by class and sex
Listing 5 uses a two-part key. Women's counts occupy columns 2–4, men's columns 5–7. Adding the two sexes recovers the class totals.
BEGIN {
FS = OFS = ","
print "Class", "Women aboard", "Women survived", "Women died", \
"Men aboard", "Men survived", "Men died"
}
NR > 1 { n[$1, $2]++; s[$1, $2] += $4 }
END {
split("1st 2nd 3rd Crew", label, " ")
for (c = 1; c <= 4; c++) {
k = label[c]
print k, n[k,"female"]+0, s[k,"female"]+0, n[k,"female"]-s[k,"female"], \
n[k,"male"]+0, s[k,"male"]+0, n[k,"male"]-s[k,"male"]
}
}
Listing 6 displays the generated table.
cat data/titanic_by_sex.csv
Class,Women aboard,Women survived,Women died,Men aboard,Men survived,Men died 1st,144,139,5,180,62,118 2nd,106,94,12,178,24,154 3rd,216,106,110,493,75,418 Crew,23,20,3,867,191,676
Three-year age bands and class
Listing 7 groups unrounded ages into [0,3), [3,6),
through [72,75), pooling sexes. The labels 1, 4, …, 73 are the middle
completed-year ages, not exact continuous interval midpoints. Class codes
1–4 identify first, second, third and crew. Unknown ages are excluded.
All 25 × 4 combinations are emitted, ordered by age and then class. Empty
cells have n=0, survivors=0 and NA, not 0%. Percentages have six decimal
places. Each age band is followed by its four class/crew categories in order.
BEGIN {
FS = OFS = ","
code["1st"]=1; code["2nd"]=2; code["3rd"]=3; code["Crew"]=4
print "age_midpoint", "class_code", "survival_pct", "n", "survivors"
}
NR > 1 && $3 != "NA" {
band = int($3 / 3); c = code[$1]
n[band,c]++; s[band,c] += $4
}
END {
for (band = 0; band < 25; band++)
for (c = 1; c <= 4; c++)
print 3*band+1, c, \
n[band,c] ? sprintf("%.6f", 100*s[band,c]/n[band,c]) : "NA", \
n[band,c]+0, s[band,c]+0
}
data/titanic_by_age_and_class.csv
cat data/titanic_by_age_and_class.csv
age_midpoint,class_code,survival_pct,n,survivors 1,1,50.000000,2,1 1,2,100.000000,10,10 1,3,45.454545,22,10 1,4,NA,0,0 4,1,100.000000,1,1 4,2,100.000000,6,6 4,3,57.142857,21,12 4,4,NA,0,0 7,1,100.000000,1,1 7,2,100.000000,6,6 7,3,20.000000,15,3 7,4,NA,0,0 10,1,100.000000,1,1 10,2,NA,0,0 10,3,23.809524,21,5 10,4,NA,0,0 13,1,100.000000,2,2 13,2,80.000000,5,4 13,3,27.272727,11,3 13,4,NA,0,0 16,1,83.333333,6,5 16,2,57.142857,7,4 16,3,26.666667,45,12 16,4,5.555556,18,1 19,1,83.333333,12,10 19,2,33.333333,24,8 19,3,21.649485,97,21 19,4,19.696970,66,13 22,1,80.952381,21,17 22,2,31.034483,29,9 22,3,23.893805,113,27 22,4,29.292929,99,29 25,1,70.588235,17,12 25,2,35.294118,34,12 25,3,39.240506,79,31 25,4,20.192308,104,21 28,1,66.666667,21,14 28,2,46.428571,28,13 28,3,21.794872,78,17 28,4,30.275229,109,33 31,1,65.217391,23,15 31,2,36.363636,33,12 31,3,29.629630,54,16 31,4,29.032258,124,36 34,1,88.888889,18,16 34,2,41.176471,17,7 34,3,24.324324,37,9 34,4,19.047619,84,16 37,1,51.851852,27,14 37,2,31.578947,19,6 37,3,23.076923,26,6 37,4,15.730337,89,14 40,1,57.142857,28,16 40,2,46.153846,13,6 40,3,10.000000,30,3 40,4,28.571429,77,22 43,1,57.894737,19,11 43,2,20.000000,10,2 43,3,5.263158,19,1 43,4,17.948718,39,7 46,1,51.515152,33,17 46,2,10.000000,10,1 46,3,30.769231,13,4 46,4,30.303030,33,10 49,1,66.666667,27,18 49,2,60.000000,10,6 49,3,0.000000,9,0 49,4,14.285714,21,3 52,1,75.000000,12,9 52,2,20.000000,5,1 52,3,0.000000,1,0 52,4,36.363636,11,4 55,1,53.333333,15,8 55,2,40.000000,5,2 55,3,0.000000,4,0 55,4,0.000000,7,0 58,1,38.461538,13,5 58,2,40.000000,5,2 58,3,0.000000,2,0 58,4,16.666667,6,1 61,1,46.153846,13,6 61,2,25.000000,4,1 61,3,0.000000,2,0 61,4,50.000000,2,1 64,1,25.000000,8,2 64,2,0.000000,3,0 64,3,25.000000,4,1 64,4,0.000000,1,0 67,1,0.000000,1,0 67,2,NA,0,0 67,3,0.000000,2,0 67,4,NA,0,0 70,1,0.000000,3,0 70,2,NA,0,0 70,3,0.000000,1,0 70,4,NA,0,0 73,1,NA,0,0 73,2,0.000000,1,0 73,3,0.000000,1,0 73,4,NA,0,0
Three-year age bands for age trends
Listing 9 uses the same 25 three-year intervals as the age/class
summary: [0,3), [3,6), through [72,75). Labels 0-2, 3-5, …,
72-74 describe completed years; fractional ages are grouped without rounding.
There is no combined 60+ group. All classes, crew and both sexes are pooled,
and the same two unknown ages are excluded. Percentages have one decimal
place; empty groups use NaN to distinguish an undefined percentage from zero.
BEGIN { FS = OFS = ","; print "Age", "People", "Survived %" }
NR > 1 && $3 != "NA" {
g = int($3 / 3)
n[g]++; s[g] += $4
}
END {
for (g = 0; g < 25; g++)
print sprintf("%d-%d", 3*g, 3*g+2), n[g]+0, \
n[g] ? sprintf("%.1f", 100*s[g]/n[g]) : "NaN"
}
Listing 10 displays all 25 bins.
cat data/titanic_by_age.csv
Age,People,Survived % 0-2,34,61.8 3-5,28,67.9 6-8,22,45.5 9-11,22,27.3 12-14,18,50.0 15-17,76,28.9 18-20,199,26.1 21-23,262,31.3 24-26,234,32.5 27-29,236,32.6 30-32,234,33.8 33-35,156,30.8 36-38,161,24.8 39-41,148,31.8 42-44,87,24.1 45-47,89,36.0 48-50,67,40.3 51-53,29,48.3 54-56,31,32.3 57-59,26,30.8 60-62,21,38.1 63-65,16,18.8 66-68,3,0.0 69-71,4,0.0 72-74,2,0.0
Checksums
Listing 11 creates a manifest covering all five CSV files. Hashes detect file changes; they do not prove historical accuracy.
set -o pipefail
(
cd data
if command -v sha256sum >/dev/null 2>&1; then
sha256sum titanic.csv titanic_by_class.csv titanic_by_sex.csv \
titanic_by_age_and_class.csv titanic_by_age.csv
else
shasum -a 256 titanic.csv titanic_by_class.csv titanic_by_sex.csv \
titanic_by_age_and_class.csv titanic_by_age.csv
fi
) | tee data/SHA256SUMS
Check the generated files :noexport
Listing 12 checks the exported totals against the source snapshot. The following checks cover row counts, empty age/class cells and valid percentages.
BEGIN { FS = "," }
NR > 1 { n += $2; s += $3; if ($2 != $3+$4) bad = 1 }
END { exit (NR != 5 || n != 2207 || s != 711 || bad) }
Listing 13 checks the separate female and male counts.
BEGIN { FS = "," }
NR > 1 {
n += $2+$5; s += $3+$6
if ($2 != $3+$4 || $5 != $6+$7) bad = 1
}
END { exit (NR != 5 || n != 2207 || s != 711 || bad) }
Listing 14 checks all 100 age/class cells.
BEGIN { FS = "," }
NR > 1 {
n += $4; s += $5; empty += ($4 == 0)
if (($4 == 0 && $3 != "NA") ||
($4 > 0 && ($3 == "NA" || $3 < 0 || $3 > 100))) bad = 1
}
END { exit (NR != 101 || n != 2205 || s != 711 || empty != 12 || bad) }
Listing 15 checks all 25 age-only groups.
BEGIN { FS = "," }
NR > 1 {
n += $2
if ($2 > 0 && ($3 == "NaN" || $3 < 0 || $3 > 100)) bad = 1
}
END { exit (NR != 26 || n != 2205 || bad) }