Titanic Passengers and Crew

Demographics and Survival Data

Purpose and outputs

This document contains 2,207 person-level records, their provenance and all calculations needed to produce the CSV datasets below. Tangle the embedded base CSV, then evaluate the AWK blocks. Org saves their CSV output using :results output file and :file. No external data download or generated export script is needed. This document is extracted from dataset-titanic.dtx. Edit the embedded records or calculations in that DTX source, not the generated files. For an independent offline rebuild, run make datasets.

Requirements: Emacs with AWK and Bash support enabled in Org Babel, Bash, AWK, head, cat, tee and either sha256sum or shasum. All paths are relative to the directory containing this document. Tangling creates data/ if needed.

Output in data/ Rows Contents
titanic.csv 2,207 Person-level class, sex, age, survival and embarkation
titanic_by_class.csv 4 Outcome counts by passenger class and crew membership
titanic_by_sex.csv 4 Outcome counts by class/crew and sex
titanic_by_age_and_class.csv 100 Counts and survival percentages by three-year age band and class/crew
titanic_by_age.csv 25 Counts and survival percentages by three-year age band
SHA256SUMS 5 Checksums of the five CSV files

All datasets are comma-separated .csv files. They use UTF-8, one header row, LF line endings and decimal points. Row counts exclude headers. NA marks unknown ages or empty age/class cells; the age-only table uses NaN for an empty group's percentage. Missing values are never zero-filled.

Provenance and limitations

Dataset collection: doi:10.5281/zenodo.22967792. This is the all-versions (concept) DOI for Titanic Passengers and Crew: Demographics and Survival Data. Use it to find the dataset collection; use a version-specific DOI when citing the exact files used in an analysis.

The direct source was Dennis Sun's Principles of Data Science Titanic CSV, accessed 25 September 2026. Selected fields were manually transcribed from the browser-readable CSV, then exported and sorted. Four crew categories were merged into Crew; gender became sex; names and identifiers were omitted. Identical five-field rows remain because they may represent different people. The historical extraction is not reproduced by this file; only transformations from the embedded snapshot onward are reproducible.

Treat this as a teaching snapshot, not a definitive historical reconstruction. Ages and transcriptions may contain errors. Crew is not a fourth passenger class. Pooling sexes or crew roles can hide differences, and small cells need their counts shown alongside percentages. These data do not establish causes.

Person-level records

Column Meaning Values
class Passenger class or crew membership 1st, 2nd, 3rd, Crew
sex Sex as recorded in the source female, male
age Reported age in years at the sinking Numeric, including fractional infant ages; NA for unknown age
survived Recorded survival outcome 1 = survived; 0 = did not survive
embarked Embarkation location B = Belfast; C = Cherbourg; Q = Queenstown; S = Southampton

The complete records are embedded in the Org source and tangled unchanged as data/titanic.csv. The following section is omitted from HTML to avoid printing 2,207 lines; the downloadable CSV retains them all.

What the snapshot preserves

Two ages are NA; other fields have no missing values. Fractional infant ages are retained. Records are ordered by class, sex, numeric age, outcome and embarkation, with missing ages last within class/sex groups. The build does not impute ages, deduplicate people or reconstruct identifiers.

Listing 1 shows ten lines without truncating the files used for calculations. Set rows to change only this preview.

Listing 1 — Preview the generated person-level CSV
head -n "$rows" data/titanic.csv
class,sex,age,survived,embarked
1st,female,2,0,S
1st,female,13,1,S
1st,female,16,1,C
1st,female,16,1,S
1st,female,17,1,C
1st,female,17,1,S
1st,female,18,1,C
1st,female,18,1,C
1st,female,18,1,C

Generate the datasets

In Emacs, first use C-c C-v t to tangle the embedded data/titanic.csv. Then evaluate the following blocks in order with C-c C-c. The AWK blocks read their CSV input through :in-file. FS = OFS = "," sets comma-separated input and output. :results output file with :file data/filename.csv tells Org to save the printed CSV as that file. The preview blocks display it. Only the previews and checksum commands use Bash; no export script is needed.

Use a decimal-point numeric locale. If necessary, set Emacs's process environment with M-: (setenv "LC_ALL" nil) and M-: (setenv "LC_NUMERIC" "C") before evaluating the AWK blocks.

Alternatively, after tangling, use C-c C-v b to evaluate all executable blocks in document order, including previews and checks. No Python, Makefile or network access is required.

Validate the snapshot

Listing 2 rejects malformed rows, unexpected categories, ages outside the documented range and changed cohort totals. Tangling preserves every record and its order. The five embedded fields contain neither quoted text nor commas, so FS = "," is sufficient here; it is not a general parser for arbitrary quoted CSV files.

Listing 2 — Validate the embedded data/titanic.csv snapshot
BEGIN { FS = "," }
 NR == 1 {
  if ($0 != "class,sex,age,survived,embarked") bad = 1
  next
}
{
  if (NF != 5 || $1 !~ /^(1st|2nd|3rd|Crew)$/ ||
      $2 !~ /^(female|male)$/ || $4 !~ /^[01]$/ ||
      $5 !~ /^[BCQS]$/) bad = 1
  if ($3 == "NA") missing++
  else if ($3 !~ /^[0-9]+([.][0-9]+)?$/ || $3 < 0 || $3 >= 75) bad = 1
  n++; survived += $4; crew += ($1 == "Crew")
}
END {
  if (bad || n != 2207 || survived != 711 || crew != 890 || missing != 2) {
    print "Invalid Titanic snapshot" > "/dev/stderr"; exit 1
  }
}

Outcomes by class and crew

Listing 3 counts everyone, including the two unknown ages. For each group, deaths are aboard minus survivors. The explicit loop fixes the display order; associative-array iteration would not guarantee it.

Listing 3 — Generate data/titanic_by_class.csv
 BEGIN { FS = OFS = ","; print "Class", "Aboard", "Survived", "Died" }
NR > 1 { n[$1]++; s[$1] += $4 }
END {
  split("1st 2nd 3rd Crew", label, " ")
  for (c = 1; c <= 4; c++) {
    k = label[c]; print k, n[k]+0, s[k]+0, n[k]-s[k]
  }
}

data/titanic_by_class.csv

Listing 4 displays the generated class-and-crew counts.

Listing 4 — Outcomes by class and crew
cat data/titanic_by_class.csv
Class,Aboard,Survived,Died
1st,324,201,123
2nd,284,118,166
3rd,709,181,528
Crew,890,211,679

Outcomes by class and sex

Listing 5 uses a two-part key. Women's counts occupy columns 2–4, men's columns 5–7. Adding the two sexes recovers the class totals.

Listing 5 — Generate data/titanic_by_sex.csv
 BEGIN {
  FS = OFS = ","
  print "Class", "Women aboard", "Women survived", "Women died", \
        "Men aboard", "Men survived", "Men died"
}
NR > 1 { n[$1, $2]++; s[$1, $2] += $4 }
END {
  split("1st 2nd 3rd Crew", label, " ")
  for (c = 1; c <= 4; c++) {
    k = label[c]
    print k, n[k,"female"]+0, s[k,"female"]+0, n[k,"female"]-s[k,"female"], \
             n[k,"male"]+0, s[k,"male"]+0, n[k,"male"]-s[k,"male"]
  }
}

data/titanic_by_sex.csv

Listing 6 displays the generated table.

Listing 6 — Outcomes by class and sex
cat data/titanic_by_sex.csv
Class,Women aboard,Women survived,Women died,Men aboard,Men survived,Men died
1st,144,139,5,180,62,118
2nd,106,94,12,178,24,154
3rd,216,106,110,493,75,418
Crew,23,20,3,867,191,676

Three-year age bands and class

Listing 7 groups unrounded ages into [0,3), [3,6), through [72,75), pooling sexes. The labels 1, 4, …, 73 are the middle completed-year ages, not exact continuous interval midpoints. Class codes 1–4 identify first, second, third and crew. Unknown ages are excluded.

All 25 × 4 combinations are emitted, ordered by age and then class. Empty cells have n=0, survivors=0 and NA, not 0%. Percentages have six decimal places. Each age band is followed by its four class/crew categories in order.

Listing 7 — Generate data/titanic_by_age_and_class.csv
 BEGIN {
  FS = OFS = ","
  code["1st"]=1; code["2nd"]=2; code["3rd"]=3; code["Crew"]=4
  print "age_midpoint", "class_code", "survival_pct", "n", "survivors"
}
NR > 1 && $3 != "NA" {
  band = int($3 / 3); c = code[$1]
  n[band,c]++; s[band,c] += $4
}
END {
  for (band = 0; band < 25; band++)
    for (c = 1; c <= 4; c++)
      print 3*band+1, c, \
            n[band,c] ? sprintf("%.6f", 100*s[band,c]/n[band,c]) : "NA", \
            n[band,c]+0, s[band,c]+0
}

data/titanic_by_age_and_class.csv

Listing 8 — Known-age counts and survival rates
cat data/titanic_by_age_and_class.csv
age_midpoint,class_code,survival_pct,n,survivors
1,1,50.000000,2,1
1,2,100.000000,10,10
1,3,45.454545,22,10
1,4,NA,0,0
4,1,100.000000,1,1
4,2,100.000000,6,6
4,3,57.142857,21,12
4,4,NA,0,0
7,1,100.000000,1,1
7,2,100.000000,6,6
7,3,20.000000,15,3
7,4,NA,0,0
10,1,100.000000,1,1
10,2,NA,0,0
10,3,23.809524,21,5
10,4,NA,0,0
13,1,100.000000,2,2
13,2,80.000000,5,4
13,3,27.272727,11,3
13,4,NA,0,0
16,1,83.333333,6,5
16,2,57.142857,7,4
16,3,26.666667,45,12
16,4,5.555556,18,1
19,1,83.333333,12,10
19,2,33.333333,24,8
19,3,21.649485,97,21
19,4,19.696970,66,13
22,1,80.952381,21,17
22,2,31.034483,29,9
22,3,23.893805,113,27
22,4,29.292929,99,29
25,1,70.588235,17,12
25,2,35.294118,34,12
25,3,39.240506,79,31
25,4,20.192308,104,21
28,1,66.666667,21,14
28,2,46.428571,28,13
28,3,21.794872,78,17
28,4,30.275229,109,33
31,1,65.217391,23,15
31,2,36.363636,33,12
31,3,29.629630,54,16
31,4,29.032258,124,36
34,1,88.888889,18,16
34,2,41.176471,17,7
34,3,24.324324,37,9
34,4,19.047619,84,16
37,1,51.851852,27,14
37,2,31.578947,19,6
37,3,23.076923,26,6
37,4,15.730337,89,14
40,1,57.142857,28,16
40,2,46.153846,13,6
40,3,10.000000,30,3
40,4,28.571429,77,22
43,1,57.894737,19,11
43,2,20.000000,10,2
43,3,5.263158,19,1
43,4,17.948718,39,7
46,1,51.515152,33,17
46,2,10.000000,10,1
46,3,30.769231,13,4
46,4,30.303030,33,10
49,1,66.666667,27,18
49,2,60.000000,10,6
49,3,0.000000,9,0
49,4,14.285714,21,3
52,1,75.000000,12,9
52,2,20.000000,5,1
52,3,0.000000,1,0
52,4,36.363636,11,4
55,1,53.333333,15,8
55,2,40.000000,5,2
55,3,0.000000,4,0
55,4,0.000000,7,0
58,1,38.461538,13,5
58,2,40.000000,5,2
58,3,0.000000,2,0
58,4,16.666667,6,1
61,1,46.153846,13,6
61,2,25.000000,4,1
61,3,0.000000,2,0
61,4,50.000000,2,1
64,1,25.000000,8,2
64,2,0.000000,3,0
64,3,25.000000,4,1
64,4,0.000000,1,0
67,1,0.000000,1,0
67,2,NA,0,0
67,3,0.000000,2,0
67,4,NA,0,0
70,1,0.000000,3,0
70,2,NA,0,0
70,3,0.000000,1,0
70,4,NA,0,0
73,1,NA,0,0
73,2,0.000000,1,0
73,3,0.000000,1,0
73,4,NA,0,0

Three-year age bands for age trends

Listing 9 uses the same 25 three-year intervals as the age/class summary: [0,3), [3,6), through [72,75). Labels 0-2, 3-5, …, 72-74 describe completed years; fractional ages are grouped without rounding. There is no combined 60+ group. All classes, crew and both sexes are pooled, and the same two unknown ages are excluded. Percentages have one decimal place; empty groups use NaN to distinguish an undefined percentage from zero.

Listing 9 — Generate data/titanic_by_age.csv
 BEGIN { FS = OFS = ","; print "Age", "People", "Survived %" }
NR > 1 && $3 != "NA" {
  g = int($3 / 3)
  n[g]++; s[g] += $4
}
END {
  for (g = 0; g < 25; g++)
    print sprintf("%d-%d", 3*g, 3*g+2), n[g]+0, \
          n[g] ? sprintf("%.1f", 100*s[g]/n[g]) : "NaN"
}

data/titanic_by_age.csv

Listing 10 displays all 25 bins.

Listing 10 — Known-age counts and survival rates
cat data/titanic_by_age.csv
Age,People,Survived %
0-2,34,61.8
3-5,28,67.9
6-8,22,45.5
9-11,22,27.3
12-14,18,50.0
15-17,76,28.9
18-20,199,26.1
21-23,262,31.3
24-26,234,32.5
27-29,236,32.6
30-32,234,33.8
33-35,156,30.8
36-38,161,24.8
39-41,148,31.8
42-44,87,24.1
45-47,89,36.0
48-50,67,40.3
51-53,29,48.3
54-56,31,32.3
57-59,26,30.8
60-62,21,38.1
63-65,16,18.8
66-68,3,0.0
69-71,4,0.0
72-74,2,0.0

Checksums

Listing 11 creates a manifest covering all five CSV files. Hashes detect file changes; they do not prove historical accuracy.

Listing 11 — Generate data/SHA256SUMS
set -o pipefail
(
  cd data
  if command -v sha256sum >/dev/null 2>&1; then
    sha256sum titanic.csv titanic_by_class.csv titanic_by_sex.csv \
      titanic_by_age_and_class.csv titanic_by_age.csv
  else
    shasum -a 256 titanic.csv titanic_by_class.csv titanic_by_sex.csv \
      titanic_by_age_and_class.csv titanic_by_age.csv
  fi
) | tee data/SHA256SUMS

Check the generated files :noexport

Listing 12 checks the exported totals against the source snapshot. The following checks cover row counts, empty age/class cells and valid percentages.

Listing 12 — Check class counts and outcome totals
BEGIN { FS = "," }
NR > 1 { n += $2; s += $3; if ($2 != $3+$4) bad = 1 }
END { exit (NR != 5 || n != 2207 || s != 711 || bad) }

Listing 13 checks the separate female and male counts.

Listing 13 — Check class-and-sex counts and outcome totals
BEGIN { FS = "," }
NR > 1 {
  n += $2+$5; s += $3+$6
  if ($2 != $3+$4 || $5 != $6+$7) bad = 1
}
END { exit (NR != 5 || n != 2207 || s != 711 || bad) }

Listing 14 checks all 100 age/class cells.

Listing 14 — Check age/class totals and empty cells
BEGIN { FS = "," }
NR > 1 {
  n += $4; s += $5; empty += ($4 == 0)
  if (($4 == 0 && $3 != "NA") ||
      ($4 > 0 && ($3 == "NA" || $3 < 0 || $3 > 100))) bad = 1
}
END { exit (NR != 101 || n != 2205 || s != 711 || empty != 12 || bad) }

Listing 15 checks all 25 age-only groups.

Listing 15 — Check age-only counts and percentages
BEGIN { FS = "," }
NR > 1 {
  n += $2
  if ($2 > 0 && ($3 == "NaN" || $3 < 0 || $3 > 100)) bad = 1
}
END { exit (NR != 26 || n != 2205 || bad) }