The Gene Channel
← All episodes

completeddiploidgenome

The Missing 8%

In 2003 the world celebrated a finished human genome. It wasn't — 8% was still blank, and the 'complete' version belonged to no one alive. This is the fifty-year climb to finally read a real person's genome, both parental copies, cover to cover.

The walkthrough

Beat by beat

completeddiploidgenome — HOOK

01HOOK

In 2003, the world threw a party. The human genome — finished. Read from cover to cover. Except it wasn't. Eight percent of the book was still blank `F10`. Whole chapters no one could read. The centromeres. The tips of the chromosomes. The parts that break in cancer. And that blank stayed blank for nineteen more years. When we finally filled it in, the finished book turned out to belong to someone in particular.

completeddiploidgenome — THE FIRST LETTER

02THE FIRST LETTER

To read a genome, first you have to read a single letter. In 1977, a quiet Englishman named Fred Sanger found the way `F1`. Chop the strand into pieces. Sort them by length. Read the ladder, one rung at a time `F1`. He'd already won a Nobel Prize once — for spelling out insulin, the first protein ever sequenced `F1`. Now he'd earn a second, for teaching the world to read DNA `F1`. Only one person in history has ever won that prize in chemistry twice. It was him. His method was slow. It was hand-done. And for thirty years, it was the only way there was `F1`.

completeddiploidgenome — THE AMBITION

03THE AMBITION

Reading one gene was a PhD. Reading all three billion letters was a moonshot. In 1990, the world took the shot: the Human Genome Project. Fifteen years. Three billion dollars. Every nation pitching in `F2`. The man who launched it was James Watson, co-discoverer of the double helix itself `F3`. But almost at once, a fight broke out. Not over science. Over ownership. The government wanted to patent raw gene fragments, thousands of them `F3`. Watson was furious. You cannot patent nature, he argued. In 1992, he quit rather than sign on `F3`. A young geneticist named Francis Collins took the wheel, and the public project sailed on: slow, careful, mapping before it read `F4`.

completeddiploidgenome — THE RACE

04THE RACE

Then, in 1998, a challenger appeared. Craig Venter, and a private company called Celera `F5`. His bet was speed. New machines had just arrived — automated sequencers that ran DNA through hair-thin tubes, hundreds at a time `F6`. Celera bought three hundred of them and read a billion letters a month `F6`. His method was reckless and brilliant: blow the whole genome into random fragments, and let the computers puzzle it back together `F5`. The public consortium mapped carefully. Celera sprinted. It became the most famous race in biology. It ended, on purpose, in a tie. June 2000, the White House — Collins and Venter shoulder to shoulder, announcing a draft `F7`. The finished papers came the next winter, back to back — the public team in one journal, Celera in another `F8`.

completeddiploidgenome — THE CATCH

05THE CATCH

Three years later, in 2003, they called it done `F9`. Ninety-nine percent of the genes, spelled out letter by letter `F9`. And that was worth every dollar. But a genome is not only its genes. About eight percent of it had never been read at all `F10`. Not because it didn't matter. Because it couldn't be done `F12`.

completeddiploidgenome — THE DESERTS

06THE DESERTS

The missing eight percent was the hardest terrain in the whole genome `F11`. The centromeres — the waists of each chromosome, the anchors that let a cell divide without losing its place `F11`. The telomeres, capping the ends. Long stretches copied and re-copied, stacked like dunes in a desert `F11`. The machines read in short snippets. In a desert where every dune looks identical, a short snippet has no landmarks, no way to know where it came from `F12`. So the assemblers gave up on the deserts, and left them white on the map. Eight percent of you, unmapped `F10`.

completeddiploidgenome — CHEAPER

07CHEAPER

Meanwhile, sequencing got cheap. Absurdly cheap. A new generation of machines read millions of tiny fragments in parallel. The price of a genome fell off a cliff — from billions of dollars toward a thousand `F13`. Suddenly a genome could be read in a day, for the cost of a laptop `F13`. But to go that cheap, these machines read in even shorter snippets `F13`. Great for the easy ninety-two percent. Useless in the deserts. The cheaper we made it, the more stubborn the blank stayed.

completeddiploidgenome — LONGER

08LONGER

The answer wasn't a faster computer. It was a longer read. A different kind of machine arrived — from companies called PacBio and Oxford Nanopore `F14`. Instead of snippets, they read DNA in long, continuous ribbons. Tens of thousands of letters in a single pass `F14`. Long enough to carry a landmark on both ends. Long enough to walk clear across a desert without ever losing your footing `F14`. For the first time, the blank could be read. Someone had to go and do it.

completeddiploidgenome — TELOMERE TO TELOMERE

09TELOMERE TO TELOMERE

2022. A team with a fitting name, Telomere-to-Telomere, finished the map `F15`. Two hundred million letters no reference had ever contained. Nearly two thousand new genes `F15`. The first gapless human genome. Every centromere and telomere, resolved end to end. And the blank turned out to be anything but junk. The centromeres hold the machinery that lets a cell split its chromosomes evenly every time it divides. The short arms hold the clustered genes that build our ribosomes — the factories for every protein you make `F11`. Filling the deserts also corrected thousands of long-standing errors, and made variants we'd been misreading for years — even clinically important ones — finally legible `F23`. But it came with a quiet asterisk. To make the puzzle solvable, they used a very unusual cell, one that carries two identical copies of every chromosome `F16`. In effect, a single genome, read twice. Cleaner than any real person. And missing a Y chromosome entirely. That one was finished, separately, a year later `F17`. A flawless reference. But not a human being. Because a human being is not one genome. It's two.

completeddiploidgenome — WHY TWO

10WHY TWO

You are diploid. Two copies of every chromosome, one from your mother, one from your father `F18`. And they are not the same. They differ, letter by letter, millions of times over. That difference is where medicine lives. A recessive disease needs both copies broken `F18`. Often it's two different faults, one on each copy, that only cause harm together `F18`. So the question that matters isn't just what variants you carry. It's which copy each one sits on `F18`. Answering that is called phasing. And for twenty years, a single-genome reference could never quite do it.

completeddiploidgenome — THE COMPLETE DIPLOID (hero)

11THE COMPLETE DIPLOID (hero)

August 2026. The same team did it for real `F19`. Not a mole. Not a composite. A living, consented donor — known to science as HG002 `F19`. Both chromosome sets. Both parents' copies. Fully separated, fully phased, telomere to telomere `F19`. Nine hundred million letters more than the best benchmark before it. Fifteen percent more genome, laid bare `F19`. Led out of Johns Hopkins and the national standards lab, published all at once as a dozen papers — with the genomes of other animals alongside, from macaque to giraffe `F19`. For the first time, this was not a human genome. It was a whole human being, written out complete.

completeddiploidgenome — WHERE IT GOES

12WHERE IT GOES

For twenty years, reading your genome meant lining you up against one borrowed reference, mostly the DNA of a handful of strangers, and noting where you differed `F21`. But any two people differ at millions of letters, and some stretches of DNA that whole populations carry are simply missing from that one reference `F20`. Measure everyone against a single stranger and you miss exactly what makes them different. That is why no one person can be the human genome. So science is building a pangenome: dozens of complete genomes, and climbing, so your DNA is read against people like you, not one borrowed map `F20`. And beyond that, the real prize. Not the reference. Yours. Every letter, both copies, yours alone `F21`. The first genome cost billions and took thirteen years `F22`. Yours will cost a few hundred dollars and take a couple of days `F22`. The book that was never finished is finally being written — a complete, personal copy for each of us, one person at a time.

completeddiploidgenome — TIMELINE + SIGN-OFF

13TIMELINE + SIGN-OFF

1977, the first letter. 2000, the draft. 2003, the deserts still blank. 2022, gapless at last. 2026, complete — and, for the first time, diploid. Half a century to learn to read ourselves, cover to cover. — The Gene Channel.

The write-up

In one line: The genome we "finished" in 2003 was missing 8% of the book and belonged to no one in particular — this is the fifty-year climb to finally read a real person's genome, both parental copies, cover to cover.


The landmark — and what it left unread

The 2003 genome was one of the great achievements of modern science, and it delivered. Even as a draft, the human reference sequence launched the entire era of genomic medicine: it let researchers pin down the genes behind thousands of inherited diseases, powered the genome-wide association studies that connect DNA to common conditions, opened cancer genomics and pharmacogenomics, and drove the cost of reading a genome from billions of dollars toward a thousand. Almost everything this channel covers stands on that foundation.

But the celebrated genome was never quite finished. About 8% was never read at all — the centromeres, the telomeres, the dense repetitive stretches that also happen to break in cancer. Those blanks stayed white on the map for nineteen more years, and when they were finally filled, the "complete" genome turned out to belong to a very particular — and very unusual — cell.

Learning to read a letter, then a genome

  • 1977 — the first letter. Fred Sanger's chain-termination method taught the world to read DNA one rung of a ladder at a time. Slow and hand-done, it stayed the only way for roughly thirty years. Sanger had already won a Nobel for sequencing insulin; DNA earned him a second — he remains the only person to win the Chemistry prize twice.
  • 1990 — the moonshot. The Human Genome Project launched: ~15 years, ~$3 billion, an international consortium. Its first head, James Watson, resigned in 1992 rather than accept a US push to patent raw gene fragments. Francis Collins took over and mapped carefully, clone by clone.
  • 1998–2001 — the race. Craig Venter's Celera bet on speed: ~300 automated capillary sequencers reading a billion letters a month, and a "shotgun" strategy that blew the genome into random fragments for computers to reassemble. The public consortium and Celera announced a draft together at the White House in June 2000, then published back-to-back the next winter in two different journals.

Why the deserts stayed blank

The 2003 "complete" genome covered ~99% of the gene-containing regions — but a genome is more than its genes. The missing 8% was the hardest terrain: satellite repeats, segmental duplications, the acrocentric short arms. Short-read machines read in tiny snippets, and in a desert where every dune looks identical, a snippet has no landmarks — nowhere to place it. Cheaper next-generation sequencing crashed the price toward the "$1,000 genome" but read in even shorter snippets, so the blank only grew more stubborn.

The fix wasn't a faster computer — it was a longer read. PacBio HiFi and Oxford Nanopore read DNA in continuous ribbons tens of thousands of letters long, enough to carry a landmark on both ends and walk clear across a repeat.

Telomere to telomere — and the quiet asterisk

In 2022, the Telomere-to-Telomere consortium delivered the first gapless human genome (T2T-CHM13): ~200 million new letters, nearly 2,000 new genes. But to make the puzzle solvable they used an unusual cell (a hydatidiform mole) carrying two identical copies of every chromosome — essentially one genome read twice, cleaner than any real person, and with no Y chromosome (finished separately in 2023). A flawless reference — but not a human being. Because a human being isn't one genome. It's two.

Why two matters

You are diploid: two copies of every chromosome, one from each parent, differing millions of times over. That difference is where medicine lives — recessive disease needs both copies hit, often by two different faults (compound heterozygosity). So the question isn't only what variants you carry but which copy each sits on. Answering that is phasing, and a single-genome reference could never quite do it.

The complete diploid genome (2026)

In August 2026 the same lineage did it for real: a complete, fully phased diploid genome of HG002 — a living, consented GIAB donor, not a synthetic or composite. Both parental chromosome sets, separated and resolved telomere to telomere; ~900 million letters (~15% more genome) beyond the prior benchmark regions. Led out of Johns Hopkins and NIST and published as a ~dozen-paper package alongside other vertebrate genomes. For the first time this wasn't a human genome — it was a whole human being, written out complete.

Why a complete, representative genome matters now

For twenty years, reading your genome meant lining you up against one borrowed reference and noting the differences. That worked well enough to build modern genetics — but the reference was assembled mostly from a handful of donors, and it skews toward one slice of human ancestry. Any two people differ at millions of letters, and some sequence whole populations carry is simply missing from that one map. When your DNA is measured against a stranger who doesn't share your ancestry, the reference misses variants, misplaces others, and quietly performs worse — and in clinical genomics that gap falls hardest on the communities already underserved by medicine.

That is the case for a pangenome: not one genome but dozens of complete ones and climbing, chosen to span the real diversity of the species, so everyone's DNA is read against people like them rather than one borrowed template. A complete, telomere-to-telomere assembly makes each of those genomes trustworthy even in the hardest regions; doing it for many people makes the reference fair. And beyond the shared map, the real prize is yours — your own genome, every letter, both parental copies, phased and complete. The first genome cost billions and took thirteen years; yours will cost a few hundred dollars and take a couple of days.

Sources

Full claim-by-claim evidence is in references.md. Primary anchors:

  • Sanger et al., chain-termination sequencing, PNAS 1977 (PMID 271968); Nobel Chemistry 1958 + 1980.
  • NHGRI Human Genome Project history: 1990 launch, June 2000 draft, 2003 completion.
  • Nurk et al., "The complete sequence of a human genome," Science 2022 (PMID 35357919); Rhie et al., complete Y chromosome, Nature 2023.
  • T2T Consortium, "A complete diploid human genome benchmark for personalized genomics," Cell 2026 (S0092-8674(26)00703-8).
  • Liao et al., "A draft human pangenome reference," Nature 2023 (PMID 37165242).

Accuracy note: "Complete" in 2003 meant the euchromatic (gene-rich) ~99%, not the whole sequence. T2T-CHM13 (2022) is gapless but comes from a near-homozygous mole cell line — "a human genome," not a diploid person. The 2026 milestone's "~15% more" is measured against prior benchmark regions; HG002 is a real consented individual (GIAB RM 8391 / GM24385), not synthetic; the companion-species count is kept qualitative in narration.

The evidence

Every claim, sourced

Each [F#] you hear in the film links to the source it came from. Nothing gets narrated until every one is checked and signed off.

Fact-gate
Open
PhD sign-off

Sign-off

  • PhD sign-off — facts above are correct; traps (F1/F3/F8/F9/F10/F13/F16/F19/F20/F22/F23) stated correctly in script.md.
  • Any numbers/dates verified (or narration kept qualitative). Open item: confirm the exact companion-species count/list from the Cell Genomics DOIs before adding a number to narration (currently qualitative).

On sign-off → run `gen-narration.mjs` (the gate opens). Then assets → Video.tsx → render → `writeup.md`.

  1. F1

    1977: Sanger chain-termination sequencing; slow, hand-done, dominant ~30 yrs. His first Nobel (1958) was insulin/protein sequencing; DNA earned a second (Chem 1980) — only person to win Chemistry twice

    Sanger et al., PNAS 1977; Nobel Chemistry 1958 (insulin) + 1980 (nucleic-acid sequences, w/ Gilbert & Berg). Trap: specify molecule/decade; "second Nobel."

  2. F2

    1990: Human Genome Project — ~15 yrs, ~$3B, international

    NHGRI HGP Fact Sheet (officially began Oct 1, 1990). Trap: $3B = projected budget; actual ≈ $2.7B (see F22).

  3. F3

    James Watson (double-helix co-discoverer) was first head of the genome office; resigned April 1992 over the NIH's push to patent gene fragments

    Watson led NCHGR from 1988; resigned Apr 1992 opposing NIH (B. Healy) patenting of EST/cDNA fragments. Trap: year 1992 (not 1990/93); institution = NCHGR; a conflict-of-interest allegation was secondary to the patenting dispute.

  4. F4

    Francis Collins took over the public effort (1993); mapped carefully / clone-by-clone

    NHGRI Collins bio (directed NCHGR from Apr 1993; renamed NHGRI 1997). Clone-by-clone (BAC map-first) vs Celera's shotgun.

  5. F5

    1998: Craig Venter / Celera; whole-genome shotgun — fragment everything, reassemble by computer

    NHGRI "1998: Company Announces Sequencing Plan"; Venter et al., Science 2001. Trap: Celera also used public HGP data — not "entirely independently."

  6. F6

    Automated capillary sequencers (ABI PRISM 3700, 1998); Celera ran ~300 of them, ~1 billion bases/month

    PE/Applied Biosystems 3700 (1998); Celera deployed 300 units. Trap: the 3700 (capillary), not the older ABI 377 (slab-gel).

  7. F7

    June 26, 2000: White House announces a draft; Collins & Venter together

    NHGRI June 2000 event (Clinton; Blair via video link). Trap: a DRAFT, not completion.

  8. F8

    The papers came back-to-back the next winter, in different journals

    HGP consortium in Nature Feb 15, 2001; Celera in Science Feb 16, 2001 (both ~90% draft). Trap: consecutive days, two journals — not "both Feb 15."

  9. F9

    April 14, 2003: HGP declared complete; ~99% of gene-containing (euchromatic) regions

    NHGRI 2003 release (2 yrs early; ~99% euchromatin at 99.99%). Trap: "complete" = euchromatic only; excluded heterochromatin.

  10. F10

    ~8% stayed blank — and it was the parts that break in cancer — until 2022

    ~8% of GRCh38 unresolved (heterochromatic repeats) through 2021; filled by T2T 2022.

  11. F11

    The blanks were the hardest terrain: centromeres (needed for a cell to divide), telomeres, stacked repeats/segmental duplications, and the acrocentric short arms carrying rDNA

    Centromere→kinetochore/segregation (Altemose 2022); seg-dups = structural-variation/evolution hotspots (Vollger 2022); rDNA on acrocentric short arms of chr 13/14/15/21/22 (Nurk 2022). Trap: rDNA on those five acrocentrics specifically.

  12. F12

    Short reads can't span repeats — no landmarks in a desert of identical dunes

    Nurk et al. 2022 framing; short reads don't align uniquely in satellite arrays / segmental dups.

  13. F13

    Next-gen/short-read sequencing crashed the price toward the "$1,000 genome," in a day — but read in even shorter snippets

    Illumina HiSeq X Ten, Jan 2014, "$1,000 genome." Trap: the $1,000 was a reagent cost at high volume, not all-in clinical cost; contested at the time.

  14. F14

    Long reads (PacBio HiFi, Oxford Nanopore) — tens of thousands of letters per pass; span the repeats

    PacBio HiFi ~15–20 kb (>Q20); ONT ultra-long often >100 kb. Combined for T2T.

  15. F15

    2022: T2T-CHM13 — first gapless genome; +~200 Mbp, ~2,000 new genes

    Nurk et al., "The complete sequence of a human genome," Science 376:44 (2022); +200 Mbp, 1,956 gene predictions.

  16. F16

    CHM13 has two identical chromosome copies — "a single genome read twice," not a person; no Y

    Complete hydatidiform mole (single sperm duplicated → essentially haploid/homozygous; 46,XX, no Y). Central trap: not a diploid individual; paper says "a human genome."

  17. F17

    The Y chromosome was finished separately, a year later (2023)

    Rhie et al., "The complete sequence of a human Y chromosome," Nature 2023 (Y from HG002; ~62.4 Mb). Trap: CHM13 lacked a Y by design (46,XX line), not by error.

  18. F18

    You are diploid — two differing copies, one per parent; recessive disease needs both copies hit, often two different faults (compound het); which copy each sits on = phasing

    NHGRI Talking Glossary; GeneReviews glossary. Compound het = two different pathogenic variants in trans; phasing = cis vs trans assignment.

  19. F19

    Aug 6, 2026: complete diploid genome of HG002 — both haplotypes fully phased, telomere to telomere; +900M letters / ~15% more genome; led by JHU + NIST; a ~12-paper package incl. other vertebrates

    T2T Consortium, "A complete diploid human genome benchmark for personalized genomics," Cell 2026; +900 Mb / 15.3% vs prior benchmark regions; senior Adam Phillippy (JHU), co-senior Justin Zook (NIST). Traps: HG002 = real consented GIAB individual (RM 8391 / GM24385), not synthetic; 15% is vs prior benchmark regions; "8 species" appears only in secondary press → narration kept qualitative ("from macaque to giraffe").

  20. F20

    Any two people differ at millions of letters, and sequence whole populations carry is missing from one reference → a single reference can't represent everyone → a pangenome of dozens of complete genomes

    Liao et al., "A draft human pangenome reference," Nature 2023 — 47 individuals / 94 haplotypes; the draft pangenome adds ~119 Mb of sequence (largely structural variation) absent/under-represented in the single GRCh38 reference and reduces reference bias. Any two human genomes differ at ~4–5 million sites (SNVs) plus thousands of structural variants (well established; 1000 Genomes Project). Traps: 47 individuals (94 haplotypes), a draft; don't say "94 individuals." Narration stays qualitative ("millions of letters").

  21. F21

    Paradigm shift: from comparing you to one borrowed reference → reconstructing your own complete genome

    Phillippy: "a paradigm shift from … differences between your genome and a reference to actually reconstructing your complete, unique genome."

  22. F22

    HGP: ~13 yrs, billions; today a few hundred dollars in a couple of days

    HGP ≈ $2.7B actual (1990–2003, 13 yrs); NHGRI cost data ≈ $525/genome (2022); clinical WGS commonly <$1,000. Trap: $3B = projected budget; "$5B/$5,000" are inflation-adjusted/retail — narration stays qualitative.

  23. F23

    The 8% wasn't junk — what reading it revealed: centromeres = the machinery of chromosome segregation; acrocentric short arms = the rDNA genes that build ribosomes; and completing it corrected thousands of errors and improved variant calling, incl. clinically relevant regions, across a globally diverse cohort

    Centromere biology/segregation + rDNA on the five acrocentric short arms (Altemose 2022; Nurk 2022 — see F11). Error correction + universally improved read-mapping/variant detection across all 3,202 1000-Genomes samples: Aganezov et al., "A complete reference genome improves analysis of human genetic variation," Science 2022. Trap: "corrects thousands of errors" refers to fixes vs GRCh38; keep "junk" as the old view being overturned, not a claim the 8% is functionless.