What is a genome? the full library, nuclear and mitochondrial.

What is a genome? The complete genetic library of an organism: nuclear DNA plus mitochondrial DNA, about 3.2 billion bases in one haploid human copy.

A genome is the whole library you inherited

You asked what a genome is because DNA, gene, and genome get used as if they were one word. They are not. DNA is the molecule. A gene is a used stretch. A genome is the whole library.

Think of a library of 23 volumes in the nucleus, plus a pamphlet in each mitochondrion. The genome is every letter in that collection, genes and the vast noncoding remainder alike.

The real term is genome: the complete genetic material of an organism. For a human, that means nuclear DNA and mitochondrial DNA. Leave either out and the library is incomplete.

Nuclear plus mitochondrial is the honest total

The nuclear genome is the long DNA packed into chromosomes. One haploid copy is about 3.2 billion bases. Typical cells are diploid, so they keep two copies of autosomes, plus sex chromosomes.

The mitochondrial genome is a circle of about 16,600 bases. It encodes 37 genes and exists in many copies per cell. It is tiny next to the nucleus and not optional if you want the word genome to mean what biologists mean.

People sometimes say my genome when they mean a chip file of a few hundred thousand sites. That file is a sample from a genome. It is not the library.

  • Nuclear genome: chromosomes, about 3.2 billion bases haploid
  • Mitochondrial genome: a small circle, about 16,600 bases, many copies
  • Diploid cell: two copies of autosomes, plus X and Y as the person has them
  • A genotyping chip: a sample, not a genome

How large a human genome is when you stop rounding into myth

Three billion is the number that stuck. The more precise public figure is about 3.2 billion bases for one haploid nuclear genome. Two people still share more than 99 percent of that sequence.

Most of those bases are not protein-coding genes. About 20,000 genes code for proteins. The rest is noncoding sequence, including switches, repeats, and RNA genes. A genome is not a gene list.

Size is not destiny. Some plants have larger genomes than humans. Salamander genomes can dwarf ours. What matters for a body is what is encoded and how it is read, not who wins a base-pair contest.

A reference genome is a shared map, not a person

When a file says a variant is on chromosome 6 at position 26,090,951, that number only means something on a named map. The map is a reference genome: a composite sequence used so labs can talk about the same coordinates.

No living person is the reference. Early references leaned on a small number of donors. Later releases added patches, alternate loci, and more diversity. The map improved. It did not become you.

Your genome is diploid and personal. The reference is a haploid-like scaffold with notes. Comparing the two is how variants are called. The comparison is not a moral ranking, and the reference letter is not always the common letter in every population.

GRCh38 is the assembly hiding in most files

GRCh38 is the current human reference assembly you will meet in VCF headers, genome browsers, and many clinical pipelines. It is also called hg38. If a file does not name its assembly, treat coordinates as untrustworthy.

Many consumer chips still report positions on GRCh37, also called hg19. The same rsID can sit at different numeric positions on those two maps. Mix them and you are reading the wrong base.

A newer telomere-to-telomere assembly, T2T-CHM13, filled remaining gaps. It is a scientific landmark. GRCh38 remains the assembly most tools, ClinVar exports, and hospital files still assume. Meet the file where it lives. Lift-over tools exist. They introduce their own errors. Prefer an rsID when you have one, and read the assembly line in the header when you do not.

  • GRCh38 (hg38): current workhorse reference in most modern files
  • GRCh37 (hg19): still common in consumer raw-data tables
  • T2T-CHM13: gap-filled assembly, not yet the default in most clinics
  • rsIDs travel across assemblies better than raw chr:pos numbers

What the Human Genome Project actually finished

The Human Genome Project was an international effort to sequence a human reference, roughly 1990 to 2003. A draft appeared in 2001. An essentially complete version was announced in 2003. It did not sequence every person. It built a map.

One public surprise was the gene count. Predictions near 100,000 protein-coding genes collapsed toward about 20,000. Another was how much of the genome was not a tidy gene. The project earned its history because it gave biology a shared coordinate system.

Later work closed remaining gaps and started representing more than one haplotype at once. The Project is the beginning of a reference, not the last word on human variation.

Genome, gene, chromosome, DNA: four words, four jobs

DNA is the chemical. Genes are used stretches of that chemical. Chromosomes are how the nuclear genome is packaged so a cell can move it. The genome is the collection.

If you confuse the words, you will confuse the tests. A chromosome count is not a gene sequence. A gene panel is not a genome. A genome is not a diagnosis.

Keep the library metaphor if it helps, then put it down: 23 nuclear volumes, a mitochondrial pamphlet, millions of untranslated pages, and a catalog of about 20,000 protein-coding genes.

A DNA test is usually not a genome

Ancestry and health chips read hundreds of thousands of prechosen sites. Exome sequencing reads most protein-coding exons, a few percent of the nuclear genome. Whole-genome sequencing reads most of the letters, with coverage that still has holes.

Each of those assays can be useful. None of them is the word genome by itself. A chip file named genome.zip is a marketing filename. Ask what was measured.

Clinical care, when it needs a genome, uses a validated assay and a clinician. A curiosity lookup of a public identifier is a different job. This page is informational, not a diagnosis.

Public identifiers, private sequence

You can ask what a gene or an rsID means in public databases without uploading a genome. genome.sh is built for that split: identifiers over the network, sequence on your disk if you have sequence at all.

A genome remains a private object. The map it is compared to is public. Mixing those two is how people accidentally send a life to a server that asked only for a name.

If a service will only explain GRCh38, an rsID, or a gene after you upload raw data, it is bundling a public question with a private file. Unbundle them. The map is public. The sequence is yours.

Questions

What is a genome vs DNA?

DNA is the molecule. A genome is the complete collection of genetic DNA in an organism, nuclear plus mitochondrial in humans. Genes are used stretches inside that collection.

How big is the human genome?

About 3.2 billion bases in one haploid copy of the nuclear genome, plus a mitochondrial circle of about 16,600 bases. Diploid cells keep two copies of the autosomes.

What is GRCh38?

The current human reference assembly most modern files and browsers use, also called hg38. Coordinates only make sense on a named assembly. Many consumer files still use GRCh37.

Do humans have one genome or two?

You have one nuclear genome, present in two copies for autosomes, plus a mitochondrial genome present in many copies. People also say a maternal and paternal haplotype, which are the two editions of the nuclear library.

What is the Human Genome Project?

An international project that produced a human reference sequence, with a draft in 2001 and an essentially complete version in 2003. It built a map for a species. It did not sequence you.

What is the difference between nuclear and mitochondrial genome?

Nuclear DNA is packed into chromosomes in the nucleus and holds almost all of the sequence. Mitochondrial DNA is a small circle in mitochondria, inherited almost always along the maternal line.

Is a DNA test a whole genome?

Usually not. Chips sample prechosen sites. Exomes read protein-coding regions. Whole-genome sequencing reads most letters and still has limits. Read the assay, not the marketing filename.

What does reference genome mean?

A shared map, currently GRCh38 for most tools, used so positions can be compared. It is a composite, not a person, and the reference letter is not always the most common allele in every population.

genome.sh reports public annotations. It is informational software, not a medical device or a substitute for clinical care.