You already have a table, not a mystery
How to read your DNA, in the sense people actually mean after a spit kit, is not how to become a clinician. It is how to look at the download without letting the company's story do all the work. The download is a spreadsheet. The story was software.
A typical 23andMe or AncestryDNA export has on the order of 600,000 rows. Comment lines at the top explain the genome build, usually GRCh37. Data lines are the measurements: a probe, a place, two letters. That is the object. Everything else is interpretation layered on top.
You do not need to upload the file again to understand a row. The rsID is a public name. The two letters are the private fact. Keep them on different paths. Literacy is knowing which is which, and knowing what the chip never measured.
What one raw-data row actually is
A row is a site the chip was designed to call. In most consumer exports it looks like this: an rsID such as rs429358, a chromosome, a position, and a genotype of two characters. Some probes have a vendor id, an i followed by digits, instead of an rs number. Those internal ids may not exist in ClinVar under an rsID. Skip them rather than inventing a dbSNP name.
The two characters are DNA letters, A, T, C, or G, or a no-call such as --. DNA letters pair A with T and C with G in the double helix. The genotype is not a base pair in that sense. It is the two alleles at that site, one from each parent, as the chip reported them on a stated strand and assembly.
Position without assembly is a trap. The same rsID has different numeric coordinates on GRCh37 and GRCh38. Prefer the rsID when you have one. It is the stable public handle. A raw chromosome-and-position pair is only as good as the build written in the header.
- rsID: public name of the site, when the probe has one
- Chromosome and position: usually GRCh37 in consumer files
- Genotype: two letters, or a no-call
- Vendor i-id: a chip probe that may have no rs number
Two letters, in plain language
Homozygous means both alleles match. AA, GG, TT, or CC at a site: the chip thinks you carry the same letter from both parents. In everyday language you are matching at that spelling. It does not mean healthy. It does not mean diseased. It means same-same.
Heterozygous means the two letters differ. AG, CT, and the other mixed pairs: one parent-sized copy of each spelling. Again this is not a moral category. Recessive conditions often need two copies of a disease allele. Dominant conditions can follow a single copy. The words homozygous and heterozygous do not tell you which world you are in until you know the gene and the evidence.
People also write 0, 1, or 2 copies of a particular letter, or they write A/G with a slash. Those are the same fact in different clothes. If a blog decoder disagrees with the letters in your file, check strand and build before you trust the blog. The file is the measurement. The decoder is an opinion.
- Homozygous: both letters the same, for example GG
- Heterozygous: two different letters, for example AG
- No-call: --, 00, or a blank the vendor uses for not measured
- Not in the file: also not called, not a secret healthy result
Missing means not called
This is the rule that saves people from inventing a clean bill of health. If the row is absent, the chip did not measure the site. If the genotype is -- or 00, the chip tried and did not assign letters. Both cases are not called. Neither case is wild type.
A consumer chip samples hundreds of thousands of SNPs, not three billion bases. Most of BRCA1 is not on it. Most ClinVar pathogenic variants are not on it. A 23andMe file is not a BRCA test. Writing normal in the gap is a story the assay cannot support.
The same rule applies to famous tutorial sites. rs334, rs1800562, and rs429358 are often on arrays, and still not always. Often-on-the-chip is not always-on-the-chip. If you cannot find the row, stop. Do not paste a reference letter from a website and call it yours.
What the two letters are not
They are not a diagnosis. Disease is a clinical conclusion that needs a person, a phenotype, and usually a different assay. They are not an ethnicity. Ancestry estimates are models built on thousands of SNPs at once, not on one entertaining row. They are not a drug dose. Pharmacogenetic guidelines are written for specific variants and clinical contexts, not for a screenshot.
They are not the rest of the gene. A SNP is one bookmark. Sequencing a gene reads the chapter. An exome reads most chapters that code for protein, about 20,000 of them. A genome tries to read the book. Your export is the bookmark list.
They are not public just because the rsID is public. Sending rs429358 to a database is sending a word that is already on the internet. Sending CT next to it is sending a fact about a person. Keep that split in your hands.
Look the public record up without uploading the file
Pick an rsID from a row you actually have. Ask dbSNP where the site is and which alleles it knows. Ask ClinVar whether labs have submitted an interpretation, and what the review status is. Ask gnomAD how often the alternate letter appears in sequenced people. Those are public questions. They do not require your genotype.
Then glance back at your file, on the same machine, and see whether the site was called. If it was, you now have two objects: a public annotation and a private pair of letters. If it was not, you have only the first object. Do not glue them together in a cloud form that wants the whole export in order to explain one name.
Promethease-style sites and random upload boxes are a different product. They may be convenient. They are also a second copy of an identifying file. You already have the table. The public databases already have the rsID. The upload is optional, and it is a privacy decision, not a literacy requirement.
Three rows people actually look up
rs1800562 is HFE C282Y. If your file called it, you have two letters at a site that is discussed in hereditary hemochromatosis. You do not have a ferritin. You do not have a diagnosis. Zygosity and blood tests still belong to a clinic.
rs429358 is one of two SNPs that tag common APOE haplotypes. The other is rs7412. If either row is missing, you cannot call the common haplotype honestly. Association with late-onset Alzheimer disease risk is a population finding. It is not a forecast. The status is sensitive. It is a poor thing to paste into a group chat.
rs334 is the HBB sickle site that tutorials love because the public record is famous. A homozygous disease genotype and a heterozygous trait genotype are different clinical objects, and a chip is still not a hemoglobin electrophoresis. Use these rows as practice for the method: name, letters, public record, stop. Not as a home clinic.
One local lookup, then stop before a diagnosis
When you want the public record for a named rsID without sending the file, identifier lookup can stay on this machine. genome.sh is a late tool for that job. It prints what ClinVar, dbSNP, and related sources say about the id. It does not diagnose. It does not fill missing chip rows with wild type.
The importer at /import parses a consumer file in the browser tab so the export does not have to travel. That is the matching half: your letters stay local, the annotations are public. A genotyping chip remains a genotyping chip either way.
This page is not medical advice. If a row worries you, a clinician and the right assay are the next step, not a more dramatic decoder. Read the table. Look up the name. Leave the gaps as gaps.
cargo install genome-sh genome db install lite genome query rs429358
Questions
How do I read my DNA raw data?
Treat the file as a table. Each row is an rsID and two letters, or a no-call. Look the rsID up in public databases. Do not treat missing rows as normal, and do not treat a chip as a whole genome.
What does homozygous mean in a DNA file?
Both letters at that site match, for example GG. It means same-same, not healthy and not diseased. Meaning still depends on the gene and the evidence.
What does heterozygous mean in a DNA file?
The two letters differ, for example AG. You carry both spellings at that site. It is not automatically a carrier result and not automatically a dominant result.
What does -- mean in 23andMe raw data?
A no-call. The chip did not assign a genotype. Missing rows are also not called. Neither is wild type.
How do I look up an rsID from my file without uploading it?
Copy the rsID only. Open dbSNP, ClinVar, or gnomAD, or run a local identifier query such as genome query rs429358. Keep the two letters on your machine.
Why is a gene missing from my raw DNA file?
Consumer chips sample a designed list of SNPs, not every base of every gene. Absence is a sampling gap. A 23andMe file is not a BRCA test and not an exome.
Can I diagnose myself from raw DNA data?
No. A row of letters is not a diagnosis. Clinical care needs a clinician and an assay that can see the relevant variants. This explainer is not medical advice.
Is reading my DNA file the same as whole-genome sequencing?
No. The export is hundreds of thousands of predetermined SNPs. Whole-genome sequencing attempts to read nearly all three billion bases. They are different tests.