A public name for a spelling difference in DNA
DNA is a long text written in four letters: A, C, G, and T. Most of that text is the same from person to person. A SNP is a place where one letter commonly differs. An rsID is the public catalog number for that place.
rs429358 does not mean “Alzheimer’s.” It means “this site in APOE.” The two letters you carry there, if a test called them, are a genotype. A disease is a clinical conclusion. Those three facts are easy to mash into one word. They are not the same fact.
That is why people search “what is an rsID” after they download a 23andMe file. The file is full of those numbers. The number is a handle. It is not a verdict.
- DNA: the sequence of bases
- SNP: one common one-letter difference
- rsID: the public name of that site in dbSNP
- Genotype: the two letters a test called, if it called them
What is an rsID in dbSNP
An rsID (Reference SNP cluster ID) is the public name dbSNP assigns when overlapping SNP reports are merged into one cluster. rs1799945, rs429358, rs1800562, and rs334 are cluster names, not people.
Researchers submit variants to dbSNP. Overlapping reports on the same site get one rs number. 23andMe files, GWAS papers, and ClinVar records use that number as a shared key.
The number does not encode clinical significance. It does not encode frequency. Those facts live in other databases that point at the same id.
Cluster IDs, merges, and withdrawn numbers
Clusters can be merged. One rs number becomes an alias of another. Clusters can be withdrawn when the evidence for a unique site collapses. Prefer a current dbSNP snapshot.
That is why a five-year-old blog table can name an rsID that now redirects. genome.sh queries the snapshot you installed. genome db stats and GET /v1/sources tell you which one.
Do not treat an rsID as a permanent primary key in a clinical system without recording the dbSNP build. For personal notes, the rsID is still the best public handle you have.
rsID versus genotype versus disease
rs429358 names a site in APOE. Your file names the two bases you carry there, for example T/C. The public databases describe the site. Only the local file describes you.
A chip that does not list the rsID did not call it. Missing is not the same as homozygous reference. Write not called. Do not invent a wild-type genotype.
ClinVar may attach a submitted significance to the same rsID. That is still a statement about the variant in some clinical context, not a diagnosis of the person who grepped the number.
- rsID: public name of a site
- Genotype: the alleles called in a file, if any
- ClinVar significance: submitted interpretation of the variant
- Disease: a clinical conclusion genome.sh will not make
rsID versus HGVS versus chr:pos
HGVS names a change on a stated sequence, for example NM_000410.4:c.187C>G or p.His63Asp. The accession before the colon is not optional if you want an unambiguous coding site.
chr:pos is assembly-specific. rs1800562 has different numeric positions on GRCh37 and GRCh38. If you mix those, you look at the wrong base. Prefer the rsID when it exists.
Novel VCF alleles may have no rsID until someone submits them to dbSNP. Query those as chr:pos:ref:alt. A dot in the VCF ID column means no id was stored, not that dbSNP has no record.
Consumer chips, i-ids, and missing sites
A typical 23andMe or AncestryDNA export has on the order of 600,000 rows. Each row is an rsID or a vendor probe id, a chromosome, a position, and a genotype string.
Some chip probes have only an internal id (i followed by digits). Those may not exist in ClinVar or gnomAD under an rsID. Do not force them into dbSNP.
Most ClinVar pathogenic variants are not on the chip. Absence of BRCA1 founder rsIDs from an export is not a negative BRCA1 result. A genotyping chip is not whole-genome sequencing.
How genome.sh resolves an rsID
Look up the rsID in dbSNP for coordinates and alleles, in ClinVar for submitted clinical significance, and in gnomAD for frequency. genome.sh joins those sources in one query.
CLI: cargo install genome-sh, genome db install lite or standard, then genome query rs429358. HTTP: GET https://api.genome.sh/v1/query/rs429358 with no key. Files stay off that API.
Formats are human, json, and compact. Scripts should use JSON. After a local database install, identifier queries do not need the network.
cargo install genome-sh genome db install lite genome query rs429358 genome query rs1799945 --format json | jq . curl -s https://api.genome.sh/v1/query/rs334 | jq .
Worked rsIDs people actually type
rs1799945 is HFE p.His63Asp (H63D). rs1800562 is HFE p.Cys282Tyr (C282Y). Both are common in consumer files. Query them as public ids, then see whether your export called them.
rs429358 and rs7412 together define common APOE haplotypes. Querying only rs429358 is incomplete for haplotype calling. APOE status is sensitive; keep genotypes on disk.
rs334 sits in HBB and is the usual sickle cell tutorial site. It is a good CLI test. It is not a hemoglobinopathy assay.
genome query rs1799945 genome query rs1800562 genome query rs429358 genome query rs7412 genome query rs334
What an rsID cannot tell you
It cannot tell you zygosity unless a local file called the site. It cannot tell you penetrance for a person. It cannot replace a clinical lab report.
It cannot survive a wrong assembly if you throw the number away and keep only chr:pos. It cannot make a missing chip probe into a reference call.
genome.sh prints public annotations for the id. Clinical care needs a clinician and the right assay. Informational software is not a medical device.
- No genotype without a local call
- No diagnosis from a cluster number
- No wild-type inference from a missing chip row
- No VCF or 23andMe upload to the HTTP API
Questions
What is an rsID?
A Reference SNP cluster ID in dbSNP: a public name for a variant site, such as rs429358. It is not your genotype and not a diagnosis.
Does every SNP have an rsID?
No. Novel variants from a VCF may have no rsID until they are submitted to dbSNP. Query those by chromosome, position, ref, and alt.
Do rsIDs change?
Clusters can be merged or withdrawn. Prefer current dbSNP and ClinVar snapshots. genome db stats and GET /v1/sources show the build you are using.
What does rs429358 mean?
It names a site in APOE. Your file, if it called the site, names the alleles you carry. Haplotype calling also needs rs7412. Query the public record; keep genotypes local.
What is the difference between an rsID and a genotype?
The rsID names the site. The genotype is the pair of bases (or indel alleles) called in a sample file. Missing from the file means not called, not wild type.
rsID vs HGVS: which should I use?
Use the rsID when you have one. It is more stable across GRCh37 and GRCh38. Use HGVS when you need a transcript-specific coding change and there is no rs number.
How do I look up what an rsID is?
genome query rs334, or GET https://api.genome.sh/v1/query/rs334, or open the dbSNP page. That returns public annotation, not a personal result.
Is an rsID enough for a medical decision?
No. Clinical care needs a clinician, the right assay, and the rest of the evidence. genome.sh reports public annotations.