HGVS notation explained as a change on a named sequence, not as an rsID.

HGVS notation explained in practice: it names a change on a stated sequence. rsIDs name a dbSNP cluster, so you often need both, and a chip export usually gives you only the rsID.

HGVS notation explained as a sequence change

HGVS notation names a change on a stated sequence. The accession before the colon is part of the name. NM_000507.4:c.845G>A is not the same object as a bare c.845G>A copied from a forum.

An rsID names a dbSNP cluster at a site. HGVS names a specific edit on a specific molecule. One rsID can map to several HGVS strings across transcripts. One protein string can map to several DNA edits.

The official language lives at the HGVS Sequence Variant Nomenclature. genome.sh accepts HGVS strings in the same query command as rsIDs. If the local database cannot resolve the string, try the rsID or chr:pos:ref:alt.

  • c.: coding DNA change on a transcript accession
  • p.: predicted protein change, often ambiguous without c.
  • g.: genomic change on a named assembly sequence

c., p., g., and why the accession is not optional

A coding change looks like NM_000410.4:c.845G>A for HFE C282Y. A protein change looks like p.Cys282Tyr. The protein string is easier to remember and worse as a lookup key.

Several DNA changes can produce the same protein string. Several transcripts can number the same codon differently. If you drop the accession, you have a nickname, not a coordinate.

Genomic HGVS strings are valid. In genome.sh, chr:pos:ref:alt is the reliable genomic form. Name the assembly next to it, because GRCh37 and GRCh38 reuse numbers for different bases.

Transcripts, BRCA1, and ambiguous protein strings

BRCA1 has many transcripts in public databases. A gene search returns a catalog of variants. An HGVS string aims at one change on one accession.

ClinVar pages show the expressions they accepted. If your c. string does not match, you likely picked the wrong transcript, an older HGVS version, or a different nucleotide numbering.

Do not invent a c. string from a chip row. Consumer exports usually give rsIDs and genotypes, not HGVS. Translate through dbSNP or ClinVar.

Query HGVS, rsIDs, and coordinates with genome.sh

Install genome-sh, then a database tier. The query language accepts rsIDs, gene symbols, HGVS, and coordinates. Human output is for one site. JSON is for scripts.

HFE is a useful worked example because both the rsID and the HGVS are famous. rs1800562 is C282Y. rs1799945 is H63D, also written NM_000410.4:c.187C>G. Prefer the rsID when you have one.

The HTTP API accepts the same public identifiers and no genome files. There is no API key.

cargo install genome-sh
genome db install lite
genome query rs1800562
genome query chr6:26090951
genome query NM_000410.4:c.187C>G --format json

Consumer chips hide HGVS behind rsIDs

A 23andMe-style export is a table of rsIDs, positions, and letters. It is not a VCF of HGVS strings. If you need HGVS, look up the rsID in dbSNP or ClinVar, then keep the accession they used.

Internal 23andMe i-ids may never have an rs number. Those probes are not HGVS either. If the site is missing from the file, write not called. Missing is not the reference allele.

Most pathogenic BRCA1 variants will not be on the chip at all. An HGVS for a frameshift you cannot see is not a negative test.

  • Chip row: rsid, chromosome, position, genotype
  • Public lookup: rsID to ClinVar / dbSNP HGVS
  • Not called: probe absent or no-call, not 0/0

Assemblies, numbering, and silent mismatches

rsIDs are supposed to follow a site across assemblies. Bare chr:pos values are not. Mixing GRCh37 and GRCh38 is the usual silent error.

Protein numbering can disagree with DNA numbering when the initiator methionine is counted. Sickle cell is a classic example: people say HBB Glu6Val; some HGVS strings use p.Glu7Val. Confirm on the ClinVar or dbSNP page instead of memorizing a blog table.

If genome query cannot resolve an HGVS string, do not assume the variant is absent from ClinVar. Try the rsID, then chr:pos:ref:alt with the assembly named.

genome query rs1800562 --format json | jq .
curl -s https://api.genome.sh/v1/query/rs1800562 | jq .
genome query chr6:26090951:C:G --format json

When to prefer the rsID instead of HGVS

Use the rsID when dbSNP has one. It is the stable key across tools, chips, and assemblies. Use HGVS when the site is novel or when you must name a transcript-specific change.

GWAS papers, 23andMe files, and most genome.sh examples lead with rsIDs. Clinical reports often lead with HGVS plus an accession. You will translate both directions.

genome.sh is identifier lookup and local annotation. It is not a full HGVS mutalyzer. For hard nomenclature questions, read varnomen and the ClinVar page for that variation id.

Worked HFE strings you can type today

HFE is the usual HGVS tutorial because both the rsID and the protein nickname are famous. rs1800562 is C282Y. rs1799945 is H63D. genome query on the rsID is enough for lookup.

If you insist on HGVS, keep the accession. NM_000410.4:c.187C>G is H63D. A bare p.His63Asp is a nickname that several DNA edits could share. ClinVar will show the expressions it accepted.

Then match the rsID to a local file. If the chip row is missing, write not called. An elegant coding string does not invent a genotype.

genome query rs1800562 --format json
genome query rs1799945 --format json
curl -s https://api.genome.sh/v1/query/rs1800562 | jq .

Questions

What is HGVS notation?

A standard way to name a sequence change on a stated accession. HGVS notation explained simply: coding (c) is DNA on a transcript, protein (p) is the amino-acid change, genomic (g) is the assembly sequence, and the accession is required for an unambiguous site.

Why does my HGVS not match ClinVar?

Wrong transcript, wrong assembly, or an older HGVS version. ClinVar pages show the expressions they accepted. Prefer the rsID when you have one.

Is p. notation enough to look up a variant?

Often not. Several DNA changes can produce the same protein string. Prefer c. plus accession, or the rsID.

Can I query genomic g. HGVS in genome.sh?

Coordinates as chr:pos:ref:alt are the reliable genomic form. Name the assembly. Example: genome query chr6:26090951:C:G.

How do I convert an rsID to HGVS?

Look up the rsID in dbSNP or ClinVar, or run genome query rs1800562 --format json. Do not invent a coding HGVS string from a chip genotype row. If the local database cannot resolve a coding string, try the rsID or chr:pos:ref:alt next.

Does a 23andMe file contain HGVS?

Usually no. It contains rsIDs and genotypes. Translate through public databases. A missing rsID is not called, not wild type. Internal 23andMe i-ids may never have HGVS at all.

Should I search BRCA1 by gene or by HGVS?

Gene search returns a catalog. HGVS aims at one change on one transcript. A consumer chip is not a BRCA test either way.

Is HGVS notation medical advice?

No. It is a naming system. genome.sh prints public annotations for identifiers it can resolve.

genome.sh reports public annotations. It is informational software, not a medical device or a substitute for clinical care.