Query ClinVar command line lookups against a local index, not a scraped HTML page.

You can query ClinVar command line indexes without scraping the website. A local ClinVar snapshot is faster, scriptable, and needs no web session or API key.

Query ClinVar command line lookups without scraping NCBI

The ClinVar browser is the right place to read a full VCV record: submissions, conditions, and stars. It is the wrong place to look up 8,000 rsIDs from a VCF.

Copying fields out of HTML is brittle. Session cookies expire. Layout changes break scripts. If you want ClinVar in a pipeline, you need a programmatic path.

Two honest paths exist: NCBI E-utilities query live NCBI, while genome.sh queries a local ClinVar snapshot in SQLite. Pick live when you need today's XML. Pick local when you need speed, JSON, and a machine that still works on a plane.

E-utilities versus a local ClinVar index

NCBI E-utilities (esearch, esummary, efetch) are the official remote path. They use the same query language as the website. They also return XML, require pacing, and want an API key once you go above a few requests per second.

A local snapshot trades freshness-until-you-update for millisecond lookups. You still cite ClinVar. You just do not hit NCBI for every rsID in a chip file.

genome.sh is not a full dump of every VCV XML field. It is an index of the aggregate fields you actually filter on: identifier, gene, coordinates, significance, review status. When you need the submission table, open the ClinVar page for that variation id.

  • E-utilities: live NCBI, XML, rate limits, optional API key
  • genome.sh: local SQLite, JSON, no key, refresh on demand
  • ClinVar website: full VCV record, not a pipeline

Install a ClinVar-capable database and query it

Install the crate genome-sh; the binary is genome. Then install a database tier: lite is the ClinVar-oriented start, while standard and full add gnomAD and the rest.

After genome db install, genome query hits disk. Confirm what you installed with genome db status or genome db stats. The HTTP API is optional and still only accepts identifiers.

Human output is for one site in a terminal. JSON is for scripts. compact is the short form. Agent scripts should use --format json.

cargo install genome-sh
genome db install lite
genome db status
genome query rs80357906
genome query BRCA1 --format json | jq .

Query by rsID, gene, and coordinates

Use an rsID when you have one. rs334, rs1800562, and rs1799945 are the tutorial sites people actually type. Gene symbols return a catalog, not a personal report.

Coordinates work when the site has no rs number yet. Name the assembly in your own notes. chr17:43092919 on GRCh38 is not the same numeric position as the GRCh37 locus.

HGVS strings are accepted in the same command when the local database can resolve them. If a string misses, fall back to rsID or chr:pos:ref:alt rather than inventing a transcript.

  • genome query rs80357906: one ClinVar-heavy site
  • genome query BRCA1 --format json: variants for a symbol
  • genome query chr17:43092919: assembly-aware coordinates
  • genome query rs334 --format compact: short terminal form
genome query rs334
genome query rs1800562 --format json
genome query HBB --format json | jq .
genome query chr6:26090951

Read significance next to review status

Pathogenic is a submitted label. Review status describes how that label was reached: criteria provided, multiple submitters, no conflicts, and so on. Do not flatten both fields to a single emoji in a script.

Conflicts are common. Two labs can disagree because they used different evidence cutoffs, dates, or transcripts. genome.sh shows the public aggregate, not a new classification.

Absence from ClinVar is not evidence of benignity. Most genomic sites have no ClinVar record. A VUS is uncertain significance: do not treat it as a finding to act on from a consumer file.

Filter clinical rows in a VCF pipeline

Identifier query is for one name. File annotation is for every row. genome annotate streams a VCF against the same local ClinVar index. The file never needs to leave disk.

--filter clinical keeps rows that carry a clinical annotation in the local database. --format json is the form you pipe into jq. The HTTP API will not accept the VCF.

A genotyping chip is not whole-genome sequencing. Most ClinVar pathogenic variants will not be in a 23andMe-style export. Missing is not called, not wild type.

genome db install standard
genome annotate sample.vcf.gz --format json
genome annotate sample.vcf.gz --filter clinical --format json

Refresh the snapshot when ClinVar moves

A local index is a dated snapshot. ClinVar releases change. Re-run the database install or update path when you need a newer build. genome db stats and GET /v1/sources report what is loaded.

Do not mix a 2022 ClinVar quote with a 2026 VCF and call it current. Record the source version next to any script output you keep.

If you need the live NCBI record for one variation id, open the ClinVar page. Use the CLI to find the id fast, then read the full submission table where it lives.

What a CLI ClinVar hit is not

It is not a diagnosis. It is not a complete extract of the VCV XML. It is not live NCBI unless you also open NCBI.

It is not a BRCA test, a hemochromatosis test, or a sickle cell test. Those need clinical assays. genome.sh prints public annotations.

AlphaGenome prediction is a separate opt-in CLI command with a key. It is not the default ClinVar query path.

Questions

How do I query ClinVar from the command line?

Install genome-sh, run genome db install lite, then genome query with an rsID or gene. That hits a local ClinVar index. Example: genome query rs334 --format json.

Query ClinVar CLI without E-utilities?

Yes. genome query uses SQLite on disk after genome db install. E-utilities query NCBI live and return XML. They are different tools.

Can I query ClinVar by gene?

Yes. genome query BRCA1 returns known variants for that symbol. Use --format json to filter significance in a script. A gene catalog is not a personal report.

How do I keep only pathogenic rows?

Output JSON and filter on the significance field with jq. Keep review status in the same object. For a VCF, genome annotate --filter clinical --format json.

Is this the same as NCBI E-utilities?

No. E-utilities query NCBI live. genome.sh queries a local snapshot. Refresh the database when you need a newer ClinVar release.

Does ClinVar include every disease variant?

No. Absence from ClinVar is not evidence of benignity. Most genomic sites have no ClinVar record.

Can I query ClinVar over HTTP instead of the CLI?

Yes for public identifiers: GET https://api.genome.sh/v1/query/rs334. No key. The API still rejects VCF, BAM, and 23andMe files.

Is genome.sh a medical device?

No. It prints public ClinVar annotations. Clinical decisions need a clinician and an appropriate assay.

genome.sh reports public annotations. It is informational software, not a medical device or a substitute for clinical care.