Annotate a VCF locally instead of uploading it
A VCF lists variants. Annotation attaches public knowledge to each row: gene, rsID, ClinVar significance, gnomAD frequency, predicted missense scores. That join can happen on the same disk that holds the VCF.
Hosted annotators ask you to send the file. genome.sh does not. The HTTP API at https://api.genome.sh accepts identifiers only. genome annotate reads a path on the machine that runs the binary.
The in-browser importer at /import parses supported files in the tab. It is the web form of the same rule. Neither path is a place to park a genome in someone else's bucket.
What annotation attaches to each row
For known sites, the local database can add dbSNP ids, ClinVar significance and review status, and gnomAD AC/AN/AF. AlphaMissense scores appear when that table is installed.
That is lookup, not a full transcript consequence pipeline. VEP and SnpEff compute HGVS and effects against a reference FASTA and a transcript set. genome.sh is the jq-shaped join onto public annotation tables.
Rows with no hit stay unannotated. That is not a benign classification. Novel alleles, structural variants, and sites absent from ClinVar will often come back empty.
- dbSNP: rsID when the site is clustered
- ClinVar: significance and review status
- gnomAD: allele frequency when the tier includes it
- AlphaMissense: missense scores when the tier includes them
Install a database and run genome annotate
Install the crate genome-sh; the binary is genome. Then install a database tier that matches the annotations you need: lite is ClinVar-oriented, while standard or full add gnomAD frequencies.
Point genome annotate at a .vcf or .vcf.gz. JSON is the form for scripts and agents. human is for a quick look. compact is the short form.
Confirm the snapshot with genome db status. Record that version next to any output you keep. Annotation is only as current as the SQLite you installed.
cargo install genome-sh genome db install standard genome db status genome annotate sample.vcf.gz --format json
--filter clinical, JSON, and compact output
A whole-genome VCF is noisy. --filter clinical keeps rows that carry a clinical annotation in the local database. Use it when you want the ClinVar-touched subset, not every SNP.
Agent scripts should use --format json. Pipe into jq. Do not drop review status when you filter on the word Pathogenic. Conflicts are data, not a bug.
Identifier query is still the right tool for one rsID. genome query rs1800562 --format json does not need a VCF. Annotate when you have many rows. Query when you have a name.
genome annotate sample.vcf.gz --filter clinical --format json genome annotate sample.vcf.gz --format compact genome query rs1800562 --format json | jq .
GRCh37 versus GRCh38 in the header
Read the VCF header. contig names and the reference line tell you the assembly. hs37d5 and GRCh37 are not GRCh38. Mixing them silently maps variants to the wrong place.
Prefer rsIDs in the ID column when they are present. They travel across assemblies better than raw POS. If ID is a dot, you have only chr:pos:ref:alt.
gnomAD v2 is GRCh37. gnomAD v4 is GRCh38. Match the frequency table to the file. genome.sh reports the fields in its snapshot; check /v1/sources or genome db stats.
genome.sh versus VEP and SnpEff
VEP, SnpEff, and vcfanno are the usual heavy tools. They need reference FASTA, cache files, and a consequence model. They are the right choice when you need transcript-level HGVS for every novel missense.
genome.sh looks up public annotations for known variants. Many workflows need both: VEP for consequences, genome annotate for ClinVar and gnomAD on the same machine.
You do not need a FASTA for ClinVar and gnomAD lookup of known sites. You do need a reference genome for calling, for left-alignment, and for tools that rewrite alleles.
- Use genome annotate: local ClinVar/gnomAD join, streaming, no upload
- Use VEP or SnpEff: transcript consequences, novel HGVS, FASTA cache
- Use both: known-site annotation plus consequence on a private disk
Consumer files, BAM, and CRAM
23andMe-style exports are not VCF. They are rsid/chromosome/position/genotype tables. The browser importer accepts several consumer formats. The CLI query path looks up rsIDs from those tables without sending the file.
BAM and CRAM are alignments, not variant tables. genome extract pulls variants from CRAM or BAM locally. Extraction still runs on disk. Do not upload a CRAM to the HTTP API; there is no endpoint for it.
A genotyping chip is not whole-genome sequencing. Annotating a chip-derived VCF will not magically cover the rest of ClinVar. Missing sites stay not called.
Keep genotypes on disk
If you hand the work to a coding agent, copy the report prompt and keep the VCF on the same machine as the agent. Copy the report template out of the repo before writing personal data.
Do not paste VCF rows into a chat so the model can call https://api.genome.sh/v1/query/{id}. Send the id if you must use HTTP. Keep GT columns local.
genome.sh is informational software, not a medical device. Annotation is not a diagnosis. Clinical testing belongs in a clinical lab.
Questions
How do I annotate a VCF locally?
cargo install genome-sh, genome db install standard, then genome annotate sample.vcf.gz --format json. The file stays on disk.
Can I annotate a VCF without uploading?
Yes. That is the default. The HTTP API rejects VCF, BAM, and 23andMe files. Use the CLI or the in-browser importer at /import.
Is genome annotate a replacement for VEP?
No. VEP computes transcript consequences against a reference. genome.sh looks up public annotations for known variants. Many workflows need both.
Does it need a FASTA?
Not for ClinVar and gnomAD lookup of known sites. You need a reference genome for calling or for tools that rewrite alleles.
How do I annotate a VCF with ClinVar and gnomAD?
Install standard or full so gnomAD is present, then genome annotate file.vcf.gz --format json. Add --filter clinical to keep ClinVar-touched rows.
What does genome annotate --filter clinical do?
It keeps rows that carry a clinical annotation in the local database. It is a filter, not a diagnosis. Review status still matters.
Can I annotate CRAM or BAM?
The CLI can extract variants from CRAM or BAM locally, then you annotate the variant table. Extraction still runs on disk. The HTTP API will not take the alignment file.
Can I annotate a VCF with the HTTP API?
No. GET /v1/query/{id} is identifier lookup only. There is no POST-a-VCF route.