A local myvariant.info alternative for identifier lookup that still refuses a VCF upload.

A myvariant.info alternative for identifier lookup can be hosted or local. genome.sh is a local CLI with an optional open API, and it will not take a VCF upload.

A myvariant.info alternative can be local SQLite

myvariant.info is a hosted aggregator. One getvariant call can return many annotation collections for a public id. That is useful in a notebook that already lives on a laptop and queries rsIDs.

genome.sh is a different shape: a Rust CLI, a local SQLite snapshot, jq-friendly JSON, and an HTTP API that still refuses genome uploads. You trade some source breadth for control of the snapshot and for a file-safe workflow.

Neither product should see a raw genome if you only needed an rsID. If a hosted API also offers VCF upload, treat that path as a different product with a different threat model.

  • myvariant.info: hosted BioThings aggregator, many sources per id
  • genome.sh CLI: local ClinVar, gnomAD, dbSNP, AlphaMissense, ClinGen, PharmGKB, UniProt
  • genome.sh API: public identifiers only, no key, no VCF

What myvariant.info is good at

Breadth. If you need a source genome.sh does not ship, the aggregator may still have it. Field names are its own schema. Python and REST clients already exist.

Interactive notebooks. A few rsIDs over HTTPS is the job it was built for. Rate limits still apply. A hundred thousand ids will throttle.

It remains a useful tool. genome.sh does not wrap it. There is no proxy. If you need both, call both and map JSON explicitly.

Batch endpoints on a hosted aggregator are still remote. They still see your identifier list. That is usually fine for public rsIDs. It is not fine if you smuggle genotypes into the query string.

What genome.sh changes: CLI, snapshot, no VCF upload

After genome db install, queries hit disk. They still run on a plane. The snapshot is yours to update. You cite that snapshot, not a moving hosted blob you do not control.

Bulk VCF annotation is a local command. genome annotate file.vcf.gz --format json streams records. The HTTP API will reject that file. That split is the product.

Identifier lookup can still be remote: GET https://api.genome.sh/v1/query/rs1799945 with no key. Use it from CI when you do not want SQLite on that host.

cargo install genome-sh
genome db install standard
genome query rs1799945 --format json | jq .
curl -s https://api.genome.sh/v1/query/rs1799945 | jq .
curl -s https://api.genome.sh/v1/sources | jq .

Field names are not a drop-in replacement

Do not point a myvariant.info parser at genome.sh JSON and hope. ClinVar significance, review status, and gnomAD AF sit under different keys. Map them once in your script.

genome.sh favors a small set of sources it indexes well. myvariant.info favors many collections, some of which will be empty for a given id.

Check GET /v1/sources and genome db stats so you know which snapshot you are comparing. Mixing gnomAD versions across tools is a silent error.

Write a 20-line mapper once. Test it on rs1799945, rs334, and BRCA1. If a field is missing on one tool, keep a null rather than copying a value from the other tool.

  • Map ClinVar significance and review status explicitly
  • Map gnomAD AC, AN, AF explicitly
  • Do not assume nested BioThings paths exist on genome.sh

Bulk lookup belongs on disk

For 100,000 rsIDs, a local index wins. Remote aggregators throttle. Loops of HTTP calls are how people get blocked and how they leak files into logs.

genome annotate is the bulk path for a VCF. For a list of ids, genome query still runs locally after db install. JSON pipes into jq, Python, or an agent.

If you only have a consumer export, parse it on disk or in the browser importer. Do not POST it to myvariant.info or to api.genome.sh. The latter will refuse. The former is still an upload.

genome annotate sample.vcf.gz --format json
genome annotate sample.vcf.gz --filter clinical --format json
genome query rs1799945 rs1800562 rs334 --format json | jq .

Call the open identifier API when you have no local DB

Servers that must not hold genomes can still resolve public ids. That is the job of https://api.genome.sh. No account. No key.

Endpoints include /v1/query/:query, /v1/gene/:gene, /v1/gnomad/:rsid, /v1/sources, /v1/stats, /v1/health. Query values are rsIDs, genes, HGVS, and coordinates.

Use it for interactive lookup and small scripts. Use the CLI for files. That is the same split as the rest of genome.sh.

curl -s https://api.genome.sh/v1/query/rs1799945 | jq .
curl -s https://api.genome.sh/v1/gene/BRCA1 | jq '.[0:3]'
curl -s https://api.genome.sh/v1/stats | jq .

Use both tools when a source exists on only one

If a source is only on myvariant.info, use it for that identifier. If you are annotating a VCF, stay local on genome.sh.

Do not send a user file to either HTTP API as a shortcut. Identifier in, JSON out. Files stay on disk.

genome.sh is MIT licensed. The crate is genome-sh. The binary is genome. Install, query, annotate, and keep the aggregator for the long tail of sources.

Privacy, logs, and the file you must not paste

Identifier APIs log URLs. An rsID in a GET path is a public fact. A VCF attached to a ticket, a notebook, or a Discord paste is a genome leak. genome.sh refuses the file on the API so you have to try another product to make that mistake.

myvariant.info is a useful aggregator. Use it for public ids. Do not treat it as a VCF annotator. genome annotate file.vcf.gz --format json is the local annotator in this project.

If you already have a myvariant.info notebook, keep it. Add genome query --format json for the sources you want under a snapshot you control, then map fields. Do not dual-write a user file to both HTTP APIs.

genome query rs1799945 --format json | jq 'keys'
curl -s https://api.genome.sh/v1/query/rs1799945 | jq 'keys'
genome annotate sample.vcf.gz --format json | jq '.[0] | keys'

Questions

What is a good myvariant.info alternative?

For local, file-safe identifier lookup, genome.sh. It ships its own indexed sources in SQLite and an open API that rejects VCFs. It does not wrap myvariant.info.

Does genome.sh wrap myvariant.info?

No. It indexes ClinVar, gnomAD, dbSNP, AlphaMissense, ClinGen, PharmGKB, and UniProt itself.

Is the genome.sh API a drop-in replacement for myvariant.info?

No. Field names differ. Map JSON explicitly. Query URLs differ too.

Which is faster for 100,000 rsIDs?

A local index. Remote APIs throttle. genome db install, then query or annotate offline. A loop of getvariant or curl calls is the slow path and the log-leak path.

Can I POST a VCF to genome.sh like some hosted annotators?

No. Annotate with the local CLI or the in-browser importer. The API accepts only public identifiers. That is the same privacy split whether you compare genome.sh with myvariant.info or with any hosted annotator.

Do I need an API key?

Not for genome.sh identifier lookup. myvariant.info has its own terms; read them there.

How do I see which sources genome.sh ships?

genome db stats locally, or GET https://api.genome.sh/v1/sources.

Should I delete myvariant.info from a notebook?

No. Keep it for sources genome.sh does not ship. Use genome.sh when you want a local snapshot or a VCF that must not upload.

genome.sh reports public annotations. It is informational software, not a medical device or a substitute for clinical care.