Install genome.sh from crates.io, Bioconda, or GitHub
Install genome.sh with Rust, with Conda, or from a GitHub release binary. The crate name is genome-sh because genome was taken. After install, the command you type is genome.
cargo install genome-sh is the path if you already have a Rust toolchain. conda install -c bioconda genome-sh is the path in bioinformatics environments. Pre-built binaries live on the CLI repository releases.
You do not need an API key for ordinary query and annotate. The optional AlphaGenome predict command is the exception: it is opt-in, key-gated, and not part of install.
- Crate: genome-sh on crates.io
- Binary: genome
- License: MIT
cargo install genome-sh # or conda install -c bioconda genome-sh genome --help
The crate is genome-sh, the binary is genome
Search logs say install genome and then fail to find a crate of that name. Search for genome-sh. Run genome. That split is documented on the homepage on purpose.
genome --help lists query, annotate, compare, extract, predict, db, and config. query is identifier lookup and annotate is VCF streaming. db installs the snapshot. predict is the opt-in model call.
config sets default format, reference genome, and an optional AlphaGenome key. Leave the key unset until you mean to call predict. Ordinary install and query do not need it.
genome --help genome query --help genome db --help genome config --help
Install a database tier: lite, standard, or full
Identifier queries need a local annotation database. genome db install downloads it. After that, query does not need the network.
lite is the ClinVar-oriented start. standard adds population frequencies. full is the complete local set, including the large gnomAD and dbSNP tables. Pick the smallest tier that answers your questions.
genome db stats reports what you installed and how fresh it is. genome db status is the short health view used in the agent guide. Re-run install or the documented update path when you need a newer ClinVar release.
Disk space is the usual surprise. full is large because gnomAD and dbSNP are large. Start with lite, run genome query rs334, then upgrade the tier when a JSON field you need is missing.
- lite: ClinVar-oriented, smallest start
- standard: ClinVar plus gnomAD frequencies, usual personal-report tier
- full: complete local set, largest download
genome db install lite genome db stats # later, if you need frequencies: genome db install standard genome db stats
First query and first annotate
If genome query rs1799945 prints a record, install worked. rs1799945 is HFE H63D, a public id. You did not upload a genome.
Human output is the default. --format json is for scripts and agents. --format compact is the short form. Pipe JSON into jq.
Annotation is a separate verb. Point it at a file on disk. The HTTP API will not accept that file.
If annotate errors on a missing file, you pointed at a path that is not on this machine. The CLI will not fetch a VCF from a URL as a substitute for a local path. That is intentional.
genome query rs1799945 genome query rs1799945 --format json | jq . genome query rs334 --format compact genome annotate sample.vcf.gz --format json genome annotate sample.vcf.gz --filter clinical --format json
Formats, filters, and the commands you will actually type
query accepts rsIDs, gene symbols, HGVS, and coordinates. annotate streams a VCF. compare diffs two genomes, and extract pulls variants from CRAM or BAM. All of those stay local after db install.
--filter clinical keeps ClinVar-relevant rows when you annotate. Use it to shrink JSON, not to invent a diagnosis.
If you are building a pipeline, install once in CI and cache the database directory. Then run query or annotate as a unit step. Do not download full on every job if lite already answers the test.
genome query BRCA1 --format json | jq 'length' genome query chr6:26090951 --format json genome extract sample.cram --format json genome compare a.vcf.gz b.vcf.gz --format json
Optional HTTP API after install
You do not need https://api.genome.sh after a local install. Local query is the default. The API is for identifier lookup on a host that must not hold SQLite or genomes.
There is no API key. There is no VCF POST. GET /v1/query/rs1799945 is the hosted form of genome query.
Use the API from CI smoke tests. Use the CLI for files. Mixing those jobs is how people leak exports into HTTP logs.
If you only needed to try one rsID, the website /query page uses the same identifier API. Install is for files, bulk, and offline use. Do not skip db install and then wonder why query has no local hits.
curl -s https://api.genome.sh/v1/query/rs1799945 | jq . curl -s https://api.genome.sh/v1/sources | jq . curl -s https://api.genome.sh/v1/health | jq .
Update the snapshot and inspect stats
A local index is not live NCBI. When you need a newer ClinVar or gnomAD release, update the database with the db commands in --help and check stats again.
Cite the snapshot in any report you generate. GET /v1/sources on the public API tells you what the server is running, which may differ from your laptop.
If stats look empty, install did not finish. Re-run genome db install for the tier you wanted. Do not assume lite contains gnomAD.
genome db stats genome db status curl -s https://api.genome.sh/v1/stats | jq .
Common install failures
cargo: command not found means you need a Rust toolchain, or you should use conda or a release binary instead. conda: package not found usually means the bioconda channel is missing.
genome: command not found after cargo install means ~/.cargo/bin is not on PATH. Add it, or hash -r in bash, and try genome --help again.
Queries that return no gnomAD fields on a lite install are expected. Install standard or full when you need frequencies. Queries that fail on a VCF path are annotate, not query.
Questions
How do I install genome.sh?
cargo install genome-sh, or conda install -c bioconda genome-sh, or a GitHub release binary. Then genome db install lite, standard, or full.
Why is the crate not named genome?
The crate is genome-sh. The binary is genome. That split is documented on the homepage.
How large is the genome.sh database?
Depends on the tier. lite is the small ClinVar-oriented start. full is much larger because of gnomAD and dbSNP. genome db stats reports what you installed. If you only needed ClinVar, stay on lite.
Do I need the HTTP API after installing?
No. Local query is the default. The API is for identifier lookup without a local DB. It still will not accept a VCF.
Do I need an API key to install genome.sh?
No. Query and annotate are local. Only the optional AlphaGenome predict command needs a key and explicit consent.
What is the difference between lite, standard, and full?
lite is ClinVar-oriented. standard adds population frequencies. full is the complete local set. Pick the smallest tier that answers your queries. You can install a larger tier later without reinstalling the binary.
How do I update ClinVar after install?
Use the genome db update path in --help, then genome db stats. A local snapshot is not live NCBI.
Can I install genome.sh without Rust?
Yes. conda install -c bioconda genome-sh, or download a pre-built binary from GitHub releases.