genome CLI jq pipelines start with --format json and stay on disk.

genome CLI jq pipelines start with --format json. Human output is for one rsID in a terminal, and JSON is for filters, counts, notebooks, and agents.

genome CLI jq pipelines start with --format json

genome CLI jq usage is the product’s Unix shape. genome query --format json prints a record you can filter. jq is the obvious companion, not a requirement. Any JSON tool works.

Human output is the default and the wrong form to parse. compact is a short display. Scripts should pass --format json every time.

The pipeline is local after genome db install. Piping a local command does not upload a genome. Piping curl talks only to the identifier API.

  • human: one variant in a terminal, default
  • json: scripts, jq, notebooks, agents
  • compact: short display, still not a parser target
cargo install genome-sh
genome db install lite
genome query rs334 --format json | jq .
genome query rs334 --format json | jq 'keys'

Install once, then pipe forever

If jq is missing, install it from your package manager. genome.sh will still print JSON without it. You will just read less well.

lite is enough for ClinVar-only filters. standard or full is required when your jq path reads gnomAD fields. Check genome db stats before you debug an empty .gnomad.

Keep identifiers in the query. Keep files on disk with annotate. Do not jq a pasted 23andMe file in a chat log.

Put --format json in aliases and in agent prompts. A forgotten flag is how an agent starts regexing human text. The CLI will not save you from that if you asked for human.

genome db stats
genome query rs1799945 --format json | jq '{rsid, clinvar, gnomad}'
genome query rs1800562 --format json | jq '.clinvar'

Filter ClinVar significance without dropping review status

The usual bug is jq '.significance' and then a grep for Pathogenic. That drops review status, conflicts, and conditions. Keep a small object instead.

When you filter a gene catalog, preserve the id so you can open ClinVar later. A boolean mask with no rsID is how reports become un-auditable.

Absence of a field means the snapshot does not have it, not that the variant is benign. lite will not grow gnomAD keys because you wished hard.

Conflicts are a string, not a boolean. If you drop rows where significance is not Pathogenic, you also drop conflicting interpretations. Keep a conflicts field if the JSON has one.

genome query rs334 --format json | jq '{id: .rsid, significance, review_status}'
genome query BRCA1 --format json | jq '[.[] | {id: .rsid, significance, review_status}] | .[0:10]'
genome query BRCA1 --format json | jq '[.[] | select(.significance == "Pathogenic")] | length'

Gene catalogs, counts, and compact output

genome query BRCA1 --format json returns a list. jq length is the first sanity check. jq '.[0]' is the second. Then filter.

HBB is a smaller catalog and a better fixture than BRCA1 when you are learning the pipeline. rs334 is the single-site form of the same lesson.

compact is for humans watching a log. JSON is for the next process. Do not parse compact with cut.

genome query HBB --format json | jq 'length'
genome query HBB --format json | jq '[.[].rsid] | .[0:20]'
genome query rs334 --format compact
genome query BRCA1 --format json | jq '[.[] | .significance] | group_by(.) | map({sig: .[0], n: length})'

Annotate a VCF and jq the stream

genome annotate sample.vcf.gz --format json writes one JSON object per record, or an array, depending on the CLI version you installed. Read --help once and jq accordingly.

--filter clinical shrinks the stream to ClinVar-relevant rows. It is a convenience filter, not ACMG. Missing chip or VCF sites are still not called.

The HTTP API will reject this file. If your jq is running on curl output of a VCF POST, you are in the wrong product.

Large VCFs should stay gzipped. genome annotate reads the stream. jq can slurp too much if you force an array in memory. Prefer per-record JSON if --help says the CLI emits it, and test on a 20-row slice first.

genome annotate sample.vcf.gz --format json | jq 'length'
genome annotate sample.vcf.gz --filter clinical --format json | jq '[.[] | {id: .rsid, significance, review_status}]'
genome annotate sample.vcf.gz --format json | jq '[.[] | select(.gnomad.af != null and .gnomad.af < 0.001)] | length'

curl is the hosted form of the same idea

curl -s https://api.genome.sh/v1/query/rs334 | jq . is genome query over HTTPS for a public id. No key. No file.

GET /v1/gnomad/rs429358 is frequency-only. GET /v1/gene/BRCA1 is the catalog. GET /v1/sources tells you the server snapshot, which may differ from your laptop.

Rate limits exist. For 100,000 rsIDs, install a local database. A loop of curl | jq is how you get throttled.

curl -s https://api.genome.sh/v1/query/rs334 | jq .
curl -s https://api.genome.sh/v1/gnomad/rs429358 | jq .
curl -s 'https://api.genome.sh/v1/query/BRCA1?limit=5' | jq .
curl -s https://api.genome.sh/v1/sources | jq .

Agents that can run jq do not need screenshots

A coding agent that can run genome query --format json on your machine does not need a screenshot of the ClinVar website. Give it the agent guide. Forbid upload. Forbid pasting genotypes into the chat.

Ask the agent to keep significance and review status together, to write not called for missing chip sites, and to refuse medical advice. Those rules are in the report prompt.

jq is how the agent picks fields. genome.sh is how it gets them. Your disk is where the VCF stays.

  • Agent input: local file path, never a paste of the file
  • Agent command: genome query or genome annotate --format json
  • Agent filter: jq objects that keep review status

TSV, grep, and other companions

Need TSV? Use jq -r to emit tab-separated columns. Prefer JSON while ClinVar is nested. Flatten only at the edge.

grep still works on human output for a single rsID. It is a bad parser for catalogs. If you grep Pathogenic on human text, you will miss review status again.

The whole point of the genome CLI jq pairing is that the JSON is stable enough to script. Read --help when fields move between snapshots, and cite genome db stats in the output.

head, rg, and miller are fine companions once you flattened TSV. Until then, stay in jq objects. Nested ClinVar review text does not survive a naive cut -f2.

genome query rs334 --format json | jq -r '[.rsid, .significance, .review_status] | @tsv'
genome query HBB --format json | jq -r '.[] | [.rsid, .significance] | @tsv' | head

Questions

How do I pipe the genome CLI into jq?

genome query rs334 --format json | jq . After genome db install. Human output is not JSON. compact is not JSON either.

Is jq required for genome.sh?

No. It is the obvious companion. Python, R, and any JSON tool work. Scripts should still request --format json.

Can I get TSV from genome query?

Yes, via jq -r and @tsv, or other format flags in --help. Prefer JSON for nested ClinVar fields, then flatten at the edge.

Does piping genome into jq upload my DNA?

A local pipeline does not. curl | jq talks to the identifier API only. Never pipe a VCF into curl against api.genome.sh; the API rejects files.

How do I filter Pathogenic variants with jq?

Keep review status in the object you store. Example: genome query BRCA1 --format json | jq '[.[] | select(.significance == "Pathogenic") | {id: .rsid, significance, review_status}]'. Do not grep human text for the word Pathogenic.

Why is .gnomad null in jq?

You probably installed lite. Use standard or full, then genome db stats. Absence of frequency is not a classification. Mixing a lite CLI with an API that has gnomAD will also confuse your mapper.

Can I jq the HTTP API the same way?

Yes. curl -s https://api.genome.sh/v1/query/rs334 | jq . Field names match the CLI JSON for the same snapshot family. Map explicitly if you mix tools.

Should a coding agent parse human output?

No. Tell it --format json and jq. Human output is for you, watching one rsID.

genome.sh reports public annotations. It is informational software, not a medical device or a substitute for clinical care.