Interpret 23andMe raw data locally on this machine, without another upload.

Interpret 23andMe raw data locally by treating the export as a table of rsIDs and letters. You can look those rsIDs up in ClinVar and gnomAD without sending the file to a third party.

What a spit-kit file actually is

A 23andMe or Ancestry raw-data file is not your genome. It is a table of a few hundred thousand common SNPs: an rsID, a chromosome, a position, and two letters. The rest of the three billion bases were never looked at.

Those letters are a genotype at sites the chip happened to probe. They can tell you something about ancestry, a handful of well-studied variants, and very little about the rest of medicine. Most pathogenic BRCA1 variants will not be in the file. Most of your DNA will not be in the file.

People upload that table because they want a story. The safer story is: look up the public record for the rsIDs you care about, keep the file on your machine, and write “not called” when a site is missing.

  • Chip: a sample of common SNPs, not whole-genome sequencing
  • Row: rsID plus two letters, if the site was called
  • Missing row: not called, not proof of the reference allele

Interpret 23andMe raw data locally, not on a stranger's server

The export is a table of rsIDs and letters. The public question is what ClinVar and gnomAD say about those rsIDs. The private question is which letters are in your file.

You do not need to upload the export to answer either question. Query public ids over HTTPS if you want. Match genotypes on disk or in the browser tab.

genome.sh is built around that split. The HTTP API never accepts a 23andMe file. The CLI and the importer at /import are the file paths.

What the export actually contains

A typical consumer export has on the order of 600,000 rows: rsid, chromosome, position, genotype. Comment lines start with #. The genome build is usually GRCh37, even when current gnomAD is GRCh38.

It is a genotyping chip, not whole-genome sequencing. It is not an exome. It is not a gVCF. Sites that were never on the array are absent, not sequenced as reference.

Some rows use vendor i-ids instead of rsIDs. Those probes may have no ClinVar or gnomAD record under an rs number. Skip them rather than inventing a dbSNP id.

  • rsid: dbSNP id or a vendor i-id
  • chromosome and position: usually GRCh37
  • genotype: two letters, or a no-call such as -- or 00
  • coverage: a few hundred thousand sites, not the genome

Chip coverage versus ClinVar

Most ClinVar pathogenic variants are not on the chip. A clean consumer report is not a clinical exome. It is not a complete BRCA1 or BRCA2 test.

Founder variants appear in tutorials because they are famous, not because the array is complete. If those rsIDs are absent from the export, they were not called.

Treat every missing marker as not called. Do not infer the reference allele. Do not write wild type in a PDF because the row was missing.

Look up rsIDs from the file

Pick an rsID from the export and query it. rs429358, rs1800562, rs1799945, and rs334 are the ones people actually type. The lookup returns public annotation, not a diagnosis.

Install the crate genome-sh; the binary is genome. genome db install lite is enough for ClinVar, while standard adds gnomAD frequencies. Then genome query the rsID.

You can also GET https://api.genome.sh/v1/query/rs429358 with no key. Send the rsID, not the file. For thousands of rows, stay on the CLI or the importer rather than looping HTTP.

cargo install genome-sh
genome db install standard
genome query rs429358
genome query rs1800562
genome query rs1799945
genome query rs334 --format json

The in-browser importer at /import

The local analysis page parses supported files in the tab. Processing stays in the browser. The file is not uploaded to genome.sh.

That is the right first pass when you do not want to install Rust. For pipelines, notebooks, and agents, use the CLI on disk.

The HTTP API behind /query is still identifier-only. Importing a file in the tab does not open a POST-a-genome route.

AncestryDNA and MyHeritage exports

The same workflow applies to AncestryDNA and MyHeritage raw-data tables. They are still chip exports: rsIDs, positions, letters. They are still not VCF, and they are still not sequencing.

Build and probe sets differ by vendor and by chip version. An rsID present on one array can be absent on another. Missing remains not called.

If you convert an export to VCF, keep the assembly in the header. Annotate that VCF locally with genome annotate. Do not upload the converted file either.

genome annotate converted-chip.vcf.gz --format json
genome annotate converted-chip.vcf.gz --filter clinical --format json
curl -s https://api.genome.sh/v1/query/rs1800562 | jq .

Read ClinVar and gnomAD without turning them into a diagnosis

ClinVar significance is a submitted interpretation plus a review status. gnomAD frequency is a population observation. Neither is medical advice.

rs1800562 and rs1799945 are HFE sites that show up constantly in chip reports. Frequency in European ancestry gnomAD is not tiny. Zygosity and ferritin still belong in a clinic, not in a CLI.

rs429358 is one APOE SNP, and haplotype calling needs rs7412 as well. APOE status is sensitive, so prefer local lookup. Do not paste those genotypes into a prompt.

Promethease-style reports without a marketplace upload

People want a readable report from a 23andMe file. Promethease and several web tools do that after an upload or a paid report. genome.sh does the match on disk.

The printable A4 template lives in the CLI repository. An agent can fill it if you copy the prompt, copy the template out of the repo, and keep the file local. The PDF lists structured findings, not 600,000 rows.

You get public annotations for sites that were called. You do not get a literature dump for every GWAS hit unless you add that work. You do not get a clinical test.

  • Local CLI: genome query and genome annotate on disk
  • Browser: /import parses in the tab
  • Agent: local CLI plus the report template, no upload
  • HTTP API: rsIDs only, never the export

Questions

How do I interpret 23andMe raw data locally?

Keep the export on disk. Query rsIDs with genome query, or parse the file in the /import tab. Do not upload the file to genome.sh. Missing markers are not called.

Interpret 23andMe without uploading?

Yes. Identifier lookup can use https://api.genome.sh/v1/query/{rsid}. File matching stays in the CLI or the in-browser importer. The API rejects the export.

Is this a Promethease replacement?

It is a local lookup and annotation path, not a commercial report marketplace. See the Promethease alternative guide for the comparison.

Will it find BRCA1 founder variants?

Only if those rsIDs are on the chip and called. Many important BRCA1 variants are not on consumer arrays. Absence is not a negative BRCA test.

Does genome.sh store the 23andMe file?

No. The website importer does not upload. The CLI reads from disk. The HTTP API never accepts the file.

How do I look up 23andMe raw data in ClinVar?

Take rsIDs from the export and run genome query rs1800562, or annotate a converted VCF with genome annotate --filter clinical --format json.

What if an rsID is missing from my 23andMe file?

Write not called. Do not infer the reference allele. Consumer chips miss most ClinVar pathogenic sites.

Does AncestryDNA use the same local workflow?

Yes. AncestryDNA and MyHeritage exports are also chip tables. Query rsIDs locally, and treat missing rows as not called.

genome.sh reports public annotations. It is informational software, not a medical device or a substitute for clinical care.