What is a VCF file, column by column
Variant Call Format is a text (or bgzipped text) table defined by the hts-specs. Header lines start with ## for metadata and #CHROM for the column names. Data lines follow.
The fixed columns are CHROM, POS, ID, REF, ALT, QUAL, FILTER, INFO, then FORMAT, then one column per sample. ID is often an rsID when the site is in dbSNP. It can also be a dot.
POS is 1-based on the named contig. REF and ALT are the alleles as written for that assembly. If the header says GRCh37 and you treat POS as GRCh38, you are reading the wrong base.
The columns that actually matter for lookup
CHROM and POS locate the site on an assembly. ID is the public handle when it is an rsID such as rs1799945. REF and ALT are the alleles you must keep if you query chr:pos:ref:alt.
INFO holds site-level annotations from the caller: DP, MQ, allele counts, and anything else the pipeline wrote. FORMAT names the per-sample fields. GT is the genotype.
FILTER is not decoration. PASS means the caller kept the site. A site that failed filters is not a clean ClinVar lookup target until you understand why it failed.
- CHROM, POS, REF, ALT: the site on a named assembly
- ID: rsID or a dot
- QUAL, FILTER: caller confidence
- INFO: site-level tags
- FORMAT and sample columns: GT and the rest
Genotypes, GT, and the meaning of a missing row
GT encodes the called alleles relative to REF and ALT: 0/0 homozygous reference, 0/1 heterozygous, 1/1 homozygous alternate, and ./ . not called. Phased calls use a bar instead of a slash.
Most VCFs only list sites that differ from the reference, or sites the caller chose to emit. A missing row is not automatically 0/0, especially on a chip or a targeted panel.
If you convert a 23andMe export to VCF and a marker was never on the array, it will not appear. Write not called. Do not fill the reference allele as if it had been sequenced.
- 0/0: homozygous reference, when the site was called
- 0/1: heterozygous
- 1/1: homozygous alternate
- ./.: not called
- Absent row: not called unless you have a gVCF block that says otherwise
A VCF is not a whole genome
A VCF is a list of variant records. It is not a FASTA of your chromosomes. It is not proof that unlisted bases were reference.
gVCF includes blocks of reference confidence so you can tell a no-call from a reference call. Ordinary VCFs often omit those blocks. Consumer chip tables never had them.
A panel VCF covers the panel, an exome VCF covers the bait, and a chip-derived VCF covers the array. ClinVar has millions of interpreted variants. Coverage is the first limit, not the annotation tool.
VCF versus 23andMe and versus BAM
A 23andMe or AncestryDNA raw-data file is not a VCF. It is a simpler rsid/chromosome/position/genotype table, usually on GRCh37, with a few hundred thousand rows.
BAM and CRAM are alignments. They store reads, not a finished variant table. genome extract can pull variants from CRAM or BAM locally. The HTTP API will not accept those files.
You can look up VCF IDs and 23andMe rsIDs the same way: genome query rs429358. You annotate a VCF with genome annotate. You parse a consumer export in the importer at /import or by reading the table on disk.
Assemblies, contig names, and compression
chr6 versus 6 versus NC_000006.11 are contig naming problems. Tools that join ClinVar and gnomAD need a consistent chromosome naming scheme. Check the header before you blame the annotator.
bgzip plus tabix is the usual on-disk form: file.vcf.gz and file.vcf.gz.tbi. Plain .vcf works for small files and is painful for whole genomes. genome annotate accepts gzipped VCF.
GRCh37 versus GRCh38 is the other silent error. rs1800562 is a stable key. The numeric POS is not. Lift over if you must mix, or query by rsID.
What to do with a VCF on this machine
Look up IDs and coordinates in ClinVar and gnomAD. Annotate the file locally. Do not upload it to a random website if you care about privacy.
Install path: cargo install genome-sh, genome db install standard, then genome annotate sample.vcf.gz --format json. Add --filter clinical when you want ClinVar-touched rows.
Single-site lookup does not need the file: genome query chr6:26090951 or genome query rs1800562. The REST API at https://api.genome.sh/v1/query/{id} is the same idea for identifiers only.
cargo install genome-sh genome db install standard genome annotate sample.vcf.gz --format json genome annotate sample.vcf.gz --filter clinical --format json genome query chr6:26090951 genome query rs1800562
Common VCF mistakes
Opening a VCF in Excel mangles contig names, scientific notation in POS, and header lines. Use bcftools, the genome CLI, or a proper parser.
Pasting the file into a chatbot sends genotypes off the machine. Send an rsID if you need a public lookup. Keep the VCF on disk.
Treating every missing BRCA1 row as negative is the clinical-shaped error. A chip VCF is not a BRCA test. genome.sh prints public annotations. It is not a medical device.
curl -s https://api.genome.sh/v1/query/rs1799945 | jq . curl -s https://api.genome.sh/v1/sources | jq .
Questions
What is a VCF file?
Variant Call Format: a standard text table of SNPs and indels. One row is a site. Sample columns hold genotypes such as 0/1. See the VCFv4.2 spec.
Is a 23andMe file a VCF?
No. It is a text export of chip calls (rsid, chromosome, position, genotype). genome.sh can look up the rsIDs, and the browser importer accepts several consumer formats.
Why is my VCF GRCh37 when gnomAD is GRCh38?
Assemblies differ. Lift over or query with the assembly that matches the file. Mixing them silently maps variants to the wrong place. Prefer rsIDs.
Can I open a VCF in Excel?
You can, badly. Use bcftools, the genome CLI, or a proper parser. Header lines start with #, and POS is not a spreadsheet number.
What does GT 0/1 mean in a VCF?
Heterozygous: one reference allele and one alternate allele at that site, when the site was called. ./ . means not called.
VCF vs BAM: what is the difference?
BAM/CRAM store aligned reads. VCF stores called variants. genome extract can derive variants from CRAM or BAM locally. Do not upload either file to the HTTP API.
How do I annotate a VCF file?
genome annotate sample.vcf.gz --format json after genome db install standard. Use --filter clinical for ClinVar-touched rows. The file stays on disk.
Does a missing VCF row mean wild type?
No. Unless you have gVCF reference blocks, a missing row means the site was not emitted. On a chip or panel, treat it as not called.