Skip to the content
Georgi DimitrovdaTuzzo

GeneScope

Reads a consumer DNA export against ClinVar and PharmGKB, then four agents interpret it

Role
Solo, directing agents
Status
Paused
Source
Public on GitHub
Stack
TypeScriptNode.jsNext.jsClinVarPharmGKBClaude Code skills

In numbers

308K

clinically significant ClinVar variants kept from the full download

4

Claude agents interpreting in parallel, plus an orchestrator

3

consumer DNA formats detected automatically

The problem

A raw DNA file from a consumer test is about 600,000 genotype rows with no interpretation attached. The public evidence to read it exists, in ClinVar for disease variants and in PharmGKB and the CPIC guidelines for drug-gene interactions, but it is spread across databases nobody queries by hand. My degree is in pharmacology, and the drug-metabolism genes (CYP2D6, CYP2C19, DPYD, TPMT and others) are the part I studied, with published guidelines behind them.

The approach

Deterministic code does the matching and agents do the reading. A TypeScript pipeline parses the file and matches every genotype against the databases first; only then do agents research what the matches mean for this person, and an orchestrator merges what they find.

How it works

  1. Matching before interpretation

    Setup downloads ClinVar, about 415 MB, and filters it to about 308,000 clinically significant variants; PharmGKB's clinical annotations, with evidence levels 1A to 2B, ship in the repo. The genome becomes queryable JSON keyed by rsID on GRCh37. Exports from MyHeritage, 23andMe and AncestryDNA are detected and loaded natively.

  2. An intake like a clinic

    Before any analysis the tool asks what a doctor would: age, sex, ancestry, medications, family history, lifestyle and the person's own concerns. The pharmacogenomics agent checks its findings against the medications the person takes.

  3. Four specialists at once

    Four Claude agents research in parallel: pharmacogenomics, verified against CPIC and dbSNP; methylation and neurotransmitter genes, read as pathways instead of single SNPs; disease risk, with ClinVar findings checked again and set against family history; and lifestyle and nutrition. An orchestrator writes the cross-domain synthesis, and a follow-up loop asks targeted questions about the person's own results.

  4. Local by design

    The genome file, the findings and the reports stay in a folder on the machine, and nothing is uploaded. Each person gets an isolated profile folder, so two people's results never mix. A static V1 report generator stays beside V2 as a baseline.

What I chose, and what lost

Chose

An agent team that researches each finding

Over

V1's static database of hardcoded descriptions

A template prints the same paragraph for a variant whatever the person's medications or history are. The agents check current sources against the intake.

Outcome

The command-line flow through Claude Code is the working interface. A Next.js report viewer exists, built for V1's report layout and not yet updated for V2's five-file output, and the README says so. The tool is for information and education; its disclaimer rules out clinical diagnosis and sends medical decisions to a genetic counsellor or a doctor. It is paused, and personal results stay off this site.