Lab notes, not a polished paper — this is what the variant-scoring model actually looks like right now.

Architecture

We use gradient-boosted trees over engineered variant-interaction features rather than a deep model, mainly because our cohort sizes don't justify the extra capacity yet.

What's working

Nested cross-validation across two cohorts gives us a false-positive rate inside our target range.

What's not

The third validation cohort has different ancestry composition than the first two, and early results suggest some feature importances shift. We're treating that as a real finding, not noise to explain away.

Share: X / Twitter LinkedIn Email