Tick Mitt · Data Science & AI Engineering Intern · Summer 2026
TiCK Heatmap
A live, community-powered tick-encounter and Lyme disease-prevalence tracking platform, backed by a hierarchical Bayesian engine and built privacy-first. I was the sole engineer.

- HHS TOPxHHS Tech Sprint grant, advanced to final-stage review
- $1MHHS TOPxHHS Tech Sprint grant, advanced to final-stage review
- engineer, from product to infrastructure to model
- Soleengineer, from product to infrastructure to model
- precise coordinates ever stored or logged
- Zeroprecise coordinates ever stored or logged
The problem
Tick-borne disease risk is intensely local and changes week to week, but there is no timely, community-level picture of where ticks are actually being found. Public surveillance data is county-resolution at best and often lags by months or years.
TiCK Heatmap lets anyone report a tick encounter from their phone in about a minute, gives them something useful back right away, and aggregates the reports anonymously into a public map that updates as reports come in. The platform became Tick Mitt's central pitch for a $1M HHS TOPxHHS Tech Sprint grant, which advanced to final-stage review.
The hard parts were making it genuinely anonymous and turning sparse, self-selected reports into a signal that is honest about its own uncertainty.
My role
I was the sole engineer and product owner, responsible for architecture, frontend, backend, infrastructure, and the statistical model behind the map. I kept a written architecture decision log throughout, so the reasoning behind each design choice stays clear to whoever inherits the system.
The model: honest estimates from a biased sample
A naive "count reports per area and color by count" map would be confidently wrong. Crowdsourced reports come from a self-selected sample, volume is low and uneven, and one person may report several ticks from a single encounter. I built an empirical-Bayes hierarchical model designed around those problems:
- Two separate questions, never blended into one score. The model estimates how often ticks are being found (sighting density, a Poisson-Gamma model) and what share of tested ticks carry disease (prevalence, a Binomial-Beta model) independently.
- Multi-level geographic pooling. Estimates pool from national to climate region to state to county, with priors informed by CDC and NOAA data. A quiet county borrows strength from its surroundings instead of swinging wildly on three reports.
- Design-effect weighting. Multiple ticks from one submission are correlated, so they are weighted accordingly rather than counted as independent observations.
- Tick-selection bias correction. Not every reported tick gets tested, and the ones that do aren't a random sample, so the model corrects for that selection.
- Output as ordinal tiers, not false precision. Results are bucketed into a small set of risk tiers rather than presented as a probability the data can't support.
The model runs as a scheduled batch job rather than per request. It does a full recompute on a schedule and swaps results in atomically, so a missed run heals itself.
Privacy by design
The platform handles health-adjacent location data, so privacy is built into the structure of the system rather than left to policy:
- No precise location is ever persisted. Reports are geocoded to a ZIP code at ingestion, and the exact coordinates are discarded.
- PII and incident data live in physically separate data stores with separate access roles. No single component can write to both.
- Sparse areas are suppressed, so no low-volume region can be de-anonymized by inspection.
- Map markers are shown as approximate areas. The UI never implies precision the data doesn't have.
Architecture
A fully serverless AWS backend defined as infrastructure-as-code, with a React frontend:
- Frontend: React + Vite, including a mobile-first multi-step survey with schema-driven validation and an interactive national heatmap with county drill-down and address search (Mapbox).
- Backend: AWS CDK, Lambda, API Gateway, Aurora Serverless v2 (PostgreSQL), and DynamoDB. Slow or third-party dependencies are queued off the request path, with dead-letter queues and alarms, because the system is meant to run unattended.
- Delivery: CI/CD through GitHub Actions with keyless cloud credentials, and database migrations applied from inside the private network.

The public landing page. The map updates continuously as reports come in, so the live site will look different from this snapshot.
What I'd point to
- A hierarchical Bayesian model that stays honest about a sparse, biased sample instead of treating raw counts as truth.
- Anonymity as a structural property, not a promise: separated stores, ZIP-only resolution, and suppression of sparse cells.
- Shipping a real health-tech product end to end as one person, with a live public deployment.
The codebase is proprietary to Tick Mitt. This write-up describes the approach at a high level.