Co-founder · Forward-deployed engineering
Harsh builds and ships zero-to-one AI products with Brandlabs — architecture, evals, and production systems inside client environments. He focuses on making regulated AI measurable and operable, not demo-grade.
BAA availability, retention defaults, region options, and eval portability — how Brandlabs shortlists model providers for healthcare and regulated buyers.
How to evaluate refusal, escalation, medical-advice boundaries, and hallucinated dosages — before a health assistant meets real users.
Golden sets, adversarial sets, and ownership so your LLM suite still means something after six months of prompt thrash.
The Brandlabs Discovery Sprint pattern for setting LLM eval targets in weeks 1–3 — before feature thrash eats the timeline.
Sampling, cost caps, flake handling, and which LLM eval suites must stay required vs nightly — so production cannot silently regress.
Inputs, outputs, kill criteria, and how Brandlabs maps a Discovery Sprint into a 12–20 week zero-to-PMF engagement.
Foundation → build to alpha → production-harden → PMF hunt — with eval and compliance gates at each phase.
Escalation, localization, and content safety patterns for healthcare chat on WhatsApp — from shipping Myna Bolo in production.
Productizing 69,000+ obligations and regulations that change ~15 times a day — without turning an AI assistant into unauthorized legal advice.
Why we treat the eval suite as a first-class deliverable — and how it changes how we negotiate scope.
The cultural and technical loop that turns production incidents into permanent guardrails.
Not a project manager, not a consultant. A senior engineer who lives in your Slack and writes the hardest code.