Selected work

Outcomes, in specific numbers.

A handful of recent engagements. We can share more under NDA. Most of our work is.

Rune Capital
Fintech
Rune Capital

Re-platforming a lending engine that approves loans in 4 seconds

Problem

Rune's legacy lending engine took 11 minutes to decision a loan, with manual review on 38% of applications. Throughput capped before they could enter the Tier-2 city market.

Solution

We re-architected the decisioning service around an event-driven core in Go, replaced the monolithic rules engine with a feature store + gradient-boosted scorecards, and introduced shadow-mode model validation before any production flip.

Technologies
GoKafkaFeastXGBoostPostgreSQLAWS EKS
Results

Decision time fell to 4.1 seconds. Approval rate rose 37% within two quarters with default rates unchanged. Rune launched in 6 new states with the same risk team.

4.1s
avg. decision time
37%
lift in approvals
99.99%
uptime YTD
Veda Health
Healthcare
Veda Health

An AI triage assistant deployed across 22 clinics in nine months

Problem

Veda's outpatient clinics were drowning in repetitive intake. Clinicians spent 14 minutes per visit on documentation; patients waited an average of 47 minutes.

Solution

We built an AI triage assistant trained on de-identified intake transcripts, with strict guardrails and a human-in-the-loop review queue. Deployed across 22 clinics with rolling regional rollouts and on-prem inference for sensitive workloads.

Technologies
PythonFastAPILlama 3Postgres + pgvectorAzureFHIR
Results

Patient wait time dropped 42%. Clinician adoption hit 92% by month three. Veda passed an external HIPAA audit with zero findings.

−42%
patient wait time
92%
clinician adoption
HIPAA
audited
Northwind AI
SaaS
Northwind AI

Scaling a multi-tenant LLM platform from prototype to $4M ARR

Problem

Northwind's prototype LLM platform was a single-tenant Python service that fell over above ~120 concurrent users. Inference costs were 31% of MRR.

Solution

We split inference, retrieval and orchestration into independently scaled services, introduced semantic caching, switched to fp8 quantized models on inference-optimized hardware, and added a usage-aware autoscaler.

Technologies
TypeScriptRustKubernetesvLLMRedisPostgresGCP
Results

Throughput rose 11× on the same infra footprint. Inference cost per request dropped 68%. Northwind closed Series B at $4M ARR.

11×
throughput vs. v1
−68%
inference cost
SOC 2
Type II

Want references you can actually call?

Request references