Northlane Pay featured image
← All work
Fintech · Cloud migration · Kubernetes · FinOps

Rebuilding a payments platform for scale — without a minute of downtime

Executive summary: Northlane Pay processed growing card volume on a single-region, hand-built cloud account. Over 14 weeks we migrated the platform to a multi-account architecture on Kubernetes, cut over under live traffic, and cut infrastructure spend 43% — while availability rose to 99.99%.

−43%infrastructure cost
99.99%availability (12 mo)
deploy frequency

Client background

Northlane Pay is a payments scale-up serving mid-market merchants across Europe. Transaction volume tripled in eighteen months; the platform had been assembled quickly by a small founding team and had never been re-architected.

The challenge

Everything ran in one cloud account in one region. A regional incident meant total outage; a mis-scoped IAM change could take down production. Costs grew faster than volume because instances were sized for peak and never revisited. Compliance reviews for new banking partners kept stalling on the architecture.

What discovery found

!Approximately 38% of compute spend was idle capacity provisioned "just in case."
!No isolation between production and staging — a staging load test had already caused one production incident.
!Deploys required a 4-hour maintenance window and manual database steps.
!Recovery from region failure was estimated at 2+ days; the partner bank required under 4 hours.

The solution

We designed a multi-account landing zone with strict separation of production, staging, and tooling, running workloads on managed Kubernetes across two regions with automated failover. State moved to managed databases with cross-region replicas. Every resource was recreated in Terraform, and a FinOps baseline (tagging, budgets, rightsizing alerts) was set before migration so savings were measurable from day one.

Implementation

Wks 1–2
Audit & target architecture
Workload inventory, cost analysis, compliance gap review; architecture signed off with the partner bank's auditors.
Wks 3–6
Landing zone & IaC
Multi-account setup, network design, and Terraform modules for every existing resource.
Wks 7–11
Workload migration
Services containerized and moved in dependency order, each behind a traffic-splitting cutover with instant rollback.
Wks 12–14
Live cutover & tuning
Database cutover with under 90 seconds of write pause, then two weeks of rightsizing and failover drills.

Results

Infrastructure spend down 43% within two billing cycles, verified against the pre-project baseline.
99.99% measured availability over the following 12 months, including two cloud-provider incidents absorbed by failover.
Deploy frequency rose from weekly maintenance windows to daily releases — a 6× increase.
Region-failure recovery tested at 22 minutes, against the 4-hour contractual requirement.
Passed the partner bank's architecture review on first submission.
They found $40k a month of waste in the first two weeks, then rebuilt the platform without a single customer-facing incident. The migration plan was the most thorough document I've seen from a vendor.
D. KellerCTO, Northlane Pay