Skip to main content

Futugrid: from one VM to four EKS clusters

Futugrid's product ran as some thirty Java microservices on a single virtual machine, deployed half by hand. Since 2022 we have built and operated the AWS platform it runs on today — four EKS clusters, full CI/CD, managed databases and single sign-on.

Client
Futugrid Technologies OÜ
Period
Services
Stack
AWS, EKS, RDS, Kafka, AWS IoT Core, Helm, GitLab CI, Authentik, OpenTelemetry, Grafana LGTM, Tailscale, WireGuard, IPsec, Teltonika

The starting point

When Futugrid came to us in autumn 2022, the whole product lived on one virtual machine. Around thirty Java microservices ran side by side on it, and getting a new version out meant a semi-manual deployment: build, copy, restart, hope nothing else on the box noticed. It had carried the company through its early stage, but there was no separation between environments, no repeatable way to roll back, and every release depended on the person doing it.

What we built

We replaced the single machine with an AWS platform built around Amazon EKS:

  • Four Kubernetes clusters with distinct jobs: production, staging, a sandbox for experiments, and a tooling cluster that hosts the shared services the other three depend on.
  • CI/CD on GitLab Runner with a Helm chart per service. A merge builds an image, runs the tests and deploys to staging; production is the same pipeline against the same chart, so what was tested is what ships. Secrets are managed rather than copied around.
  • Managed databases on RDS, with backups and point-in-time recovery handled by AWS instead of by cron on a VM.
  • Kafka and AWS IoT Core for the device and streaming side of the product.
  • Single sign-on on Authentik, so the team logs into the cluster tooling, dashboards and internal services with one identity.
  • Private access over Tailscale. Internal services — cluster tooling, dashboards, admin interfaces — are not exposed to the internet at all; the team reaches them through a Tailscale network tied to the same identities as the SSO.
  • Site connectivity on Teltonika edge routers. The solar parks and the third-party providers the product exchanges data with are linked to the platform over IPsec and WireGuard tunnels terminated on Teltonika routers at the edge, so field equipment talks to the cloud over encrypted links rather than public endpoints.
  • Full observability on the Grafana LGTM stack: every service is instrumented with OpenTelemetry, and traces (Tempo), logs (Loki) and metrics (Mimir) land in one Grafana. A slow request can be followed from the API gateway through the microservices it touched down to the database query, with the matching log lines alongside.

How it runs today

The engagement did not end with the migration. We operate the platform: alerting on the same metrics and traces the developers use, cluster and dependency upgrades, cost review, and the changes that come with every new service the Futugrid team adds. Deployments that used to be an event are now a merge request, and an environment for trying something out is a namespace away rather than a second server nobody wants to touch.

Have a project in mind?

Tell us what you need and we will get back to you.