Futugrid: from one VM to four EKS clusters
Futugrid's product ran as some thirty Java microservices on a single virtual machine, deployed half by hand. Since 2022 we have built and operated the AWS platform it runs on today — four EKS clusters, full CI/CD, managed databases and single sign-on.
- Client
- Futugrid Technologies OÜ
- Period
- Stack
- AWS, EKS, RDS, Kafka, AWS IoT Core, Helm, GitLab CI, Authentik, OpenTelemetry, Grafana LGTM, Tailscale, WireGuard, IPsec, Teltonika
The starting point
When Futugrid came to us in autumn 2022, the whole product lived on one virtual machine. Around thirty Java microservices ran side by side on it, and getting a new version out meant a semi-manual deployment: build, copy, restart, hope nothing else on the box noticed. It had carried the company through its early stage, but there was no separation between environments, no repeatable way to roll back, and every release depended on the person doing it.
What we built
We replaced the single machine with an AWS platform built around Amazon EKS:
- Four Kubernetes clusters with distinct jobs:
production,staging, asandboxfor experiments, and atoolingcluster that hosts the shared services the other three depend on. - CI/CD on GitLab Runner with a Helm chart per service. A merge builds an image, runs the tests and deploys to staging; production is the same pipeline against the same chart, so what was tested is what ships. Secrets are managed rather than copied around.
- Managed databases on RDS, with backups and point-in-time recovery handled by AWS instead of by cron on a VM.
- Kafka and AWS IoT Core for the device and streaming side of the product.
- Single sign-on on Authentik, so the team logs into the cluster tooling, dashboards and internal services with one identity.
- Private access over Tailscale. Internal services — cluster tooling, dashboards, admin interfaces — are not exposed to the internet at all; the team reaches them through a Tailscale network tied to the same identities as the SSO.
- Site connectivity on Teltonika edge routers. The solar parks and the third-party providers the product exchanges data with are linked to the platform over IPsec and WireGuard tunnels terminated on Teltonika routers at the edge, so field equipment talks to the cloud over encrypted links rather than public endpoints.
- Full observability on the Grafana LGTM stack: every service is instrumented with OpenTelemetry, and traces (Tempo), logs (Loki) and metrics (Mimir) land in one Grafana. A slow request can be followed from the API gateway through the microservices it touched down to the database query, with the matching log lines alongside.
How it runs today
The engagement did not end with the migration. We operate the platform: alerting on the same metrics and traces the developers use, cluster and dependency upgrades, cost review, and the changes that come with every new service the Futugrid team adds. Deployments that used to be an event are now a merge request, and an environment for trying something out is a namespace away rather than a second server nobody wants to touch.