Building an Auto-Scaling, On-Demand Minecraft Network on Kubernetes
GamerHaus has been a gaming community since 2007. We've run servers for Tribes, Minecraft, and everything in between. When we decided to build a new Minecraft network, we wanted to think about what our community actually might like. We've hosted all kinds of Minecraft servers in the past, but never one unified experience. And we really love our minigame party night shenanigans, so minigames had to be a part of it. Kubernetes, or at least container orchestration, started to look like a natural choice.
Most Minecraft server networks solve the multi-game problem with a handful of fixed backend servers and a proxy. That works, but it means static allocation, so you either over-provision for peak hours or players wait. And minigames that generate procedural maps create a cleanup problem that never really goes away.
We went the Kubernetes direction. Our lobbies autoscale based on player load. Every minigame is an ephemeral container that spawns on demand, runs its round, and disappears. Survival and Creative are persistent StatefulSets backed by PVCs. We run two completely separate clusters, one for production, and a staging/pre-production environment to test changes in before we zing 'em at the players.
Here's how it actually works.
The Architecture
Hardware: A single Ryzen 9900X system with 128GB RAM. We chose to use this hardware rather than our Proxmox cluster on Dell R640s because the Ryzen's single-threaded performance matters more for Minecraft than, well, really anything else. It's been noticeably faster for this workload.
Kubernetes: k3s. We tried Talos first. Talos is great if you need serious security posture, but for a game server network that's greenfield with unknowns, it was more operational overhead than value. k3s gave us everything we needed with far less friction.
The proxy: Velocity. All servers register themselves in Velocity via plugin-k8s (one of many in-house plugins). Servers are assigned names based on labels configured at launch time. For lobbies, there's a virtual name that weights connections toward the least-loaded instance. Our plugin actively watches Kubernetes with read-only access to pods, using Kube events so server registrations and removals are processed near instantly. We intentionally kept write access to Kube out of the hands of Velocity as it's internet facing. Instead, Velocity talks to controllers which are capable of taking discrete, thought out actions within Kube.
Lobbies: Autoscaling. More players means more lobby pods. The lobby controller regularly polls Velocity to see how many players are on and decides on a target lobby count with a hysteresis band to avoid flapping. After enough players leave, lobbies will terminate and rebalance players. Lobbies are unbuildable jumping points, so this autoscaling is done with completely ephemeral lobby pods that we just delete when we're done with them.
Survival / Creative: Persistent, stateful pods backed by Kubernetes PVCs. Long-running, no autoscaling, no ephemeral nonsense. They're the anchors of the network, backed up both on and off-site routinely to avoid losing our worlds..
Minigames: There's a base container image that all minigames inherit from. Each game type (Siege, Mob Arena, Bloctionary, etc.) is its own image baked from that base. When a player queues up, a Velocity plugin publishes a command to Redis pubsub. The minigame controller receives the message, creates a new pod from the appropriate image, and sends the generated server name back to the minigame plugin. Once the server registers with Velocity, all the players in line to join it get whisked away.
When the game ends, the pod terminates itself and is fully cleaned up by the minigame controller a few minutes later. The proxy handles this shutdown gracefully, zapping everyone right back to a lobby.
Cold start: We average about 30 seconds from the queue firing to playable minigame. Paper's startup dominates this time. We've optimized what we can, and for minigames it's acceptable. You queue up, browse chat or your inventory, and you're in. You get a nice bossbar while you're at it, showing you the estimated launch time, and, if Kube is too busy, your position in a queue waiting for resources to free up.
The Polish That Makes It Invisible
The infrastructure doesn't matter if players keep getting kicked to a loading screen. We spent time on the seams.
Lobby rebalancing: When a lobby pod needs to shut down, we don't wait to drain it. We nuke it from orbit and send all connected players to the remaining lobbies. They see a quick "reconfiguring" screen and land on another lobby. In the same exact position, pitch, and yaw that they were at previously. Since lobbies aren't buildable, this creates a fairly seamless experience.
Server updates without losing players: When we bounce Survival or Creative for updates, players are sent to the lobby with a message that the server will be back shortly. When the updated pod comes online, if you're still in that lobby, you're automatically whisked back.
Velocity rollouts: This one was particularly clever. When we roll a new Velocity instance, it comes up first and registers on the Kubernetes load balancer. Only then does the old Velocity instance unregister and kick every connected player with a transfer packet during its pre-shutdown phase. The clients reconnect to the new proxy fairly transparently. No one manually rejoins, unless their client bugs out on the transfer packet, which is rare but can happen.
What Was Easy, What Was Hard
Networking was easy. It just worked the way we expected. Servers are services. Velocity does discovery. Labels handle naming. No port collisions, no proxy reconnection drama.
Kubernetes distribution choice was the hardest part. Talos ate half a week fighting with Cilium, BGP, PVCs, and a weird MTU issue. k3s solved it in an afternoon.
Cold start time is the biggest open problem. 30 seconds is fine for minigames. We considered keeping one of each minigame pre-warmed, but that'd be more to manage and this works well enough. We'll revisit pre-warming if it becomes a bottleneck.
Everything else was straightforward. Minecraft servers are Java processes. Kubernetes is good at managing Java processes. The wiring is the hard part, and once that's solid the whole thing just runs.
What This Unlocks
No world wipes, ever. Survival and Creative are persistent and they stay that way. We don't have to mess them up to run events.
Procedurally generated minigames. Every round starts from a clean image. We can throw procedural generation at the wall without fear of orphaned data, messing about with cleanup scripts, or worries about breaking the world.
Scaling that matches demand. Lobbies scale up when people are online, down when they're not. Minigames only exist while they're being played.
Full control over everything. Nearly all plugins are in-house. We took an approach inspired by microservices. Small plugins that do one thing, and do that one thing well. Tight integration between components that need it, no compatibility worries.
We're Live
The server is at gamer.haus. Survival (Terralith), Creative (amplified world, no plots), and eleven minigames: Siege (Capture the Flag), Mob Arena, Bloctionary, Rabbit (inspired by Tribes 2), Hunters (ditto), King of the Hill, Soccer (which should really be called Slimeball), Spleef, Turf War, Ice Boat Racing, and Mining Derby. All on the same cluster, all behind the same proxy, all built and maintained in-house.