6 min read

🍕 The Day Our GitLab Pipeline Queue Broke Us — And What We Did About It

🍕 The Day Our GitLab Pipeline Queue Broke Us — And What We Did About It

A story about CI/CD, patience, pizza, and eventually — Kubernetes.


Imagine a pizza restaurant.

A really good one. Popular. Always busy.

You walk in, place your order, and the waiter smiles and says: "Your pizza will be ready in 8 to 12 minutes."

You sit down. You wait. 8 minutes pass. Then 12. Then 15.

You ask the waiter. He smiles again. "We're just waiting for an oven to free up."

The pizza is eventually perfect. But by the time it arrives, you've eaten half the breadsticks and lost the will to enjoy it.

That was us. Except instead of pizza, it was deployments. And instead of ovens, it was GitLab SaaS runners.


The Beginning — When Everything Was Fine

When we first set up our GitLab CI pipelines, SaaS runners were a gift.

No setup. No maintenance. Push code, pipeline runs, done. We were a small team moving fast. Nobody wanted to think about infrastructure just to run a few tests and deploy an app.

And for a while — it was genuinely great.

The pipelines were simple. Run tests, build the image, deploy. The whole thing took 6 minutes. Developers pushed code and went for coffee. By the time they came back, it was done.

Life was good.


The Cracks Start Showing

Then the team grew.

More developers. More features. More pipelines running simultaneously. And that's when we started noticing something strange.

The pipelines weren't getting slower. They were just... waiting. Before a single test ran, the pipeline would sit in a queue. Pending. Scheduled. Waiting for a runner to pick it up.

At first it was 2 minutes. Annoying but tolerable.

Then it was 5. Then 8. Then one particularly bad afternoon — 14 minutes before the first job even started.

We had a 6-minute pipeline that was taking 20 minutes to complete.

And the thing that made it worse? The waiting wasn't consistent. Some pipelines picked up instantly. Others sat there. You couldn't predict it. You'd push a hotfix on a Friday afternoon — the kind that needed to go out now — and watch it sit in the queue while you refreshed the page every 90 seconds.

The developers started doing the thing you never want developers to do: they started working around the CI pipeline. Skipping it. Pushing directly when something was urgent. The pipeline that was supposed to give everyone confidence was quietly becoming the thing people resented.


The Deeper Problem Nobody Noticed at First

Around the same time, something else started bothering us.

Our pipelines needed to pull large Docker images. A base image for our build environment — heavy, specialised, slow to download. Every time a pipeline ran, it pulled this image fresh. On GitLab's shared runners, there was no caching. Every job started from scratch, pulling gigabytes of image every single time.

It was like going to that pizza restaurant, and every time you ordered, the chef first had to drive to the supermarket to buy ingredients. Even if you ordered the same pizza you ordered yesterday.

Then came the security conversation.

Our pipeline needed to connect to an internal service — an artifact registry that sat behind our private network. On SaaS runners, that was simply not possible. The pipeline ran on GitLab's infrastructure. It had no way to reach anything inside ours.

We had three choices: make the internal service publicly accessible (no), pass everything through environment variables and workarounds (painful), or change how our runners worked.

We finally had our reason.


The Conversation That Changed Things

I remember the meeting clearly.

Someone shared a screen showing the pipeline timeline. There was a thick grey bar at the start of every pipeline — the queue time. It was longer than the actual pipeline.

Someone else did the math out loud. If every developer pushed 5 times a day, and every push waited an average of 8 minutes before anything happened — we were collectively spending hours every day just waiting. Not blocked by tests. Not blocked by slow builds. Blocked by a queue.

The question someone asked: "What if we just ran our own runners?"

The immediate response: "That sounds like more infrastructure to maintain."

Which was true. But so was the alternative — a team slowly losing faith in their own deployment pipeline.

We decided to try it.


Moving to Self-Hosted — The Honest Version

We already ran Kubernetes. Our applications lived there. Our monitoring lived there. Our entire platform was Kubernetes.

So the idea was: run the GitLab runners inside our own cluster. Each CI job would become a small container — spin up, do its work, disappear. Clean. Isolated. On our own compute.

The setup wasn't instant. There was a weekend of work. Configuration, permissions, figuring out how the runner talked to GitLab, making sure secrets were handled correctly, ensuring jobs couldn't accidentally touch things they shouldn't.

But then we turned it on.

The first pipeline ran. No queue. Zero seconds of waiting. The job started the moment someone pushed code.

We ran it again. Same thing.

Someone on the team sent a message in Slack: "Did the pipeline always start this fast?"

No. It didn't. We had just forgotten what it felt like when it wasn't broken.


What Actually Changed

The queue disappeared entirely. Our runners lived in our cluster. When a job came in, a runner picked it up immediately. No shared pool. No waiting for GitLab's infrastructure to find us a slot. Ours. Instant.

Images cached on the nodes. The heavy base image we pulled every single time? Our cluster nodes started caching it. First pull on a fresh node — slow, as expected. Every pull after that — instant. The pipeline that used to spend 3 minutes just pulling the image was now spending 10 seconds.

Private network access worked. Our internal artifact registry, our private services, our internal tooling — the pipeline could reach all of it. Because the pipeline ran inside our network. No workarounds. No exposing things publicly. It just worked.

Developers started trusting the pipeline again. This was the one nobody expected. When the pipeline is fast and consistent, people stop working around it. They wait for it. They trust it. That cultural shift is worth more than any technical metric.


The Things We Got Wrong

It wasn't all smooth.

The first week, a developer accidentally pushed a change that caused every pipeline to fail with a confusing error about permissions. Turns out the runner needed specific access to certain cluster resources for deployment jobs — and we had locked it down too tightly in the name of security.

We spent half a day debugging something that would never have been an issue on SaaS runners. Because on SaaS runners, someone else handles that complexity.

That's the honest trade. When you bring infrastructure in-house, you own its problems too.

There was also the time a node ran out of disk space — all the cached images had quietly filled it up. Pipelines started failing in a way that looked completely unrelated to disk space. It took an embarrassingly long time to find the cause.

And runner version updates. On SaaS, GitLab just handles it. On self-hosted, you need a process for keeping the runner version current. We learned this the hard way when a mismatch between runner version and GitLab version caused intermittent job failures for a week before we connected the dots.


The Restaurant Analogy Revisited

Going back to the pizza restaurant.

SaaS runners are like a restaurant that's great when it's not busy. Convenient, zero effort, always available. But on Friday night when everyone wants pizza — you wait. And you can't control the wait. You're at the mercy of how many ovens they have.

Self-hosted runners on Kubernetes are like building your own kitchen.

More work upfront. You have to buy the oven, learn to use it, maintain it. But when you want pizza — you make pizza. No queue. Your oven. Your rules. And if you need it hotter, faster, or with a specific ingredient that the restaurant doesn't stock — you can.

The catch is: you're now responsible for the kitchen. If the oven breaks, that's on you.

Neither is wrong. It depends on how often you need pizza, how specific your requirements are, and whether you have someone who knows how to maintain a kitchen.


How We Run It Today

We didn't abandon SaaS runners entirely.

Simple, lightweight jobs — quick linting checks, basic unit tests, small validation tasks — still run on SaaS. There's no queue because they finish before the queue has time to matter. And we don't want to waste our own compute on work that doesn't need it.

Heavy jobs — building Docker images, running integration tests, deploying to internal environments, anything that needs to reach our private network — those run on our self-hosted runners.

Each job knows which runner to use. Developers don't think about it. It just routes correctly.

The pipeline that used to take 20 minutes from push to done — counting the queue — now takes 6. The same 6 minutes it always actually took.

We just stopped wasting the other 14.


If You're Sitting on That Fence

If your pipelines are fast, your team is happy, and nothing is blocking you — stay on SaaS runners. The operational simplicity is real and genuinely valuable.

But if you're watching that grey queue bar at the start of every pipeline. If you're working around your own CI because it's too slow. If you need to reach something inside your own network. If your team has stopped trusting the pipeline —

You already know what needs to change.

The kitchen is worth building. You just have to be ready to maintain it.


I write about real infrastructure decisions, DevOps, and the honest version of platform engineering at:

👉 blog.rootbysatya.in

— Satya 👨‍💻 Platform Engineer | GitLab CI | Kubernetes | Occasionally maintains the kitchen


💬 What pushed your team to move away from SaaS runners — or what's keeping you on them? Drop it in the comments. Genuinely curious. Follow RootBySatya for more honest takes on real infrastructure.