🤖 How a DevOps Engineer Should Actually Use AI — Not Just Vibe and Deploy
AI is a powerful tool. It's also a very confident liar. Here's how to tell the difference.
Everyone in tech is using AI now.
Copying Terraform from ChatGPT. Asking Claude to write CI pipelines. Pasting error messages into a chat box and deploying whatever comes back.
And honestly? Sometimes it works perfectly.
But in DevOps — where a wrong config doesn't just crash your app but can take down infrastructure, expose secrets, break deployments across environments, or silently corrupt your pipeline — "sometimes it works" isn't good enough.
I've been using AI daily as a Platform Engineer. Managing GKE clusters at work, running a 5-node k3s homelab at home, writing Terraform, building CI/CD pipelines, debugging incidents. This is what I've actually learned — where AI genuinely helps, where it confidently lies, and how to use it without getting burned.
🧠 First, Understand What AI Actually Is
AI doesn't know DevOps. It has seen millions of pages of documentation, blog posts, GitHub issues, Stack Overflow answers, and tutorials — and it predicts what a helpful response looks like based on that pattern.
That distinction matters enormously in our field.
It means:
- It can sound completely correct while being completely wrong
- It doesn't know your environment, your versions, your cloud provider quirks, your team's conventions
- It has a knowledge cutoff — it may not know about breaking changes in the latest release
- It will fill gaps in its knowledge with plausible-sounding answers
- It has no idea what's running in production right now
In DevOps, context is everything. And AI has almost none of yours.
✅ Where AI Actually Helps (Use It Freely Here)
1. 📝 Boilerplate and First Drafts
AI is genuinely excellent at generating starting points — Terraform modules, Kubernetes manifests, Dockerfiles, CI pipeline stages, Ansible playbooks, Helm values files.
It saves you from staring at a blank file and gets you to 70% in seconds.
# Good use — let AI scaffold this, then review every block
resource "google_container_cluster" "primary" {
name = "my-cluster"
location = "us-central1"
initial_node_count = 3
node_config {
machine_type = "e2-medium"
disk_size_gb = 50
}
}
Treat it as a template. Never as a finished config.
2. 🐛 Debugging Starting Point
You have a failing pipeline. A pod in CrashLoopBackOff. A Terraform plan that's doing something unexpected. Paste the error and logs into AI — it gives you 5 possible explanations in 10 seconds.
You still verify each one. But it narrows the search space fast, which is genuinely valuable at 2am.
3. 💡 Understanding Concepts Quickly
Need to understand how Terraform state locking works? What's the difference between blue-green and canary deployments? How does a service mesh handle mTLS? What does depends_on actually do in Terraform?
AI explains these faster and more clearly than docs in most cases. Great for onboarding junior engineers. Great for quick refreshers before a meeting.
4. 🔁 Writing Repetitive Automation
Bash scripts, Python tooling, GitHub Actions workflows, Makefile targets — AI handles repetitive automation well. Give it clear input/output requirements and it usually gets the structure right.
Review it. Test it. But don't write it from scratch when you don't have to.
5. 📖 Translating Documentation
"Explain this Helm chart's values.yaml in plain English." "What does this Prometheus alerting rule actually trigger on?" "Translate this bash one-liner into something readable."
AI is a brilliant translator between dense technical docs and human understanding.
6. ✍️ Writing and Reviewing
PR descriptions, runbook drafts, incident postmortems, architecture decision records (ADRs), team documentation. AI is fast at first drafts of anything written. You edit, you own it — but the blank page problem disappears.
❌ Where AI Gets Dangerous (Slow Down Here)
1. 🔐 IAM, RBAC, and Security Configs
This is the most dangerous area across the board — whether it's AWS IAM policies, Kubernetes RBAC, GCP service account permissions, or GitHub Actions secrets.
AI defaults to the path of least resistance. That usually means overly permissive configs that make things work but open up security holes.
# AI might give you this — works, but dangerously broad
resource "aws_iam_role_policy" "example" {
policy = jsonencode({
Statement = [{
Effect = "Allow"
Action = "*"
Resource = "*"
}]
})
}
Always apply least privilege. Define exactly what actions, on exactly which resources, under exactly which conditions.
2. 📦 Deprecated Syntax and Old Versions
AI has a training cutoff. It frequently generates configs using deprecated or removed syntax — old Terraform provider versions, removed Kubernetes API versions, deprecated Docker Compose fields, outdated GitHub Actions syntax.
Always verify against the actual version you're running. Don't assume AI's output is current.
# Always check what you're actually running
terraform version
kubectl version
docker --version
3. 💾 Resource Sizing and Limits
AI guesses. It doesn't know your workload, your traffic patterns, your node sizes, or your budget. Wrong resource configs cause real production problems:
- Too low → OOMKills, CPU throttling, pod evictions
- Too high → wasted spend, noisy neighbour problems
- No limits → one misbehaving service starves everything else
Base resource decisions on actual profiling and metrics, not AI suggestions.
4. 🌐 Networking Configs
Networking is deeply environment-specific. VPC peering rules, security group configs, CNI-specific NetworkPolicies, firewall rules, load balancer annotations — what works on AWS doesn't work on GCP. What works on Calico doesn't work on Cilium.
AI gives generic networking advice. Your environment has specific constraints AI doesn't know about. Always validate against your actual cloud provider and tool docs.
5. 🔄 CI/CD Pipeline Logic
AI can write pipeline YAML. It cannot reason about your deployment strategy, your environment promotion rules, your rollback triggers, or your secret management approach.
A pipeline that deploys to prod without proper gates, approvals, or rollback logic is worse than no pipeline. Review every stage. Understand what each step does before it runs against real infrastructure.
6. 🏗️ Infrastructure State and Drift
Never let AI make decisions about existing infrastructure state. Terraform import, state manipulation, terraform destroy, resource moves — these require deep understanding of what's already running.
AI doesn't know your state file. It doesn't know what's in prod. It will confidently suggest things that will cause real drift or data loss.
🛠️ How to Actually Use AI Well in DevOps
The mental model that works: AI is a senior pair programmer who knows everything in theory but has never seen your environment.
Give It Full Context
Bad prompt:
"Write me a CI pipeline"
Good prompt:
"Write a GitLab CI pipeline for a Node.js app that builds a Docker image, pushes to GCR, and deploys to GKE using kubectl. We use environment-specific kubeconfigs stored as CI variables. Needs separate stages for build, test, and deploy. Deploy only runs on main branch."
The more context you give, the closer the output is to something you can actually use.
Always Review the Critical Fields
Before applying anything AI-generated:
✅ Versions — correct for what you're actually running?
✅ Permissions — least privilege applied?
✅ Secrets — not hardcoded anywhere?
✅ Resource limits — based on actual data?
✅ Rollback — is there a way back if this goes wrong?
✅ Blast radius — what breaks if this fails?
Use It for Explanation, Not Decision
If you're unsure about a behaviour, ask AI to explain it. Then verify in official docs or test in a non-prod environment. Use AI to understand your options. Make the decision yourself.
Test in a Safe Environment First
Before applying AI-generated configs to production, test them somewhere safe. A staging environment. A homelab. A throwaway cloud project. If it breaks something there, you learn — without the incident.
This is one of the underrated values of running a homelab. It's a consequence-free environment to test things that AI told you would work.
🚦 The Rule I Actually Follow
AI handles the 80% that's generic and repetitive. I own the 20% that's specific to my environment, my team, and my production systems.
That 20% is where incidents happen.
AI doesn't know you have a custom admission controller that rejects certain annotations. It doesn't know your VPC has a specific CIDR that conflicts with what it suggested. It doesn't know your team disabled public ECR access last month. It doesn't know your prod DB is in a different region than your app.
You know these things. AI doesn't. That's why you're still the engineer.
💬 Final Thought
I use AI every single day. It makes me faster. It handles the boring parts. It explains things clearly. It unblocks me when I'm stuck at step one of a problem I've never solved before.
But I've never let it make a production decision unsupervised.
The engineers who will get burned by AI aren't the ones who use it — it's the ones who trust it without understanding what it's doing and why.
Use it as a tool. Stay the engineer.
— Satya 👨💻 Platform Engineer | GKE | k3s | GitOps | Terraform | Using AI without being replaced by it
Member discussion