Reasoning Authority vs Production Authority in AI Infrastructure Work
A working outline for the operating-model boundary between AI agents that can reason about infrastructure and systems that hold authority to change production.
Infrastructure engineering — AWS architecture, Terraform, Kubernetes, platform patterns.
A working outline for the operating-model boundary between AI agents that can reason about infrastructure and systems that hold authority to change production.
We built a full EKS platform for a high-volume content distribution service and decided not to promote it to production. The decision came down to operational ownership, edge compute, and failure domains.
ECS Fargate egress through an inspection VPC cost ~$1,300/day. The fix: public subnets and a local IGW. The Terraform changes, apply sequence, and the gotcha.
When a Terraform module creates too much by default, the caller works against it instead of with it. The fix is a module design rule: every optional resource defaults to false. The caller opts in explicitly.
How we migrated production services from EKS to ECS Fargate by running both orchestrators simultaneously during a soak period — and why the bridge pattern is often the right call.
How to cut over DNS from one ALB to another with zero downtime: pre-provision everything first, lower TTL 48 hours early, and treat rollback as a scheduled option, not an emergency.
Changed-files detection in Terraform CI silently skips stacks that consume modified modules. Here's the pattern that actually works, plus three supporting fixes.
S3 Table Buckets encrypt, wire up IAM, and coexist with standard S3 differently than you'd expect. Five Terraform gotchas that cost real debugging hours.
ECS vs EKS isn't a features comparison — it's an operational capacity question. Decision framework including ECS Anywhere, with real production examples.
How to break the apparent IAM trust policy circular dependency in Terraform by constructing role ARNs deterministically — no two-pass apply needed.
The morning after go-live, one service was running at 99.8% CPU across 9 tasks with a 1,010ms p50 latency. Here's what the data said to do next — and why the sequence mattered.
You just merged a PR. Now you copy the URL, open Jira, paste it in a comment, change the status to Done, and update the deployed field. 5 minutes wasted, 20 times a week, every week. Here's how to eliminate it completely with GitHub Actions.
Every new Claude Code session starts cold. After 3 months of daily use across 20+ infrastructure repos, I built a three-tier context hierarchy that loads team patterns, project state, and infrastructure conventions automatically — no copy-paste, no ramp-up.
We were deploying IAM permission sets and finding the bugs in production — 5 iterations to get one permission set right. Here's the testing framework that caught 95% of permission issues before they hit prod.
S3 Table Buckets are purpose-built for Iceberg metadata operations — 10x faster than regular S3 for query planning. Here's how we built a three-zone medallion lakehouse with Terraform, including what's different about Table Buckets that most posts don't mention.
Our network topology had three standalone Transit Gateways that couldn't talk to each other. Here's how we redesigned the full platform — hub-spoke TGW with inline Network Firewall inspection — and cut over without touching a single application.
I was managing EKS add-ons with kubectl and YAML files. No version control, no drift detection, no rollback. Here's how I fixed it with a single Terraform module.
How I eliminated the 5-minute deployment tax with GitHub Actions so git push just works
Why I automated my side hustle infrastructure with Terraform instead of clicking through AWS console like a normal person
How Terraform eliminated 15 manual steps from SSL certificate validation across two domains