Deployment: IaC, CI/CD, environments, the minimum shape
The smallest setup that lets you ship an AI MVP without breaking things. CDK, GitHub Actions, three environments, migrations and secrets handled honestly.
There's a thing that happens about three weeks into shipping an AI MVP. The product works. The prompt is dialled in. The first pilot users have started clicking around. And you've quietly accumulated a small zoo of resources in the AWS console that you created by hand, in no particular order, while you were trying to get the thing to work. There's a Lambda you can't remember configuring. There's an RDS parameter group with a name like default-pg15-2. There's an S3 bucket whose lifecycle rule you set up at midnight one Tuesday.
This is the part where most MVPs either get on top of their deployment story or get crushed by it. Not because the AWS console is bad. Because the second environment is when the bill comes due. The day you want a staging environment that mirrors production, you discover that you don't actually know what production looks like, you only know what's currently running, which is not the same thing.
The fix is boring and the fix is well-known: infrastructure as code, version-controlled, deployed by a pipeline. The interesting question is how small you can make that setup and still ship safely. This piece is the minimum I'd actually do for an AI MVP. Not the gold-plated version. Not the "what FAANG does" version. The version that fits a one-person or three-person team and a real budget.
If you've been following the architecture spine through this series. Lambda, API Gateway, Cognito, RDS with pgvector, S3, Bedrock, EventBridge, SQS, CloudWatch, and a Mac Studio on the local side, this piece is how you actually wrap a build-and-deploy pipeline around it.
Pick one tool, and pick CDK
Terraform and CDK both work. They both produce reproducible AWS infrastructure. They both have warts. The argument over which is "better" has been running for years and will run for more. For an AWS-native MVP, pick CDK and stop debating.
The reason isn't religious. It's that CDK lets you write infrastructure in TypeScript (or Python, but I default to TypeScript here), which means your Lambda code and your infrastructure code share a language, a linter, a test runner, and a type system. The Constructs library (the standard CDK pre-built bundles) removes a stack of boilerplate. A RestApi construct wires API Gateway, CloudWatch logging, throttling, and stages in twelve lines. A DatabaseCluster construct wires RDS, security groups, subnet groups, and parameter groups in twenty. CDK stands for Cloud Development Kit; it's AWS's official "write infra as code in a real programming language" tool, if you want to look it up later.
You also get the synth-then-apply rhythm, cdk synth outputs the CloudFormation, cdk diff shows you what's about to change, cdk deploy applies it. The diff step is the single most valuable habit you can build. Read the diff every time. The day you skip reading it is the day you delete a production database.
Terraform isn't worse, exactly. It's broader, it works across clouds, has a stronger community in multi-cloud shops, and the state-file model is more transparent. If your team already knows Terraform cold, use Terraform. If you're choosing fresh and you're on AWS, CDK pays back faster because the Constructs library is genuinely good and the type safety catches mistakes the YAML world doesn't.
Pick one. Don't run both. Don't half-CDK-half-console. The split makes the pipeline brittle and the audit trail useless.
Three environments, the way they actually need to differ
The standard answer is dev / staging / prod, and that's right, but the standard explanation undersells how different they need to be from each other. Let me walk through what they're really for.
Dev is one developer's playground. Every developer should be able to spin up a personal copy (cdk deploy --context env=dev-sid, for example) and tear it down without affecting anyone else. The data is fake. The Bedrock calls are real but cheap (point at Haiku models, low rate limits). RDS is a tiny instance. There's no Multi-AZ. The point is step-by-step speed, not durability.
Staging is one shared environment that mirrors production's shape. Same instance sizes (cheaper tier if you must, but same topology). Same secrets pattern. Same observability wiring. The difference is the data (staging gets synthetic data or anonymised production samples) and the customers, of whom there are zero. Staging exists so the CI pipeline can deploy to it, run the eval harness against it, run integration tests against it, and let humans poke at the actual UI. Staging is the thing CI breaks if it's going to break.
Prod is what customers touch. Stricter alarms. Full backups. Whatever Multi-AZ or read replicas you've decided you need. Locked down IAM. The blast radius matters here in a way it doesn't in dev or staging.
The shape of the differences matters more than the count of environments. I've seen teams run six environments where they were all subtly different from each other in ways nobody documented. Three environments that are deliberately the same in everything that matters beats six that drift.
The CI/CD pipeline. GitHub Actions, three jobs, no drama
The pipeline I land on for this kind of MVP has three GitHub Actions workflows.
pr.yml runs on every pull request. Lint, unit tests, cdk synth, cdk diff against staging. The diff gets posted as a PR comment so reviewers see exactly what infrastructure is about to change. No deploy happens.
deploy-staging.yml runs on merge to main. It deploys CDK to staging, runs the eval harness against staging Bedrock endpoints, runs integration tests, and if everything passes, tags the commit staging-passed. The eval harness is part of the gate, not a separate concern. A regression in model quality is a regression. You can read more about that bit in the eval harness piece.
deploy-prod.yml runs on manual trigger, a workflow_dispatch with a commit SHA. It only accepts SHAs that carry the staging-passed tag. It deploys, runs a smoke test, and pings a Slack channel. Manual trigger because production deploys should be a person saying "yes, do it now," not a side effect of a merge.
That's the whole CI/CD shape. Three workflows. The most complex one is maybe two hundred lines of YAML. Don't build the elaborate version on day one. Most of the hard problems people solve in CI/CD (canary deployments, automated rollback, blue/green) are solving problems you don't have yet at MVP scale. Add them when you feel the pain.
Want to see how the deploy gate interacts with prompt versioning? The prompt-versioning approach in prompts as code and the audit story in observability and audit, not later both depend on the deploy pipeline being able to roll back cleanly. They're sibling concerns to this one.
Migrations on RDS, the part everyone hand-waves
The CDK side of RDS is easy. The migrations side is where I see MVPs hurt themselves.
Here's the rule I'd ask of any team: migrations are not part of the CDK deploy. They are a separate step, run as a job, with its own logging and its own retry semantics. CDK builds the database; a migration tool changes the schema inside it.
I default to flyway or migrate (the Go one) for this, pick whichever your team already knows. The migrations live in db/migrations/ in the same repo, numbered sequentially, with up-only files (down-migrations look great in slides and ruin you in production). The CI pipeline has a separate job, migrate-staging, that runs after the CDK staging deploy but before the eval harness, so the eval runs against the new schema. The same shape exists for prod: migrate-prod runs before the prod deploy is considered done.
The reason migrations need to be their own job is that they have failure modes the CDK deploy doesn't. They can deadlock against running queries. They can run for forty minutes on a big table. They can succeed but leave the application's data in a state nobody expected. Wrapping all of that inside cdk deploy makes the failure mode opaque and the rollback impossible. Pulling it out gives you a job you can rerun, monitor, and reason about independently.
For an IT ops consultant productizing their triage tree, the kind of migration that gets you is something like "add a severity column to the tickets table." It's three lines of SQL. The thing that ruins you isn't the SQL, it's that the new application code expects the column to exist, gets deployed before the migration runs, and starts throwing 500s on every customer query. The fix is the boring rule: migration runs first, app code runs second. Bake the order into the pipeline.
Secrets, three places, one pattern
Every environment has the same kind of secrets: database passwords, Bedrock API quotas, third-party API keys, signing keys for the audit trail, SMTP credentials for transactional email. The MVP-grade answer is Secrets Manager, with a clean naming convention.
The naming convention I use is eotm/<env>/<service>/<key>, so eotm/prod/rds/admin-password, eotm/staging/bedrock/throttle-config, and so on. CDK creates the secret resources; the actual values get rotated through Secrets Manager directly, never through CDK and never through a checked-in file. Secrets Manager is AWS's hosted store for sensitive config values; it integrates with KMS for encryption and IAM for access control, if you want to look it up later.
The application Lambdas reference secrets by ARN, not by value. The value is fetched at cold-start (cached for the warm-start window) using the AWS SDK. The Lambda's execution role has IAM permission to read only the secrets for its environment. A prod Lambda cannot read staging secrets, full stop. This separation matters more than people think; the day someone runs a staging test that hits a prod resource by mistake is the day this boundary saves you.
For local development, the pattern is the same with one substitution: developers use named AWS profiles to assume a dev role, the application Lambdas (when run locally via sam local or equivalent) fetch the dev secrets, and nobody (ever) copies a secret value into a .env file in the repo. If you need a local-only value for a workflow that runs entirely offline (say, a Mac Studio job for the legal pro auto-reviewing contracts against their playbook), it lives in the Mac Studio's local keychain, not the repo.
The cross-environment story matters: every secret used in prod has a staging twin, and the staging twin is what dev work points at. A new secret added in prod without a staging twin is a deployment that will break when it goes through the pipeline. Make the absence of a staging twin a CI failure. It's a six-line check; it saves you a 2 a.m. incident.
What this gives you
Land all this and you've got a setup that does what it needs to: one CDK app, three environments with deliberate parity, a three-workflow GitHub Actions pipeline that runs evals as a deploy gate, RDS migrations run as their own auditable job, secrets in Secrets Manager with a clean per-environment IAM split.
You don't have canaries. You don't have automated rollback. You don't have a multi-region failover. You don't have blue/green. That's fine. You can ship without those for a long while. What you do have is the smallest setup that lets you change infrastructure on purpose, deploy on purpose, and trace what was deployed when, which is the bar an MVP actually needs to clear.
The day you outgrow this is the day you have ten customers, the eval harness is catching real regressions, and you start to want a canary because rolling back is too slow. That's a good problem. The setup above is the foundation you'd extend toward that, not throw away.
If you're starting today: pick CDK, write three GitHub Actions workflows, treat your migrations as a first-class job, put your secrets in Secrets Manager with a per-environment naming scheme. That's the minimum shape. It's not much. It's enough.