What I'd cut, what I'd keep: the actual MVP cutline
Closing the AI MVP series. What you can safely skip on day one, what you absolutely can't, and what the first 30 days of customers will teach you that nothing else can.
This is the closer for the MVP series. Eighteen pieces ago I started with a question, what does "MVP" actually mean when the value of your product is the AI doing something useful?, and we worked through it: the hybrid cloud-and-local split, the AWS-native shape, auth and multi-tenancy from day one, the secret-sauce capture loop, the three-loop product flow, prompts as code, evals, the approve-deny gate, audit, failure modes, cost, deployment, hybrid sync. A lot of ground.
Today is the part where I'd take all of that and cut it down to the version you actually ship in eight weeks with a small team. Not the version that's "good enough for now and we'll fix it later" (that version makes you cry) but the version where every piece you build is earning its keep, and the pieces you skip are the ones that genuinely don't bite until you have customers telling you they bite.
I'll do this in three parts. What I'd cut on day one. What I absolutely wouldn't. And what the first thirty days of real customer use will teach you that you cannot, no matter how clever you are, predict in advance.
What I'd cut
These are the things people build into MVPs because they feel important, and which you can almost always defer until the product has earned the right to need them.
A pretty admin UI. Whatever supervisory work the consultant has to do (review the queue, approve diagnoses, mine for patterns) can run on a stripped-down internal tool for the first hundred customers. Retool, an admin-style React page, even a couple of Postgres views and a CLI. The customer-facing surface gets the polish budget. The supervisor surface earns its polish later, when the consultant tells you which three actions they do twenty times a day and you build a button for those three things.
Anything other than email for notifications. SMS, push notifications, in-app real-time toasts, Slack integrations, webhooks for customers, all good ideas, all later. Email is universal, asynchronous, and works. SES from AWS gives you the first sixty-two thousand emails a month free, which covers a real pilot. SES, the Simple Email Service, is AWS's outbound email transport, if you want to look it up later. Build the notification layer as a single "send-event" function that today only knows how to send email. The day you need SMS, you add a path. You don't add five paths and use one.
Multi-region anything. One region, the one closest to your pilot users. The day you have a customer in Singapore complaining about latency, you have a real reason to think about a second region. Until then it's expense and complexity for a problem you don't have.
Caching layers beyond what AWS gives you. No Redis, no ElastiCache, no in-front-of-everything caching layer. Use API Gateway's caching for the obvious GETs. Use CloudFront for static assets and obvious cacheable endpoints. Lambda's own warm execution acts as a small cache. That's enough for an MVP. The day you can prove you have a hot read pattern that's costing you, you add Redis with a clear purpose. Adding it speculatively gives you a cache invalidation problem on top of all your other problems.
A microservices split. One Lambda monorepo, one CDK app, one RDS database. Two or three Lambda functions for the customer-facing API. Maybe a separate Lambda for the heavier async work, but the line is "different concurrency requirements," not "different team owns it." There is no other team. You are the team. Distributed systems problems are the most expensive problems in software, and you do not need to invite them in until they're forced on you.
Custom dashboards and BI. CloudWatch dashboards are ugly and they are sufficient. The metrics that matter (call volume, error rate, p95 latency, Bedrock spend per day, eval pass rate) all fit in a single CloudWatch dashboard you build in twenty minutes. You don't need Datadog or a custom Grafana for the first six months. When you do, you'll know exactly which ten metrics you need on it because you'll have stared at the CloudWatch one daily for six months.
Fancy A/B testing infrastructure. A feature flag library that lets you set a percentage rollout per environment is enough. LaunchDarkly is great for the day you have a real product team running real experiments. Your early A/B is one new prompt vs the old one, fifty users vs fifty users, evals telling you which won. You can do that with a feature flag and a column in the audit log. Don't buy LaunchDarkly in month one.
Most of your "what if" features. The features that the consultant brainstormed in the second discovery call but no customer has actually asked for. Cut those. Build the three things every pilot customer has asked for in the same words. The other ten things are real maybe, and you'll know which ones to build when customers tell you.
What I absolutely would not cut
There are four things where deferring them is the most expensive thing you can do. They show up in the audit logs of every team I've watched fail, marked "we should have done this from the start."
Auth and multi-tenancy. Cognito on day one, tenant ID on every row, row-level security in Postgres on day one. The day-zero version is small, maybe a hundred lines of CDK and a few hundred lines of application code. The bolt-on version, after you have customers and data and assumptions baked in, is a months-long migration that occasionally leaks one customer's data into another customer's view. Don't do that to yourself. The auth and multi-tenancy piece earlier in this series walks the cheap version in detail; that's the floor.
A real audit trail. Every meaningful decision the AI makes (and every action a human takes on top of it) written into an audit table with the actor, the action, the evidence, the timestamp, the outcome. This is not a logging concern. CloudWatch logs disappear. The audit table doesn't. The day a customer asks "why did your system reject my application?" or a regulator asks "show me how this decision was made," you point at a row, and the row has the answer. The observability and audit piece details the row shape; the bar is non-negotiable.
An eval harness. Even a small one. Even fifty golden examples and a script that runs them. The point isn't full coverage on day one; the point is a habit. The eval harness is what tells you the prompt change you just shipped didn't make things worse on the cases you already cared about. Without it, you're shipping prompt changes by vibes, and vibes regress silently. The eval harness piece describes the minimum shape, fifty examples is enough to get the discipline started.
A documented escalation path. When the AI doesn't know, when it returns a low-confidence answer, when the customer asks something off-pattern, there has to be a path that gets a human eyeball on it within a day. For a one-person shop, "human" might be the consultant themselves checking a Slack channel each morning. For a small team. It's a queue with an SLA. Either way, the path exists, the customer knows it exists, and the AI knows when to invoke it. Without an escalation path, the only failure mode is "AI gives a wrong answer and nobody notices until the customer churns." That's the failure mode that ends MVPs.
If you're reading this and your draft architecture is missing any of those four (auth, audit, evals, escalation) go put them in before you ship to your first customer. The other things on the cut list, you can add as needed. These four, you can't bolt on.
What 30 days of real customer use will teach you
Here's the part nobody can predict for you, no matter how good a brief I write or how clever your architecture is. The first thirty days of real customers using the product will teach you four things, and you will not see any of them coming.
The questions they actually ask are not the questions you built for. You designed the discovery prompt for the sales consultant assuming customers would upload transcripts. Half of them paste raw notes from memory. You designed the IT ops triage flow assuming customers would describe symptoms. Half of them paste the entire stack trace and ask "what is this?" Your retrieval, your prompts, your tone, all calibrated to inputs you guessed at. The first thirty days show you what the inputs actually look like, and the gap is always larger than you expected.
The fix is cheap if you're set up for it: capture the actual queries (PII-stripped, in your audit table), categorise them, and update your retrieval and prompt structure for what you're actually seeing. The fix is expensive if you're not set up for it, meaning, if you didn't keep the audit trail in shape and you don't have a way to safely look at customer queries, you're flying blind.
The places it fails are not the places you tested. You hammered the diagnose flow. The thing that fails is the file upload, because the consultant's customers are uploading PDFs three times the size you tested with. You stress-tested the Bedrock path. The thing that breaks is the Cognito password-reset email, because you used the default sender domain and customers' spam filters are eating it. The first thirty days expose the unsexy operational gaps, the ones that have nothing to do with the AI and everything to do with the rest of the product being a real piece of software that real strangers are using.
The fix is alarms and a habit of looking at them. Page yourself on error rate spikes. Look at the dashboard daily, really daily, not "when I think of it." Customer-facing breakage is invisible to you and visible to them.
The AI will be wrong in a way that surprises you. Not the failure modes you anticipated and planned around. A new one. The legal pro's contract-review tool will confidently say a clause is fine when it's actually missing. The financial advisor's portfolio diagnosis will recommend a rebalance based on a misread of the customer's risk profile. The medical specialist's second-opinion review will agree with the original diagnosis when the original was wrong. Whatever the new failure mode is, it will be the one your eval harness didn't have an example for, because if you'd had an example you'd have caught it in dev.
The fix isn't to prevent it (you can't prevent the unknown unknown) but to catch it within a day. Approve-deny gate stays on for the first thirty days, no exceptions. The consultant is in the loop. Patterns get added to the eval set. The eval set grows from fifty to two hundred examples in those first thirty days, every one of them a real near-miss that taught you something.
Customers will tell you what to build next, and they will be partly right. They'll ask for features. They'll ask for fields. They'll ask for integrations. About sixty percent of what they ask for will be the right thing to build, usually a smaller, more specific version of what they asked for. About forty percent will be a misdiagnosis of their underlying need, where they're describing the solution they imagined and the actual job is something else.
The skill is listening to what they're trying to do underneath the feature ask. A customer who says "I need a Salesforce integration" might actually need "I need to stop manually re-entering customer info." Those are very different things and lead to very different builds. The thirty days teach you that translation skill. There is no shortcut.
Where this hands off
The architecture you've shipped at the end of this series (small, hybrid, intentional) gets you through the first thirty days. The next series, Operating an AI product, year one, starts tomorrow. Four pieces, running daily through May 30. We pick up where this one ends: real customer load, real Bedrock model selection on real eval evidence (first piece tomorrow), the five-minute onboarding path for new tenants, the two-surface UI split between customer and consultant, and the pricing model that makes the unit economics work. Operating, not building. The architecture is settled; the question becomes how to run it well.
The MVP series was about getting the shape right. The operating series is about keeping it alive while customers actually use it.
If you've built along with this series, you have the spine. The Lambda, API Gateway, Cognito, RDS-with-pgvector, S3, Bedrock, EventBridge, SQS, CloudWatch on the cloud side. The Mac Studio with mflux, mlx-lm, whisper, and the SQS-poller-batch-runner on the local side. The hybrid sync wired through SQS, EventBridge, S3, and signed manifests. The eval harness, the audit table, the approve-deny gate, the prompt versioning. The cost model in your head. The deployment pipeline that doesn't break things.
That's an AI MVP. That's the shape. The secret sauce sitting on top, that's yours, and it's the only part of the system that's actually you. Everything else generalises across consultants, verticals, products. Architecture is the protagonist of these eighteen pieces. Your secret sauce is the protagonist of your product.
If you're shipping yours soon, my one ask: send me the URL when it's live. I want to see what you built.
Tomorrow we start operating it.