The cost model: what you pay before you have customers
Runway math for an AI MVP with zero customers. What the AWS free tier actually covers, what bites you, and when local pays back.
The number that matters before you have customers is not your revenue, your conversion rate, or your activation rate. You don't have any of those yet. The number that matters is how long your bank account lasts at the burn rate the architecture is quietly setting for you.
I've watched founders get into a real bind here. They picked a stack that looked free on the marketing page, shipped a working MVP, and three months in they're staring at a bill that says they're spending five hundred dollars a month to serve fourteen pilot users, twelve of whom are friends. The bill isn't huge in absolute terms. It's huge relative to the runway, and it's huge relative to the revenue, which is zero. They are paying retail to subsidise other people's curiosity.
This piece is the runway math. I'll walk through what the AWS free tier actually covers for the AI MVP shape we've been building in this series. Lambda, API Gateway, Cognito, RDS with pgvector, Bedrock, S3, CloudFront, EventBridge, SQS, CloudWatch, and where the bill suddenly stops being "rounding error" and starts being "we should talk." Then I'll show where the Mac Studio side of the split earns its keep, and where it doesn't.
If you missed the earlier piece, the hybrid split between cloud and a local rig is the framing this whole article assumes. Customer-facing stuff lives in the cloud. Training, batch, eval, and back-office image generation live on the Mac Studio. The cost story is a big reason that split exists.
The free tier, what it actually gets you
The AWS free tier is more generous than people think for the always-free chunks, and meaner than people think for the twelve-month chunks. Let me sort which is which for the services this MVP uses.
Lambda is genuinely free for ages. A million requests a month and 400,000 GB-seconds of compute, always free, no twelve-month timer. For an AI product where the customer-facing API is mostly thin handlers calling Bedrock (the model does the work, the Lambda just orchestrates) you can run a real pilot inside that envelope. A Lambda that takes 800 ms and uses 512 MB of memory is about 0.4 GB-seconds per call. That gets you roughly a million calls a month before Lambda itself charges you a cent.
API Gateway will take a bite earlier than you think. The free tier is one million REST API calls a month for the first twelve months, then it costs about $3.50 per million. It's not a lot. But if you're wiring a chatty front end that hits five endpoints to render one screen, you'll burn through the free million faster than your Lambda compute. HTTP APIs are cheaper than REST APIs, about a third the price. For an MVP I default to HTTP API and don't look back unless I need a feature that's REST-only.
Cognito is the surprise in your favour. It gives you 50,000 monthly active users (MAU) free, forever. That's an enormous amount of headroom for a pilot. If you have 50,000 active users and you're worried about the Cognito bill, you have other things to worry about, and they're good things. MAU here means anyone who logged in at all that month, if you want to look it up later.
RDS Postgres is the line item that gets people. The free tier offers 750 hours per month of a db.t4g.micro for the first twelve months (basically a single small instance running 24/7) plus 20 GB of storage. After twelve months, that small instance becomes about $12-15 a month. That's still nothing. The trap is what people pick instead. They look at the chart of instance sizes, decide they want to "leave headroom," pick a db.m6g.large, turn on Multi-AZ for "production readiness," and now they're at $250 a month for a database serving zero queries an hour.
For an MVP, start at db.t4g.micro or db.t4g.small. Don't turn on Multi-AZ until you have a customer who'd notice it being off. Use the free backups. The pgvector extension runs fine on the small instances; the cost of indexing 50,000 embeddings is not your problem at MVP scale.
S3 is cheap until egress. Storage is twenty-three cents per GB per month for standard, less for infrequent-access tiers. You can store gigabytes of model artifacts, training corpora, transcripts, and eval sets and barely notice. What costs is moving the data out. Nine cents per GB egressed to the internet, less to other AWS services in the same region. If your CloudFront cache is doing its job, S3 egress stays small. If you're serving raw S3 URLs to users, you're paying the worst version of this bill.
CloudFront is where bandwidth lives or dies. First terabyte of egress per month is free, forever. After that it's about $0.085 per GB to North America and Europe, more to Asia-Pacific. For static assets and API caching it's a good deal. For streaming video, image-heavy pages, or large model downloads, it goes up fast.
EventBridge and SQS are basically free at MVP scale. EventBridge custom events are a dollar per million. SQS standard queues give you a million requests a month free, forever. You will not notice these on the bill. Don't optimise them; optimise things that matter.
CloudWatch is the silent killer. Free tier gives you ten custom metrics, ten alarms, a million API calls, 5 GB of log ingestion, and 5 GB of log storage. Sounds like a lot. Then you turn on verbose Lambda logging across six functions, and a single bad day of debug logs ingests 8 GB and blows the free tier in one afternoon. Log ingestion is fifty cents per GB and storage is three cents per GB per month after that. Set retention on your log groups before you ship anything. Seven days for dev, thirty for prod. The default is "never expire" and it will quietly cost you.
Secrets Manager is forty cents per secret per month. Trivial. KMS keys are a dollar per key per month plus three cents per 10,000 requests. Trivial. These add up if you create a hundred of them by accident, so don't do that.
Want a deeper tour of the cloud shape these line items belong to? I sketched the box-and-line view in the AWS-native shape I actually start with. This piece is the price tag on each of those boxes.
The line item that actually bites. Bedrock
Everything I just listed is rounding error compared to the model bill.
Bedrock charges per token, in and out, and the prices vary by model. A rough mental model for May 2026: a Sonnet-class model is in the $3-per-million-input-tokens: $15-per-million-output-tokens neighbourhood. Haiku-class is roughly a tenth of that. Opus-class is roughly four times Sonnet. Tokens are the chunks the model reads and writes, a token is about three-quarters of a word, if you want to look it up later.
The shape of the bill follows the shape of the call. A typical "diagnose" call in a consultant-vertical product, say, the sales consultant's discovery framework asking the model to extract pain points from a transcript, runs maybe 5,000 tokens of input (the system prompt, the retrieved RAG context, the transcript) and 800 tokens of output (the structured pain-point list). On Sonnet that's about 1.5 cents. On Haiku it's about 0.15 cents. On Opus it's about 6 cents.
Now multiply by how often that call fires per real customer interaction. If your product flow is "user uploads a transcript, system runs three Sonnet calls to triage, diagnose, and propose actions," you're at four and a half cents per interaction. Five interactions per pilot user per week, fifty pilot users, that's about $45 a month on Bedrock alone, before anyone has paid you a dime. Manageable. Now imagine each interaction fires six calls because you went wide on retrieval, or you used Opus because "it gives better answers," or your prompt grew from 5,000 tokens to 25,000 because you started shoving the whole knowledge base into context instead of doing real retrieval. You're at $400 a month for the same fifty users.
The cost-as-a-design-input principle: every prompt you write should be paired with a per-call cost number. Not someday. The first time you ship the prompt. If a flow is too expensive, you fix it the way you'd fix a bug, with intent and a stopwatch. Reaching for a smaller model, tightening retrieval, caching where you can, and using the eval harness to prove the cheaper version is still good enough.
The marketing strategist productizing their brand-positioning method is a great example: the first version of the prompt asks Sonnet to "review this brand" and dumps in twelve pages of brand voice notes. Cost per call: thirty cents. The shipped version retrieves the three most relevant sections, runs Haiku as a router to decide which brand-positioning lens applies, then runs Sonnet on a focused 3,000-token prompt for the actual review. Cost per call: under two cents. Same output quality, the eval harness said so. Fifteen times cheaper.
When the local rig pays back
The Mac Studio side of the split is fixed cost. You buy it once, it sits in the corner, the electricity bill goes up by maybe twenty dollars a month if it's running serious workloads, and that's it. There's no per-token, no per-request, no egress fee. The cost-per-inference for a workload it can run trends to zero as you use it more.
Three workload classes earn back the rig fast.
Image generation for back-office assets. mflux on Apple Silicon will turn out marketing illustrations, blog hero images, internal slide art, and product mockups all day. The cloud equivalent (Bedrock image models or third-party APIs) can run two to ten cents per image. If you're producing a hundred images a week for marketing pages and internal use, that's $30-100 a month for something a Mac Studio does for free. mflux is a Mac-native runtime for image generation models, if you want to look it up later.
Batch inference and eval runs. When you're regression-testing prompts against your golden eval set, you might fire 500 calls in a single eval run. On Sonnet that's $7-8 every time you run evals. Run evals nightly, that's $200-250 a month. Run them locally on a smaller-but-good-enough open model (mlx-lm with a Llama or Qwen variant) and that bill is zero.
Fine-tuning the secret-sauce model. This is the big one. Training even a small model on AWS (SageMaker, Bedrock custom models, or just GPU instances) is real money. A few hundred dollars per training run, easily, when you're iterating. mlx-lm fine-tuning on a Mac Studio handles small-to-medium models on annotated examples for the cost of an evening's electricity. The trained artifact gets uploaded to S3, picked up by cloud Lambda on next cold-start, and customers get the secret-sauce model without the cloud-training bill.
For an HR consultant productizing their interview rubric, the workflow looks like: gather a few hundred annotated interview transcripts, fine-tune a small model locally to classify candidate responses against the rubric, push the model artifact to S3. Cloud-side, Lambda loads the artifact and uses it for the cheap classification step before any expensive Bedrock call fires. The fine-tune happened for the cost of running the rig overnight. The cloud equivalent of that workflow easily runs into four-figure monthly bills if you do it through managed services.
When cloud is genuinely cheaper
It's not free-versus-paid; it's a matrix. A few cases where cloud wins on cost.
Anything customer-facing with real latency requirements. Round-tripping a customer query to a Mac Studio in your house adds 200-500 ms of network latency, plus the home internet connection's variability. The customer-facing path lives in the cloud. Always. The Mac Studio is for batch and back-office.
Models you don't have hardware for. A frontier-class model (anything in the Claude Opus or GPT-5 class) won't run locally on a Mac Studio at meaningful quality and speed. If your product needs that tier of model, you're paying Bedrock or OpenAI. The mitigation is using the cheaper models where they suffice and reserving the expensive model for the calls that genuinely need it.
Spiky workloads. A Lambda that runs once a week is essentially free. Buying a server for it is silly. The cloud is excellent at "do nothing most of the time, scale when needed." Local is excellent at "always be doing something."
The runway picture
Pull all of this together for a representative MVP, fifty pilot users, a hundred Bedrock calls a day across them, modest static traffic, a Mac Studio doing back-office work, and the AWS bill ought to land somewhere in the $40-80 a month range. Most of that is Bedrock. The infrastructure pieces. Lambda, Cognito, RDS small, S3, CloudFront, EventBridge, SQS, CloudWatch with retention set, sum to maybe $20-30 of that.
The Mac Studio amortizes against eighteen-to-thirty months of cloud workloads it would have otherwise replaced, image generation, evals, fine-tuning. The electricity is rounding error.
This is the burn rate before you have a customer paying you. If your pricing has a Bedrock-cost-per-customer in mind from the start, and the pricing piece coming up in the follow-on series goes into how to do that, your unit economics survive contact with the first paying customer. If they don't, you'll find out in week two and be glad you found out cheaply.
The cost model isn't a side concern. It's a design input. Every prompt has a price. Every retrieval has a price. Every audit log line has a storage price. You don't need to obsess. You do need to know the number for each of them, the same way you know whether a function returns the right answer. If you're starting an AI product this quarter, my one ask: write down the per-call cost of your three most-used flows before you ship them. That single discipline saves the runway.