Observability and audit, not later
CloudWatch, structured logs, a real audit table, and trace IDs that follow a request through every Lambda hop and every back-office Mac Studio job. Day one, not a thing you bolt on later.
The cheapest hour I ever spent on an AI product was the one where I added a single field to a logger. The most expensive was the one eighteen months earlier when I shipped without it and spent the next nine months stitching log fragments together every time a customer asked "why did the system do that?"
That field is a trace ID. With it, observability is a query. Without it, observability is archaeology, and the dig site is your customer's loss of patience while you reconstruct what happened.
This piece is about the day-one observability shape for an AI product, the thing I now refuse to ship without. Not because I love telemetry. Because I want to answer the customer's question, the auditor's question, and my own debugging question with the same query, in seconds.
Layman version. A sales consultant has productized her discovery framework, the AI scores client conversations, suggests follow-ups, drafts summaries. Three months in, a client asks: why did your system tell my rep to ask the budget question on call two instead of call three? That's a real question with a real answer. The system did something. There's a chain of decisions: which prompt fired, which rules matched, what the model returned, who approved it. Pull that chain back in twenty seconds and you keep the customer. Can't and you've lost their trust no matter how good the actual answer was. Observability is the receipt the system hands you on its way out the door.
The two things people conflate
Observability and audit are related, share infrastructure, but answer different questions.
Observability is the engineer's question. What is the system doing right now? Where is the latency? Which Lambda is failing? Diagnostic. Audience: the team; retention: days to weeks; format: structured logs, metrics, traces.
Audit is the customer's question, the regulator's question. What was decided? By whom? When? On what evidence? Evidentiary. Audience: non-engineers, sometimes lawyers; retention: years; format: a database table you can query with SQL and hand to a non-technical person.
These share a backbone (the trace ID, the structured-log discipline) but they should be separate stores. The observability store rotates and gets pruned. The audit store does not. The observability store holds things you can lose in a retention policy. The audit store holds things you'd be in serious trouble for losing.
Build both. They're cheap on day one. Unrecoverable later.
The trace ID, which is everything
Every customer-facing API request gets a trace ID at the edge, generated at the API Gateway authorizer, attached to request context, propagated to every downstream call. The format is a UUID, because UUIDs are unique without coordination and easy to grep for.
It rides through every hop. The handling Lambda adds it as a structured-log field on every line. Downstream Lambda invocations carry it in the message envelope. Work queued for the Mac Studio side carries it in the SQS message; the local worker logs everything under the same ID. Results pushed back to the cloud carry it through to S3, Postgres, and the audit row.
Done consistently, a CloudWatch Logs Insights query takes ten seconds to write and returns the entire timeline of a customer's request (across every service, every queue, every back-office worker) in chronological order. The query is short because the discipline was long: every log line on every machine includes the trace ID. No exceptions.
Surface the ID in customer-facing error messages (reference ID abc123-...), when a customer reports a problem the first thing they paste is the exact key you need. Surface it in the consultant's queue UI too. Trivial cost, huge leverage.
Structured logs, not stringified essays
Every log line is a JSON object: trace ID, timestamp, service, operation, tenant ID, user ID, and event-specific data. The message field is a short label, "model_call_complete", "retrieval_failed", "approval_recorded", not a paragraph.
Logs are queries. The moment your logs are free text, you've lost the ability to filter, group, and combine. "Give me every model_call_complete event for tenant X in the last hour grouped by latency bucket" is one Insights query with structured logs. Without them. It's a person and an afternoon.
For a medical specialist running a second-opinion review service, this matters more than usual. When something looks off, the structured log lets you reconstruct the exact prompt fired, the retrieval results returned, the model output, the reviewer action. Stringified logs are stories. Structured logs are evidence.
I lean on a small library every Lambda imports, fifty lines of code wrapping the standard logger, baking in the trace ID and tenant context automatically. It's the most-used import in the codebase. Build the equivalent on day one. Don't let anyone log strings.
The audit table
Separately from the observability stream, a Postgres table whose purpose is to record decisions. Not events. Decisions.
The schema is boring. Columns: primary key, trace ID, tenant, actor (user ID for humans, rule ID for auto-resolved), action type, subject, evidence (JSON blob of what the AI saw, retrieved docs, model output, confidence, rationale), decision, timestamp. A "supersedes" column for when a decision gets reversed by a later one; the chain is preserved.
This is the table you point at when someone asks why did your system tell my rep to ask the budget question on call two? Query by tenant and time range, find the row, read the evidence column. Human-approved? Actor field has the user ID. Auto-resolved? Actor field has the rule ID and version. Need more depth? Pivot on the trace ID into the observability logs.
The audit table doesn't get pruned because storage is cheap and regret is expensive. Backups are real backups, and encryption at rest is on by default. KMS, not "we'll get to it." Schema migrations are themselves audit events with a reason.
The most important rule: every state-changing API has to write to the audit table before returning success. If the audit write fails, the API returns an error and the action doesn't happen. This is harsh on purpose. The day you let an action happen without an audit row is the day your audit story has a permanent hole, and you won't know which day it was until a regulator finds the missing row.
The hybrid wrinkle. Mac Studio side
The cloud half is one trace continuum. The Mac Studio half, the back-office worker draining SQS, running whisper transcriptions, running eval batches, fine-tuning the small model, is a separate machine, and observability across the boundary is what teams fail at most.
The fix is the same fix: trace ID rides with the work. SQS messages carry the originating trace ID, and the local worker logs every step under it. Result manifests pushed to S3 carry the ID, and EventBridge events back into the cloud carry it. The downstream Lambda that updates the audit table writes a row with the ID still attached.
The local worker ships its logs back to CloudWatch via a small daemon, under the same log-group conventions the cloud Lambdas use. The Insights query that returns the request timeline includes the local worker's hops as if it were just another Lambda. The customer doesn't know there's a Mac in a closet. The trace doesn't either.
Want to go deeper on the cloud-local mechanics? The wiring is in the hybrid sync pattern. The point here: the trace ID and the structured-log discipline cross the boundary unchanged. If your observability falls apart at the edge of your network, you have a cloud monitor and a separate local monitor and a habit of swiveling between them.
CloudWatch and what it isn't
CloudWatch is the default for AWS-native systems and a fine one. Logs flow there from every Lambda. Metrics flow from API Gateway, Bedrock, SQS without you doing anything. Insights queries are fast at small scale.
What CloudWatch is not is the audit store. Its retention is rarely "forever." Its query model is awkward for pulling a specific decision and showing it to a non-engineer. Its access controls are coarser than you want for "legal can read these rows but not those." Use CloudWatch for the observability stream and Postgres for the decision record. Engineering surface vs governance surface. Same trace ID, different stores.
Dashboards on day one are small and pointed: request rate by endpoint and tenant; model-call latency distribution; model-call cost per tenant per day; error rate by service; SQS queue depth; local-worker heartbeat; eval-suite pass rate over time. Seven dashboards, single screens, readable in fifteen seconds each. Premature dashboards are a way to feel observant without being observant.
The parts that bite
PII in logs is a permanent problem if you don't head it off on day one. Customer queries contain personal data, financial details, regulated content. The structured logger has to know what fields to scrub before serialization, and the scrub list is in version-controlled config. Once PII is in CloudWatch, getting it out is messy. Don't put it there.
The trace ID has to survive serialization round-trips. SQS messages, EventBridge events, S3 metadata, Postgres JSON columns. Every one is a place a sloppy serialization can drop the field. Test the round-trip with an integration test. Ten minutes to write, saves a debugging session.
Audit writes have to be idempotent. Lambdas retry. SQS at-least-once means duplicate processing. Use an upsert keyed on (trace ID, action type, subject) so a retry doesn't double-record. Otherwise the count of "decisions made" diverges from the actual count, and the report you give the consultant is silently wrong.
Structured-log fields drift. latency_ms, latencyMs, latency all end up in the same log group. Lock the field names with a schema and lint for them.
Day one, not later
You don't build observability later because the decision points where it would have been cheap (how to log, what fields are universal, where the trace ID gets generated) pass quietly during the first week. By the time you wish you had them, retrofitting means touching every service, every Lambda, every worker. You'll do most of the work but not all of it, and the gaps will bite you on the day you least want to be bitten.
The audit table gets harder to add in proportion to traffic handled without it. On day one the table is empty and the schema is your decision. On day three hundred it's full of rows you wish were structured differently.
So: day one. Trace ID at the edge. Structured logs everywhere. A Postgres audit table every state-changing API writes to before returning. Both layers crossing the cloud-local boundary unchanged. Seven CloudWatch dashboards. PII scrub list locked. Retry idempotency. The whole pile, in the first week, before there's anything to observe. The day a customer asks the question (and they will) you want the answer to be a query, not a dig.