AIP-C01 Production Capstone: Build a Governed Support Assistant
A manufacturing support team spends too long finding current repair instructions, matching them to the correct machine and preparing service requests. A generative-AI assistant could shorten that work, but the business cannot tolerate the model inventing an approved repair, exposing another customer’s equipment history or creating duplicate replacement orders.
This AIP-C01 capstone turns the official AWS domains into one defensible production design. The intended outcome is not a chatbot demonstration: it is an authorized, observable workflow whose technical and business state can be tested. All identities, machines and records below are fictional, and any practical exercise belongs in an explicitly approved AWS account with controlled costs.
A technician enters a machine serial number and asks whether an observed fault is covered by the current service bulletin. A read-only response can cite the approved troubleshooting instructions. A more consequential request might create a replacement-part proposal, but the inventory service must check stock and the customer agreement before anything is committed. Some faults require escalation to an authorized engineer.
Agree on success criteria before configuring a model: correct machine match, authorized document selection, supported citation, safe response when no current procedure exists, confirmed ticket state, latency for interactive work and an owner for escalation. If a process is fully deterministic, a conventional API route may be more reliable than agent planning.
A release should fail if the assistant invents a service authorization or attempts an out-of-scope order, even when other responses sound helpful.
Approved repair bulletins may be stored in S3 and ingested into Amazon Bedrock Knowledge Bases. Machine ownership, warranty status and replacement inventory should remain in authoritative transactional services reached through validated APIs. The model can explain those facts but cannot become their source of truth.
Give each source a stable identifier, data owner, classification and update expectation. A newly uploaded draft bulletin should not become active merely because it has the newest timestamp. The ingestion process should select approved policy versions; metadata can identify product series, effective date and jurisdiction.
For retrieval, compare semantic search with a hybrid strategy when exact machine codes matter. A cited passage that describes a neighboring model may be relevant language but an incorrect technical instruction. Verify the product identifier before selecting the final clause.
The security test should include permission revocation, not just two users who have never shared access. Suppose a technician changes teams after receiving an earlier ticket summary. Their old conversation may still contain that customer ID, and a cache may remember previously allowed source records. The next API request must be evaluated against current access, not the permissions present when the session began. Define how the application expires authorization decisions and which downstream service owns revocation. A content filter can reduce unsafe text, but the absence of restricted documents from retrieval is what proves the customer boundary held.
A field technician may service customer A but not customer B. The application should derive that scope from authenticated context and enforce it before returning protected machine records or customer-specific bulletins. The AWS IAM role used by an integration controls its access to AWS services, but that role does not automatically encode every technician-to-customer relationship.
A narrow retrieval or tool API can check this relationship on each request. Amazon Bedrock Guardrails can help with supported input and output safety controls; they are not a replacement for customer-level authorization or the transactional system’s approval rules.
Test a permitted request and the same request for another customer’s serial number. If a restricted record enters model context before the response is filtered, the architecture has already failed an important boundary.
Use a representative pilot dataset to compare supported Bedrock models for factual answer quality, supported tool behavior, latency, output constraints and total operating cost. A general benchmark cannot establish whether a model reliably extracts the right machine variant or preserves a critical repair warning.
Define an answer contract that distinguishes verified bulletin facts, current warranty information, an unresolved question and a proposed action. A structured response can help the application verify required fields before presenting it to technicians. If a source is unavailable, the model should explain the limitation rather than invent a document citation.
Include a documented fallback route only if the alternate model meets the same security, data-location and quality criteria. Failure to reach an approved model can be handled as a controlled unresolved task.
A service-ticket tool can take a validated machine ID, diagnosed symptom, requesting technician and a stable operation ID. The service records whether the request was accepted, rejected, pending review or committed. A model-generated sentence such as “the replacement was ordered” is not proof that any order exists.
For a replacement-part request, require the inventory or approval service to check entitlement and policy before changing state. A timeout after the tool commits must not lead to an unguarded duplicate request. Use an idempotency key or equivalent backend rule, and provide a status lookup for uncertain outcomes.
AgentCore or another supported orchestration layer may coordinate the work, but it should be judged by how reliably it preserves the business contract, not by the number of tools or agents involved.
A short bulletin explanation is interactive. A nightly synthesis of many resolved tickets is asynchronous. The application can route them differently: a supported Bedrock inference path for live queries, and a queued job with appropriate workers for batch work. An API boundary validates callers and limits request complexity before model invocation.
Track token usage and response latency, but also the percentage of cases correctly resolved without recontact. Retrieval, repeated model calls, safety checks and downstream tool retries all contribute to cost. A cheaper model that produces more wrong machine matches may increase the real cost of service.
Separate private content from general cache entries. A cached response generated for one customer’s machine should not be served to another user without the same authorization check and an appropriate invalidation policy.
Use a fixed set of test records with an approved source revision, a stale revision, a customer who owns two similar machines and another customer whose records should never be visible to that user. Save both the expected retrieval IDs and the expected final business states. A test that confirms only the generated answer can miss an authorization problem if the model saw forbidden data and happened not to repeat it. Conversely, a retrieval result can be correct while the model generates a misleading service recommendation. Reviewing these conditions separately points to the component that actually needs repair and avoids accidental improvements on one metric at the expense of a more important control.
Test the workflow against outcomes the business actually cares about:
| Scenario | Required behavior | Proof |
|---|---|---|
| Current approved bulletin | Accurate answer for the exact machine series | Correct source ID and policy version |
| Superseded repair instruction | Use current content or request review | Retrieval results exclude irrelevant old version |
| Other customer’s equipment | Deny restricted retrieval | No unauthorized record reaches model context |
| Replacement requires approval | Return pending status; no immediate order | Business system shows only a review request |
| Tool times out after commit | Reconcile without duplicate action | Exactly one durable operation ID |
| Model route is unavailable | Use approved fallback or safe escalation | Verified data location, quality and user message |
Evaluate retrieval and generation separately from tool execution. An answer that sounds correct but selects the wrong source or changes the wrong case fails the test. Calibration with human reviewers matters for ambiguity; machine-generated evaluation scores should not silently overrule a business requirement.
Every accepted request needs a correlation ID that connects authentication, retrieval, model invocation, tool calls and confirmed business state. CloudWatch and application tracing can identify latency, throttling, errors and unusual tool sequences, but they must be configured to avoid unnecessary exposure of personal customer or machine data.
A meaningful alert could indicate a spike in unavailable bulletin lookups or uncommitted replacement requests rather than only elevated model latency. Name the team that will respond. A trace should tell them whether the failure belongs to source ingestion, authorization, model inference or the business API.
When incidents lead to new regression cases, add those cases to the approved evaluation dataset. The production system should improve through verified failures, not only by editing the prompt when users complain.
One release drill should deliberately change the tool schema before the agent has been updated. The request may still reach the backend but fail validation, or the model might fill an absent field with an invented default. Catch the mismatch in the validation environment and document the deployment order that prevents it. Rehearse a second change in which the policy source updates while the model stays the same. If the correct answer changes in a way the business expects, that is a healthy data-driven update; if unrelated answers drift, investigate retrieval and context assembly. These are distinct regression cases with distinct rollback paths, which is why the entire support assistant must be managed as a versioned application rather than a prompt stored in isolation.
Version the prompt, selected model, knowledge configuration, tool schema, IAM permissions and application code as a coherent release. Deploy in a test environment with synthetic tickets and run permitted/denied tests under the actual technician persona. Do not verify only through an administrator whose privileges hide data-access defects.
During controlled rollout, compare user outcomes with baseline service quality, log failures and watch for unexpected order creation or data exposure. An unsafe tool should be disableable without necessarily taking the entire read-only assistant offline.
Rolling back an agent configuration does not undo a replacement order already committed. Keep an explicit reconciliation and customer-notification process for any affected business transactions.
A reviewer should receive the source map, permissions, business rules, model decision, tool contract, evaluation ledger, failure handling and named operational owner. Ask what happens if a bulletin is corrected, a technician changes accounts, a model route becomes unavailable or an approved part order’s acknowledgment is lost.
If the design cannot answer those questions, it is not ready for a production customer. AIP-C01 professional judgment is the ability to integrate AWS managed services with ordinary software-engineering controls so the system gives supported answers, acts only within authority and can recover when parts of the workflow fail.