Microsoft AB-100: Azure OpenAI Models and Deployment
An AB-100 architect should not begin model strategy by asking which model is most capable. The starting point is the business workload: what decisions the solution supports, what level of quality is required, how much latency users tolerate, where data may be processed, how much the interaction is worth, and what operational controls the organization can sustain.
The AB-100 exam expects solution-level reasoning across AI-powered business applications. As of October 3, 2026, the current exam still uses the pre–October 14 blueprint. Microsoft has published the upcoming update, but candidates today should not treat those future changes as already effective.
Model architecture is therefore best studied as a set of business decisions rather than a list of Foundry deployment features.
Different AI interactions have different value. A model call that classifies an internal request may be worth fractions of a cent. A model that helps resolve a high-value customer case can justify more expensive reasoning. An autonomous action that changes a business record may need extra validation even if the model itself is cheap.
This is why model strategy belongs inside architecture. Total cost includes inference, retrieval, data preparation, integration, monitoring, evaluation, human review, and support. A technically impressive model can still be the wrong choice if it makes the operating model uneconomic.
Start with the unit economics of the process, then set model cost and latency budgets within that boundary.
A solution architect should insist on workload-specific evidence before standardizing on a model. Public benchmarks are useful context but do not capture company terminology, retrieval quality, tool schemas, or the consequences of failure.
The approach in foundation-model evaluation is directly applicable: define representative tasks, expected outputs or rubrics, safety criteria, and acceptance thresholds. The model that clears those thresholds with acceptable cost and latency is a viable candidate.
Without an evaluation standard, model selection becomes preference rather than architecture.
Foundry model routing lets a solution use different underlying models without hard-coding every routing rule in the application. That can support a tiered design in which simple work is handled efficiently while complex work receives stronger reasoning.
The business architecture question is whether the variability is acceptable. Some regulated or highly deterministic processes may require a fixed model and validated version. Other workloads may benefit from dynamic selection and automatic fallback.
The broader principle from agentic operations applies: intelligence should be allocated according to task complexity and risk, not uniformly across every step.
If the platform chooses among models, operations teams need evidence about those choices. Which model served the request? Did fallback occur? Was the more expensive model used because the prompt was difficult or because another deployment was unavailable? Did quality improve enough to justify the additional cost?
An architecture should capture routing metadata alongside latency, token usage, error rates, quality scores, and business outcomes. Otherwise the organization cannot explain sudden cost changes or investigate why similar requests produced different behavior.
Routing is valuable when it is measurable. Hidden routing behavior creates governance problems.
Foundry supports deployment choices with global, data-zone, and regional processing characteristics. An architect must translate legal, contractual, and organizational requirements into a processing boundary before the development team selects a deployment.
A global deployment can offer broad capacity flexibility, but it may not satisfy a workload that must process data in a specified boundary. A data-zone or regional design can be more appropriate when residency requirements dominate.
This is an example of why cloud architecture knowledge remains important in AI work. Model quality does not override data-governance requirements.
Pay-per-token standard deployments suit many variable workloads because cost follows usage. Provisioned throughput can make sense when demand is predictable and the organization needs reserved capacity or more consistent performance. Batch options fit workloads that can wait.
The architect should ask how traffic behaves across the day, how large prompts and outputs are, what peak concurrency looks like, and how painful throttling would be. Those answers matter more than the prestige of the model itself.
Capacity strategy is part of user experience. A brilliant model that cannot answer during a peak business hour is not a successful architecture.
Multi-agent solutions create an opportunity to specialize model choice. A coordinator that interprets ambiguous goals may need stronger reasoning. A retrieval worker or structured extractor may perform well with a faster, less expensive model. A safety-critical action may use a separate verification step.
This aligns with the architecture of AI agents: the value comes from combining roles and tools, not from forcing one model to perform every task identically.
AB-100 candidates should be ready to justify the additional operational complexity. Model specialization is useful only when the quality or cost gain exceeds the burden of managing more components.
When the business problem depends on proprietary knowledge, stronger retrieval can be more valuable than a stronger base model. A well-designed retrieval-augmented generation layer can supply current evidence, reduce unsupported answers, and make responses traceable to enterprise sources.
The architect should therefore evaluate retrieval and generation separately. If the system lacks relevant evidence, changing models may only change the style of the wrong answer.
Model strategy is inseparable from data strategy.
Organizations need consistent rules for who can deploy models, who can invoke them, how application identities are authenticated, and which teams can change model configurations. Ad hoc keys and shared administrative access are difficult to audit.
Entra ID and Azure RBAC provide a strong basis for Azure-side permissions. The architecture should also define ownership for quotas, cost centers, model approvals, and exceptions.
These controls become more important when a central Foundry environment serves several business units.
Models are versioned dependencies with their own retirement and upgrade schedules. A production solution needs an owner who watches lifecycle notices, tests replacements, and coordinates changes before deadlines force an emergency migration.
Use the release discipline from DevOps: maintain a test suite, stage upgrades, compare quality and performance, monitor after release, and define rollback criteria.
A newer model can change tool behavior, response structure, safety characteristics, token usage, or latency. “Newer” is not a sufficient validation result.
Some workloads need more than content filtering. They may require explanations, human review, confidence thresholds, restricted tools, or a model that performs reliably under a particular safety rubric.
Responsible AI practices should therefore influence the model and workflow design. If the process is consequential, a slower path with stronger validation may be more appropriate than the lowest-latency option.
Safety requirements are architecture constraints, not values added after deployment.
For AB-100 preparation, imagine you are defining model policy for an enterprise rather than choosing one endpoint for one application. Specify approved model categories, evaluation requirements, processing-location rules, cost thresholds, routing conditions, monitoring requirements, and lifecycle ownership.
Then test the policy against three workloads: a high-volume internal assistant, a customer-facing agent, and a low-volume high-risk decision-support process. The correct model and deployment strategy should differ because the business constraints differ.
That is the level of reasoning AB-100 is trying to develop. Model strategy is not a static technology selection. It is a repeatable method for making defensible decisions as models, workloads, and organizational requirements change.
An enterprise model policy should be consistent, but not rigid. A low-risk internal summarization tool and a customer-facing autonomous agent should not be forced into identical deployment, evaluation, and approval rules. The architecture should define risk tiers and the additional controls required as impact increases.
For example, a low-risk workload might use an approved standard deployment with basic evaluation and routine monitoring. A higher-risk workload might require a fixed model version, stricter data-processing location, a documented human-approval path, deeper evaluation, and explicit change review before model upgrades.
Creating these tiers makes exceptions governable. Teams can move quickly inside an approved pattern without turning every project into a one-off security discussion.
Foundry now exposes models from several providers, which is useful but can tempt application teams to bind business logic to provider-specific behavior. Architects should identify where portability is worth preserving. Stable application contracts, structured outputs, abstracted model clients, and isolated prompt assets can reduce switching cost.
Not every feature should be abstracted to the lowest common denominator. A specialized model capability may justify deliberate coupling. The important point is to make that coupling visible and document why the business value outweighs the migration cost.
Architects should also decide who is allowed to approve a model exception. Central review for high-risk deviations keeps local teams from quietly expanding the approved model surface without security, compliance, or cost visibility.
Model strategy also needs a retirement process. When a workload is abandoned or replaced, remove unused deployments, revoke access, archive evaluation evidence, and update the approved inventory. Stale endpoints create cost and security exposure, especially when nobody remembers which application or team still depends on them.