Microsoft AI-103: Microsoft Foundry Architecture

Microsoft Foundry is easiest to understand when you stop treating it as one product screen and start seeing it as the operating environment around an AI application. A real solution can include models, agents, retrieval, data connections, evaluations, identity, network controls, SDKs, monitoring, and deployment automation. AI-103 expects you to understand how those pieces work together.

The AI-103 exam gives substantial weight to planning and managing Azure AI solutions and to implementing generative AI and agents. That means architecture questions often begin before a prompt is sent and continue after a response is returned. The model is only one component inside a larger system.

A useful mental model is to divide Foundry architecture into six layers: resource and project structure, identity and network access, models, knowledge and tools, application and agent logic, and observability. If you can trace a request through those layers, most scenario questions become easier to reason about.

Start with the resource and project boundary

Foundry projects organize the assets and connections a team uses to build AI applications. The project is not merely a folder. It becomes a practical boundary for development workflows, permissions, models, agents, evaluations, and connections to external resources.

Thinking in Azure architecture terms helps because Foundry still depends on wider cloud design. Subscription structure, resource groups, regions, policy, networking, and ownership determine how the AI project is governed.

For AI-103 scenarios, look for requirements about separation between teams or environments. A development project should not automatically have the same permissions, models, data sources, or quotas as production. Good architecture makes those differences deliberate.

Identity should be designed before connections are added

An AI app can touch search indexes, storage, databases, APIs, model deployments, monitoring resources, and other services. Every connection is an access decision. If credentials are scattered through configuration files, the architecture becomes difficult to secure and rotate.

Microsoft Entra ID and Azure RBAC provide the foundation for assigning access to people and workloads. Managed identities can allow applications to authenticate without embedding long-lived secrets. Role assignments then define what that identity can actually do.

AI-103 preparation should include at least one lab where you deliberately remove a role and diagnose the resulting access failure. Security concepts become much more durable when you see exactly which principal, scope, and permission the service requires.

Models are deployed resources, not names in a dropdown

Foundry gives access to a broad model catalog, including Azure OpenAI models and models from other providers. Architecture starts with model suitability, but deployment adds another set of decisions: capacity, processing location, availability, cost model, quota, and lifecycle.

Microsoft currently supports multiple deployment types, including pay-per-token standard options, provisioned throughput, batch processing, and additional deployment choices for certain models. The exact choice should follow workload behavior. A sporadic interactive workload has different needs from a predictable high-throughput service.

The selection discipline from foundation-model evaluation still applies. Benchmark representative work before you design the whole system around one model.

A single-model design is simple, but it can be inefficient when requests vary widely in difficulty. Foundry model routing can place an optimization layer between the application and several eligible models. The router can choose a model per request based on configured behavior and the characteristics of the prompt.

This changes the architecture conversation. Instead of asking “which model is best for our application?” you may ask “which set of models and routing strategy gives us the best quality-cost balance?” That can be useful for agent systems where one step is simple classification and another needs stronger reasoning.

Routing does not eliminate evaluation. You still need to observe which model served a request, whether fallback occurred, what quality was achieved, and whether the cost behavior matches expectations.

Knowledge is a pipeline, not a checkbox

Grounded applications often depend on enterprise content. That content must be ingested, processed, indexed, filtered, secured, retrieved, and refreshed. The retrieval layer can become the dominant source of answer quality even when the model itself is strong.

Understanding retrieval-augmented generation helps you see why. The generation stage can only use the evidence it receives. Weak chunking, missing metadata, stale indexes, or poorly tuned retrieval can produce confident but wrong answers.

For AI-103, think about knowledge architecture in terms of freshness, authorization, search quality, citations, and fallback when no trustworthy evidence is found.

Tools turn a model into an application participant

Models generate; tools allow a system to act. A Foundry agent can connect to APIs, functions, data sources, and other capabilities. Tool design therefore creates a new security and reliability surface.

The architecture should make each tool narrow and explicit. Define clear inputs, predictable outputs, permission boundaries, and failure semantics. A read-only lookup deserves different controls from an operation that changes customer data or triggers an external workflow.

The same principles discussed in Model Context Protocol are useful even when a particular tool is not implemented through MCP: capabilities should be discoverable through clear contracts, but the system still needs governance over what is exposed.

Agents add state, control and orchestration

An agent is not simply a model with a system prompt. It has a goal, instructions, tools, knowledge, memory or conversation state, and a policy for deciding what to do next. In multi-step workloads, this creates a control problem as much as a generation problem.

AI agent architecture is useful for understanding the loop: observe, decide, act, inspect the result, and continue or stop. AI-103 expects you to reason about tool schemas, conversation tracking, safeguards, approval flows, and multi-agent designs.

The strongest architecture keeps autonomy bounded. Define the maximum authority and stopping conditions before you optimize how clever the agent can be.

Foundry does not remove the need for conventional application engineering. Your web app, API, worker, or service still handles user identity, input validation, state, retries, business rules, UI behavior, and integration with the rest of the system.

Serverless and event-driven Azure components can complement AI workloads well. Understanding the differences among Azure Functions, Logic Apps, and Event Grid helps you decide which work belongs inside an agent turn and which work should be delegated to deterministic or asynchronous infrastructure.

A common design mistake is forcing every workflow step through the model because the model is available. Use AI where reasoning or language understanding adds value; keep predictable operations predictable.

Evaluation and monitoring complete the architecture

A Foundry application is not production-ready until you can tell whether it is behaving correctly. Traditional telemetry covers latency, availability, errors, resource health, and cost. AI telemetry adds answer quality, safety, grounding, tool behavior, agent trajectories, and user feedback.

Azure-style monitoring and alerting gives you one part of that picture. The AI layer requires evaluation datasets and behavioral measures as well. A healthy endpoint can still produce a poor answer.

For agents, traces are especially important because the final output hides the path taken. You need to see tool calls, handoffs, retries, model choice, latency, and failures inside the workflow.

CI/CD should treat AI assets as deployable components

Prompts, agent definitions, tool configurations, application code, infrastructure, evaluation criteria, and environment settings all change over time. A mature architecture moves those changes through controlled development, testing, and production stages instead of editing the live system manually.

The broader practices of Azure DevOps apply: version what matters, automate repeatable deployment, validate before release, and preserve rollback paths. AI adds another challenge because model behavior can change even when application code does not.

That is why continuous evaluation belongs beside continuous integration. Deployment success should mean both “the resources were created” and “the solution still meets its quality and safety thresholds.”

Study Foundry by tracing one request end to end

For AI-103 preparation, build a small grounded agent and narrate every boundary. A user authenticates. The application receives a request. The agent selects a model, searches authorized knowledge, calls a tool if needed, generates a response, and emits telemetry. The result is evaluated and the system records enough evidence to diagnose problems.

Then deliberately break each layer: remove a role assignment, deploy an unsuitable model, stale the index, make a tool fail, exceed a loop limit, block a network path, or introduce a low-quality prompt. Troubleshoot from symptoms back to the layer that caused them.

Once you can see Microsoft Foundry as a connected architecture rather than a set of features, AI-103 becomes much easier to study. The exam is testing whether you can make those connections under real constraints.

AI scenarios often focus attention on models and agents, but enterprise requirements may be decided by network boundaries. A Foundry application can depend on storage, search, APIs, monitoring, and other services that must be reachable without exposing sensitive traffic unnecessarily. Private connectivity and controlled egress can therefore shape the architecture long before prompt design begins.

For study, draw the network path of your lab as carefully as the logical AI flow. Identify which components need public access, which can use private connectivity, where DNS resolution matters, and how developers reach the environment. A solution that works only because every resource is open to the internet is not a strong production design.

img