Agentic AI on Microsoft: From Agents to Systems
Microsoft Foundry Agent Service is Microsoft’s current managed platform for building, deploying, scaling, and governing AI agents. It supports prompt agents for low-operations configurations, hosted agents for full-code agent logic, model and tool integration, tracing, evaluation, and enterprise identity. The AI-103 exam is the current Microsoft role target for developers building AI apps and agents on Azure.
The important shift is from building one chatbot to operating a system of agents, tools, identities, data sources, policies, evaluations, and business actions. Agentic architecture is mostly about controlling that system as capability and autonomy increase.
Foundry prompt agents are defined through instructions, model choice, and tools while Microsoft manages the runtime and scaling.
They are useful for internal assistants, business automations, and production agents that do not require a custom orchestration loop.
Because the behavior is mostly configuration, teams can version and review agent definitions without maintaining a separate hosting stack.
Start here when the business requirement is clear and custom code would add infrastructure without adding useful control.
Managed agent definitions can be created interactively and then moved into code-first, versioned deployment as the solution matures.
This gives teams a useful progression from rapid experimentation to reviewed production configuration without forcing an early commitment to a custom hosting architecture.
Hosted agents let developers bring code and frameworks such as Microsoft Agent Framework, LangGraph, or other supported agent SDKs while Foundry manages the container runtime, endpoint, scaling, session state, and agent identity.
Use hosted agents for custom orchestration, specialized libraries, multi-agent systems, custom protocols, or application logic that cannot be expressed through a prompt-agent definition.
Full code expands control and expands the testing and support surface.
The architecture should justify that extra ownership instead of choosing custom orchestration by default.
Hosted agents receive a dedicated Entra identity, which is important when each agent needs distinct permission to models, data, and downstream business systems.
Treat that identity like any other workload identity: apply least privilege, rotate or avoid secrets, monitor use, and remove access when the agent is retired.
Microsoft Agent Framework provides a consistent programming model for agents and workflows while integrating with Foundry-hosted or remote agent services.
Applications can also call Foundry model inference directly when the application itself owns the agent definition and tool loop.
That creates a useful architecture choice: service-managed agent definition, hosted agent code, or self-managed orchestration using Foundry models and tools.
Choose based on ownership, lifecycle, scaling, and integration requirements rather than library preference.
Durable or long-running workflows may need state and retry behavior beyond a simple request-response agent. Microsoft’s current agent framework ecosystem includes durable execution options and several hosting protocols.
The architecture should choose protocol, hosting, and orchestration separately so one client-integration requirement does not unnecessarily dictate the entire runtime.
Protocol choice should also follow the client. A web application, another remote agent, Microsoft 365 channel, or custom service may need different interfaces even when the agent logic is the same. Microsoft’s current platform separates hosting from client protocol, which helps teams avoid rebuilding the agent merely because a new consumer needs A2A, Responses, or another supported interface.
Foundry agents can use platform tools, custom functions, MCP servers, enterprise data connectors, and other actions.
Tool contracts should expose specific business capabilities rather than unrestricted shells, databases, or broad administrator APIs.
A user’s ability to ask for an action should not automatically grant the agent permission to perform it.
Use Microsoft Entra identities and least-privilege downstream access so business authorization remains enforceable outside the model.
Tool results should be treated as potentially untrusted data, especially when they contain user-generated documents, emails, web content, or external API responses.
High-impact tools should be idempotent where possible and return structured confirmation so retries do not create duplicate transactions or ambiguous completion state.
Multiple agents can help when domains have distinct tools, identities, context, or ownership, but they also add latency, state, failure modes, and debugging complexity.
Define what each agent owns, which messages can cross boundaries, how tasks are delegated, and what happens when agents disagree or fail.
The internal A2A protocol material is useful context for interoperable agent-to-agent communication.
Use a multi-agent pattern only when specialization or separation creates measurable value over one capable agent with well-designed tools.
A supervisor/worker pattern can clarify delegation, while peer-to-peer patterns can reduce central bottlenecks and complicate global state.
Define task ownership, timeout, retry, escalation, and completion rules before adding more agents. Otherwise, one business process can become a conversation among services with no clear accountable owner.
State ownership is another design issue. If several agents can update the same business record or workflow status, the system needs idempotency and conflict handling outside natural-language conversation. A coordinator can serialize high-impact steps, or downstream services can enforce optimistic concurrency and transaction rules. The model should never be the only mechanism preventing two agents from performing the same irreversible action.
Foundry observability includes tracing, metrics, evaluations, and Application Insights integration so teams can follow requests through model calls and tool use.
A final answer can look correct while the internal tool path was inefficient, expensive, or unsafe.
Production traces should identify agent version, model, user/session context, tool calls, errors, latency, and important action outcomes.
The goal is an evidence trail that supports debugging, quality improvement, incident response, and audit.
Trace sampling should preserve enough high-risk and failure cases for diagnosis even when full production tracing would be expensive.
Separate user-visible quality metrics from internal operational metrics: a task can complete successfully while using too many tool calls, or it can fail gracefully while preserving security and evidence.
Agent evaluation should measure task completion, tool selection, argument accuracy, safety, latency, and domain-specific quality.
Keep representative regression datasets and rerun them when instructions, model, tools, data sources, or orchestration logic change.
A new model version can improve general reasoning and alter tool behavior in ways that matter to the application.
Treat agent releases like software releases: compare evidence, approve the change, deploy gradually when appropriate, and keep rollback available.
Use both deterministic tests and model-based evaluators. Deterministic tests are strong for required fields, permissions, tool arguments, and business rules; model-based evaluation can assess quality dimensions such as relevance or task adherence.
Keep evaluator versions fixed during regression comparisons so a score change represents the agent, not a moving judge.
Use separate evaluation slices for ordinary users, privileged users, unusual tool failures, and adversarial instructions. A high average score can hide a serious regression in one small but important population. Release gates should therefore include minimum performance on critical slices instead of relying on one aggregate metric that makes rare high-impact failures disappear.
The AB-620 exam represents a business-oriented agent-builder branch.
The AB-100 exam is the expert agentic business-solutions architecture path.
The AI-300 exam covers MLOps and GenAIOps lifecycle operations.
AI-103 is the custom AI apps and agents engineering role, AB-620 is closer to agent building in business solutions, AI-300 focuses on lifecycle operations, and AB-100 spans the broader architecture.
Use the paths as responsibility boundaries rather than as one mandatory ladder.
Professionals can enter the agent ecosystem from several directions: software developers, low-code builders, AI operations engineers, or business-solution architects.
The best certification path follows the system you are accountable for after deployment rather than the number of agent technologies you touch during one project.
The Azure AI Apps and Agents Developer Associate certification provides the credential context for AI-103.
The Microsoft exam inventory can help with internal navigation.
An enterprise agent inventory should record owner, purpose, model, tools, identity, data sources, evaluation thresholds, environment, and retirement plan.
As agents become more autonomous, governance should increase observability and control without forcing every low-risk workflow into the same heavyweight process.
The internal agentic operations material can provide broader context.
Governance should include discovery and retirement as well as approval. An abandoned agent with working credentials and stale tools is still a production risk even if no team remembers using it.
Maintain an agent registry and periodically verify owner, purpose, identity, data access, model, tools, evaluation status, and business criticality.
Publishing also creates a distribution problem. An agent exposed in Teams, Microsoft 365, a web application, or an API can reach very different audiences even though the backend definition is identical. Governance should record where each version is published and which users can discover it. Retirement is incomplete until stale channels, endpoints, credentials, and tool access have all been removed.
Keep production agents tied to an explicit owner and review date so capability, access, and business purpose remain current as the portfolio grows.
Keep ownership explicit.