AB-100 Agent Interoperability: MCP, A2A and Trust Boundaries

AB-100 Agent Interoperability: MCP, A2A and Trust Boundaries

A contact-center agent knows how to summarize a case, but an external fulfillment service knows the actual shipping status. Connecting them sounds straightforward until the architect must decide whether the fulfillment capability is a tool, a specialist agent, or an independent business system. Those distinctions determine identity, failure handling, auditability and who may perform an action.

The AB-100 exam includes Model Context Protocol (MCP), Agent2Agent (A2A) interoperability and multi-agent design. The real architecture question is not how many protocols a solution can use; it is which authority crosses each boundary and how that authority remains controlled as agents and tools evolve.

Decide whether the dependency is a tool or an agent

A shipping-status API performs a bounded operation on a known order. That is often best exposed through a typed connector or tool. A specialist service that independently plans several subtasks, uses its own knowledge and returns a reasoned recommendation may warrant an agent-to-agent interface. Mixing the two can make ownership harder to understand and can turn a deterministic lookup into an unpredictable conversation.

Classify a dependency by the decisions it owns. If the order system owns shipping state, an AI agent should not reinterpret a stale cache as a competing source of truth. If a separate compliance agent owns policy interpretation, establish which version of the rule it uses and whether its output is advice or an enforceable business decision.

A simple workflow may be superior to multiple autonomous components. Complexity should buy separation of responsibilities, reuse or reliability—not just a fashionable diagram.

MCP is a tool and context integration boundary

Model Context Protocol can standardize how clients discover and invoke supported tools or retrieve resources from a server. That is useful when a business wants a consistent way to expose a controlled capability to different AI clients. The server’s advertised description and schema help the host choose an operation, but they do not replace the service’s own authorization and input validation.

For the fulfillment lookup, expose only the operations the agent needs: verify an order reference and read its permitted shipping status. A schema should describe required parameters and the response fields precisely. A loosely documented general API with broad mutation capability would give the agent unnecessary choices and make review difficult.

Treat discovered tool descriptions as configuration that deserves scrutiny. A changed endpoint or newly exposed action can enlarge capabilities after an agent has already been approved. Use a controlled inventory and review process for what a production agent is allowed to discover and execute.

A2A connects independently owned agents with contracts

Consider a support coordinator requesting that a specialist agent review a disputed delivery charge. The coordinator should provide a permitted task identifier, a bounded summary of verified facts and an expected result format. The specialist returns its decision or a reason it cannot proceed; it does not automatically gain permission to change the finance record. The receiving agent’s team may deploy new versions independently, so the contract should define how incompatible fields, missing capabilities and changed data classifications are handled. Before upgrading, test a known valid task, an unauthorized request and a timeout that occurs after partial work. A cross-platform protocol is useful precisely because it exposes a boundary that needs ownership; it does not make independently operated systems equally trustworthy.

The Agent2Agent protocol provides a pattern for agents or orchestrators to exchange work across system and organizational boundaries. A contact-center coordinator may delegate a specialized task to an order-dispute agent operated by another team. In that case, the receiving agent has its own lifecycle, knowledge and behavior; the caller should not assume it is simply a function that follows every instruction verbatim.

Define an explicit task contract: purpose, allowed data, identity, result status, timeouts, error states and the meaning of completion. If a downstream agent reports that a dispute has been reviewed, that is not the same as confirming that a payment reversal has been committed in the transactional system.

Platform support for MCP and A2A channels can depend on product version, environment and preview availability. Design the integration around verified current capabilities rather than promising every Copilot Studio environment can expose every protocol combination today.

Make identity delegation a conscious design decision

A user may ask an agent to find their order, but the coordinator’s own service identity may have broader access than the customer. Passing a natural-language request to another agent does not automatically convey the customer’s business rights. The receiver must understand which principal is making the request and which resources that principal may access.

Use platform-supported delegated credentials where appropriate, or a tightly scoped service identity with server-side customer and record checks. Avoid embedding long-lived tokens in the conversation transcript or giving an entire integration broad administrator permissions just to simplify development.

Test a legitimate lookup and the same lookup for a different customer’s order. The denial should occur at an enforceable boundary, not because the downstream language model happens to be polite.

Limit context transfer and prevent information leakage

A connected agent rarely needs an entire call transcript to answer a shipping-status question. Send the validated order identifier, the minimum permitted context and a task-specific correlation token when needed. Sensitive payment details and internal employee commentary should stay out of the payload unless there is a clearly authorized purpose.

Data minimization makes troubleshooting simpler as well as safer. A smaller payload reduces accidental disclosure and helps reviewers establish why the receiving agent used particular information. It also limits the risk that irrelevant conversational content changes how the specialist interprets its task.

A receiving component should treat externally provided documents and agent messages as untrusted evidence. It must not grant new tool access because an upstream agent describes an instruction as urgent, official or higher priority.

Put business mutations behind deterministic controls

Reading order status is a different class of action from changing delivery details, issuing credits or closing a case. A tool capable of mutations should validate typed parameters, current state, caller authorization and any policy-mandated approval independently of model output. The agent can request an action; the business service decides whether it is allowed.

Design for timeouts after a commit. If a refund action succeeds but its acknowledgment is lost, an unbounded retry can duplicate the financial effect. Use a transaction identifier and an appropriate idempotency mechanism or equivalent business control so that the same logical request is not executed twice.

A response saying “approved” should correspond to an auditable committed state. Where an action requires a human, the agent should clearly indicate pending approval rather than presenting the intended future result as complete.

Trace each handoff without logging unnecessary secrets

An enterprise incident may span a coordinator, an MCP tool host, an A2A specialist and a Dynamics 365 transaction. Carry a correlation identifier, request status, component version and elapsed time so operators can reconstruct which step failed. Record the action result independently of the generated customer-facing explanation.

Limit logging to fields required for operations, security and audit. A detailed conversation log containing personal information or credentials can create a new unauthorized data store. Different teams may need different views of the same trace: latency statistics for platform engineers and specific permitted case actions for investigators.

Build alerts for repeated failures, timeout storms, unexpected tool invocations and mismatches between claimed and committed outcomes. A successful model response does not prove a downstream service actually performed the requested work.

Test loops, stale tasks and specialist failures

Also test a version mismatch between two independently deployed agents. The coordinator may expect a structured completion status while the specialist now returns a different field or an additional approval state. The safe behavior is to detect that incompatibility, retain the unresolved task and notify the responsible service owner. A natural-language summary claiming everything worked is not a substitute for contract validation. Rehearse how the team would restore the previous supported interface or deploy compatible changes in a controlled sequence while preserving case state.

Suppose a coordinator asks a specialist to check a delayed shipment. The specialist requests a customer verification from another agent, which routes the request back to the original coordinator. Without explicit stop conditions, the workflow can loop and consume resources without resolving the case. Set bounded delegations, retry limits and a clear fallback to a human.

Test a specialist returning malformed fields, an unavailable tool server, an expired delegated identity and a downstream operation that completed after the coordinator timed out. For each case, establish what the customer sees and which team owns the repair. Avoid silently substituting an invented result for missing service data.

A recovery test must confirm the correct business state after the dependency is restored, not merely that the conversation proceeds. This is what distinguishes a resilient agent architecture from a happy-path demonstration.

Review interoperability as a governed contract

An architecture decision record should state why a tool, connector, native connected agent or external A2A boundary was chosen. Include task ownership, schema/version lifecycle, authentication, permitted information, mutation approval, observability, expected failure modes and rollout controls.

The existing Copilot Studio multi-agent guide covers platform patterns. This AB-100 decision adds the enterprise contract: how independently managed components cooperate without accidentally inheriting one another’s authority.

Test one authorized and one denied request, one timeout, one duplicate-action attempt and one support escalation. Interoperability is successful when teams can understand and recover a cross-system business process without granting broad trust merely because the participants are called agents.

img