Evaluating Apps in Microsoft Foundry
Microsoft Foundry can evaluate models, agents, datasets, and captured interactions so teams can compare quality and safety before deployment and monitor behavior after release. The AI-103 exam is the current Microsoft role target most closely aligned with building AI apps and agents that need this evaluation discipline.
Evaluation is not one score. A production app can be relevant but unsafe, helpful but ungrounded, or accurate on average while failing a high-risk user segment. Foundry’s evaluation workflow is most useful when metrics are tied to release decisions and real application outcomes.
Start with what the user or business process needs to achieve.
A retrieval assistant may prioritize groundedness and citation quality, a support agent may prioritize task completion and tone, and a tool-using agent may prioritize tool choice and argument accuracy.
Write acceptance criteria that another reviewer could apply consistently.
Metrics should answer whether the application is fit for its purpose rather than reward whatever Foundry can measure most easily.
Write separate success criteria for quality and safety. An agent may satisfy the user’s request and violate a policy, or refuse every risky-looking request and fail legitimate users. Production evaluation should therefore include a minimum quality bar and a minimum safety bar. When the application uses tools, add correctness criteria for the action itself rather than assuming a good conversational response proves the underlying business task was executed properly.
Evaluation datasets should include ordinary traffic, hard edge cases, ambiguous wording, long context, adversarial instructions, and high-value business scenarios.
Keep important cohorts or risk categories labeled so results can be sliced later.
A benchmark full of easy prompts can hide the exact failures that will matter most in production.
Version the dataset so teams know whether a score changed because the app changed or because the test set changed.
Include production-like distribution and deliberate challenge cases. Keep common queries frequent enough that the aggregate reflects real usage, then create slices for rare high-impact tasks such as privileged actions, regulatory content, or customer escalation. This allows the team to optimize everyday quality without letting the average hide a catastrophic regression in one small category. Dataset curation is an ongoing product responsibility, not a one-time prelaunch exercise.
The Foundry portal can target a model, an agent, a dataset of existing outputs, or captured traces from deployed interactions.
Model evaluation is useful when choosing or configuring the foundation model.
Agent evaluation should include behavior such as task adherence, tool use, and multi-turn success.
Dataset evaluation lets teams score outputs generated elsewhere without rerunning inference.
Choose the target that isolates the engineering question you are trying to answer.
Layered evaluation speeds diagnosis. If model-only evaluation is strong but the agent fails, investigate instructions, tools, memory, and orchestration. If the agent passes on synthetic conversations and fails on production traces, investigate real input distribution, data freshness, or environmental dependencies. The evaluation target should be chosen to narrow the engineering problem. Running every evaluator against the entire stack can produce scores without explaining what to fix.
A tool-using agent can produce a good final answer after choosing the wrong tool several times or using unsafe arguments.
Foundry’s agent-focused evaluation supports rubric-based checks and built-in quality, safety, and tool-use measures.
Measure completion, tool selection, tool-call accuracy, adherence, and safety rather than evaluating only the last message.
The full conversation is the unit of behavior when the application performs several steps before reaching the user outcome.
Tool-call evaluation should include whether the chosen tool was necessary, whether arguments were correct, and whether the agent used the returned evidence properly. An agent can choose the right tool and then ignore the result, or choose an unnecessary tool that adds latency and risk. Measure step efficiency where it matters, but do not reward the fewest steps when the task genuinely needs verification or approval.
Foundry can evaluate interactions already captured in Application Insights instead of replaying every request.
This is useful for monitoring real-world traffic and for agents built with custom frameworks that emit compatible OpenTelemetry spans.
Trace evaluation is currently preview in some scenarios, so production architecture should respect Microsoft’s regional and support notes.
The design value is important: production evidence can feed the same evaluation concepts used before release.
Production traces reveal user behavior that test writers did not anticipate. Sample interactions by failure, latency, user feedback, risk class, or new feature and run targeted evaluation on those traces. Keep privacy and retention controls in the monitoring pipeline because traces can contain prompts, retrieved data, or tool arguments. The objective is to convert real production evidence into a safer regression set, not to store every interaction forever.
Some Foundry human-evaluation features are preview, but the underlying practice is valuable when domain experts must judge quality that automated metrics capture poorly.
Use human reviewers for complex professional judgment, brand tone, regulated content, or novel failure analysis.
Give reviewers a consistent rubric and sample calibration so ratings reflect the application standard rather than individual taste.
Human findings can later inform custom rubrics or automated evaluators.
Human review should also be used to calibrate model-based evaluators. Select examples across the scoring range, have experts score them independently, resolve disagreements, then compare the automated evaluator. If the judge consistently rewards verbosity while experts prefer concise correctness, change the rubric before trusting the metric. The human process provides the semantic anchor for qualities that do not have a simple deterministic answer.
Evaluation should block or flag a release when critical metrics fall below agreed thresholds.
Pin datasets, model deployments, evaluator versions, and rubric versions when comparisons must be reproducible.
A CI workflow can run a small fast suite on each change and a broader evaluation before production promotion.
Treat evaluation results like test results: attach them to the release and make regressions visible before users encounter them.
Different environments can use different gates. A development branch may allow warnings while production promotion requires all critical safety slices and a minimum task-completion threshold to pass. Store the failed examples, not just the score, so developers can understand the regression. A gate without diagnostic output creates pressure to bypass evaluation; a gate with clear failing cases turns evaluation into useful feedback.
The AI-300 exam is the Microsoft MLOps and GenAIOps operations boundary.
The AB-100 exam is the expert agentic business-solutions architecture path.
AI-103 developers build the app or agent behavior; AI-300-oriented roles go deeper into operational lifecycle, evaluation, deployment, and governance; architecture roles coordinate the larger solution.
The same Foundry evaluation capabilities can support all three responsibilities at different depths.
Role separation can also improve accountability. The application developer may own agent logic and test cases, while an AI platform or operations team owns evaluator infrastructure, deployment gates, monitoring, and model migrations. Architecture or risk teams can define mandatory safety criteria. Foundry supports these workflows in one platform, but organizations still need clear ownership of who defines the metric, who fixes failures, and who can approve an exception.
The Microsoft exam inventory can help with internal navigation.
For a final Foundry project, establish a baseline, change one component, rerun evaluation, inspect trace failures, and decide whether the new version should ship.
Keep the previous version and evaluation results available for rollback and comparison.
Evaluation becomes durable engineering when a model migration, prompt change, tool update, or retrieval change can be judged against the same explicit application standard.
Create a migration playbook: freeze a representative dataset, pin the current production model and evaluator, score the candidate model, inspect regressions, retune prompts or tools if needed, then rerun. Promote only when the new version meets the acceptance bar. After release, compare trace-based production evaluation with the test result. This lifecycle keeps model upgrades evidence-based instead of relying on provider claims that the newer model is globally better.
Add a disagreement workflow for results near the threshold. When one evaluator passes a response and another fails it, inspect the example instead of averaging the scores blindly. The disagreement may reveal a vague rubric, a missing business rule, or a tradeoff between helpfulness and safety that needs product ownership. Evaluation is most valuable when it creates a decision process for ambiguous cases, not only a numerical gate.
After deployment, sample real failures and user-feedback cases into a curated regression set. Keep privacy controls and remove examples that no longer represent current product behavior. This closes the loop between production evidence and pre-release testing so the evaluation suite becomes more realistic over time instead of freezing the assumptions the team had before launch.
Keep one small smoke-test set that runs quickly on every change so catastrophic regressions are caught before the larger suite starts.
Keep one frozen regression set for every important release. It should contain routine cases, known edge cases, previously fixed failures, and a few adversarial inputs. Re-running the same set after a prompt, model, tool, or retrieval change makes quality drift visible before a production rollout turns a subtle regression into a user-facing problem.