Know if you're actually ready. Take the Azure AI-103 quiz → get your AI readiness report.
Take the free test →Plan & Manage an Azure AI Solution: AI-103 Domain 2 (25-30%)
The second-heaviest AI-103 domain, and the operational one. Choosing services, sizing capacity, securing access, watching for drift, and governing agent behaviour.
Model selection is a fit decision
The exam consistently rewards picking the smallest model that clears the accuracy bar for the task. A frontier reasoning model on a high-volume classification step costs more and responds more slowly for no benefit; a small model, prompted or fine-tuned for that narrow job, is usually the right answer. Reserve the large model for steps that genuinely need multi-step reasoning, and remember that an agent workflow can use several models rather than one.
Capacity is measured in tokens, not requests
Generative endpoints are metered on tokens, and one request can be a hundred times another, so requests per second is a poor planning unit. Size against tokens per minute across input and output. A 429 response means the request rate has exceeded the deployment's provisioned allowance — the fixes are raising capacity, spreading load, or backing off on the client, not changing region or model version.
Predictable sustained load, strict latency -> provisioned throughput Variable or bursty load, cost-sensitive -> standard shared capacity Throughput-insensitive offline work -> batch
Keyless beats key rotation
Managed identity removes the credential from configuration entirely: tokens are issued to the workload identity at runtime, so there is nothing to commit to a repository, leak through a log, or forget to rotate. It does not change throughput or pricing — the benefit is purely that the secret does not exist. For network isolation, a private endpoint gives the service a private address inside the virtual network, and disabling public network access is what actually closes the public path; an IP allow-list still routes over the internet.
Cost control in a RAG application
Retrieved context usually dominates input tokens, so retrieving fewer but better-targeted passages cuts cost and often improves answers, because less irrelevant context competes for the model's attention. In an agent, the repeated prefix — system prompt plus tool schemas, resent on every step — is the other large cost, and prompt caching is the direct remedy. Bounding the maximum steps per request is what turns a runaway loop from an open-ended bill into a handled failure.
Monitoring what health checks cannot see
Quality drift produces no errors. A stale search index, an auto-updating model version, or user inputs moving away from what you tested all yield worse answers, successfully, with 200 responses throughout. Detecting it needs quality signals rather than health signals: a fixed evaluation set run continuously against production, sampled human review for what automated scoring cannot judge, and in-product feedback for failures nobody anticipated. Pinning a model version is what keeps an evaluation result valid between releases.
Responsible AI as configuration
Content safety filters classify prompts and completions against harm categories and block what crosses the configured thresholds — they address harm, not factual accuracy. Safety evaluations are the pre-release counterpart: they probe the system with adversarial and risky inputs to characterise how it behaves before users do. Auditing an individual automated decision months later requires that decision's own trace — inputs, retrieved sources, steps and output — because aggregate metrics cannot reconstruct one case.
Deploying changes that are not code changes
Prompt, model and retrieval changes alter behaviour without altering any code path, so unit tests pass while quality moves. A deployment pipeline for a generative application therefore needs an automated evaluation run against a fixed set as a gate — the check a conventional web pipeline has no reason to have. For comparing a candidate model against the incumbent on real traffic, shadow evaluation sends a copy of live requests to the candidate while users continue to receive the incumbent's responses.
Exam tip
When a question gives you an application that is healthy by every infrastructure metric but wrong or expensive in some way, the answer is almost never an infrastructure control. Look for the quality signal or the token-level control instead.
Further reading
Think you're ready? Prove it.
Take the free Azure AI-103 readiness test. Get a score, topic breakdown, and your exact weak areas.
Take the free Azure AI-103 test →Free · No sign-up · Instant results