Home›Guides›Azure AI-103›AI-103 Generative AI & Agents Domain Guide (30-35%)

Know if you're actually ready. Take the Azure AI-103 quiz → get your AI readiness report.

Take the free test →
Azure AI-103

Generative AI & Agentic Solutions: AI-103 Domain 1 (30-35%)

The heaviest AI-103 domain and the reason the exam was renumbered. Building agents that use tools, keep memory, orchestrate each other, and get evaluated.

Examifyr·2026·9 min read

What separates an agent from a completion

A completion call is one request and one response. An agent runs a loop: it reasons about a goal, calls a tool, reads the result, and decides what to do next — possibly many times before it answers. Almost every exam question in this domain follows from that difference. It is why cost per request is variable, why one request produces a tree of calls rather than a line, and why evaluating the final answer alone is insufficient.

Note: If a question can be answered either from general LLM knowledge or from how an agent loop behaves, the loop-specific answer is the one being tested.

Tools are described, not documented

The model chooses a tool from its schema — the name, the natural-language description and the typed parameters. That description is the interface: a vague one produces wrong tool selection no matter how good the model is. Credentials never appear in the schema; the orchestration layer holds them and executes the call the model proposes, after validating the arguments against the schema and feeding any validation error back so the model can correct itself.

Model proposes  -> { tool: "get_order", args: { id: "A-1" } }
Orchestrator    -> validate args against schema
                -> invalid? return the error to the model, let it retry
                -> valid?   execute with credentials the model never sees
                -> return the result into context

Memory is storage, not context length

Conversation state lives as long as its thread. Anything that must survive into a later session has to be written somewhere durable and retrieved back into context when the next one starts. A larger context window holds more of one conversation; it persists nothing. When a long conversation approaches the window limit, the standard compaction is to summarize earlier turns while keeping recent ones verbatim, rather than truncating the oldest messages blindly.

Retrieval, and why chunking decides its quality

RAG grounds a response in source content rather than parametric memory, which is what makes answers current and checkable. Its quality is bounded by chunk quality: chunks split on token counts alone cut across meaning and lose the heading that gave them context, so the model receives fragments it cannot interpret regardless of how many it gets. Hybrid search — semantic similarity combined with exact keyword matching — is usually the right default, because embeddings are weakest exactly where identifiers and error codes matter.

Note: Exposing retrieval to the agent as a tool, rather than always prefetching, lets it skip searching when it already knows, refine a query after a first attempt, and search several times for a multi-part question.

Multi-agent orchestration

An orchestrator holds the control flow: it routes work to the right specialist and assembles the results, which is what keeps a multi-agent system debuggable rather than emergent. Splitting one agent into several is justified when subtasks genuinely differ — different tools, different instructions, different model requirements — not because the system prompt got long or the tool list grew. Coordination is a real cost and it should buy something.

Evaluating something non-deterministic

Asserting that an agent's output equals an expected string flags correct behaviour as failure, because two correct answers can be worded quite differently. Agent evaluation needs criteria that tolerate phrasing, and it needs to look at the trajectory rather than only the final response: if retrieval fetched the wrong document and the model summarized it faithfully, the answer is wrong while generation is behaving correctly, and only a per-step view shows which half to fix. Groundedness is the evaluator that catches fabrication, because a fabricated claim is fluent, coherent and fast.

Note: Hold the evaluation set fixed between runs. A measurement whose inputs change cannot attribute a score movement to the change you made.

Guardrails belong around the action

An agent that produces a bad answer may also have sent an email or issued a refund. The gate therefore belongs on the proposed tool call, before it executes — an output filter runs after the money has already moved. The same reasoning governs prompt injection: a retrieved document carrying instructions cannot be reliably ignored by asking the model to ignore it, because the model cannot separate instructions from data in its context. What holds is least-privilege tool permissions and approval gates on consequential actions.

Exam tip

For any control question, find where the damage actually occurs. If the harmful step is an action, the answer is a tool permission or an approval gate — never an output filter, which runs after the action has happened.

Further reading

Think you're ready? Prove it.

Take the free Azure AI-103 readiness test. Get a score, topic breakdown, and your exact weak areas.

Take the free Azure AI-103 test →

Free · No sign-up · Instant results

← Previous
AI-103 Exam Guide — Domains, Weights & Format (2026)
Next →
AI-103 Plan & Manage Azure AI Domain Guide (25-30%)
← All Azure AI-103 guides