Home›Guides›Azure AI-103›AI-103 Text Analysis Domain Guide (10-15%)

Know if you're actually ready. Take the Azure AI-103 quiz → get your AI readiness report.

Take the free test →
Azure AI-103

Text Analysis: AI-103 Domain 5 (10-15%)

Turning free text into structured data, handling personal data safely, and the speech capabilities that make an agent something you can talk to.

Examifyr·2026·7 min read

Free text into structured data

Named entity recognition locates the spans that name things and labels each with its type, which is what makes prose queryable. For richer requirements — turning a support ticket into a record with category, urgency and affected product — a generative extraction prompt returning a defined JSON schema handles all the fields in one call and adapts when the taxonomy changes. Keyword lookup tables miss every paraphrase, and a classifier per field triples the training and maintenance burden.

Sentiment is not one number

Document-level sentiment collapses a review that praises one aspect and criticises another into a middle value describing neither. Aspect-level analysis keeps them separate, which is why it is the right answer whenever the question involves mixed opinions. Detecting hostility that contains no prohibited vocabulary is a related case: it is expressed through phrasing and intent, so it needs a model-based classifier judging the whole message rather than a blocklist.

Redaction happens before, not after

Personal data has to be removed before the text reaches the model and before it is written to logs or traces. A pipeline that cleans the prompt but writes the raw transcript to its trace has not achieved redaction — it has moved the exposure. This is the single most repeated mistake in this area, because traces feel like infrastructure rather than storage.

Note: Agents accumulate context, and that context is frequently personal data. Every place it lands is a place it has to be redacted.

Translation with things that must not move

General-purpose translation will localise brand and product names along with everything else. Supplying a glossary or do-not-translate list constrains those specific terms so they survive intact. Pivoting through an intermediate language compounds errors, and splitting into sentences to translate independently loses the context that made the translation correct.

Speech as an agent modality

Speech-to-text is the input leg — transcribing what the user said so the agent can reason over it — and text-to-speech is the output. A voice agent also needs barge-in detection: listening while speaking, so it can stop playback when the user starts talking. Without it, the agent talks over the person it is serving. Where recognition fails on domain vocabulary in a specific acoustic environment, a custom speech model trained on that vocabulary and those conditions is the direct remedy; no amount of downstream reasoning recovers a misheard part number.

Reasoning over audio directly

Transcription flattens speech to words and discards prosody. A model that consumes audio directly can attend to tone, hesitation and emphasis, which matters for sentiment, urgency and intent in ways the transcript cannot express. It is not universally faster and language coverage is comparable — the signal that survives is the reason to choose it.

Exam tip

Watch for questions where the naive answer is a lexicon — a blocklist for hostility, a keyword table for categories, a phrase list for injection. The exam consistently prefers a model that judges meaning over a list that matches strings.

Further reading

Think you're ready? Prove it.

Take the free Azure AI-103 readiness test. Get a score, topic breakdown, and your exact weak areas.

Take the free Azure AI-103 test →

Free · No sign-up · Instant results

← Previous
AI-103 Computer Vision Domain Guide (10-15%)
← All Azure AI-103 guides