← AI-102: Azure AI engineering, historical course
02 / 6 · 40 MIN

Generation, grounding, and evaluation

Build traceable answers and evaluate complete application behavior.

Concept and mechanism

A generative model predicts outputs from context without guaranteeing every statement matches a current procedure. RAG retrieves material to support the answer; fine-tuning adapts behavioral patterns from examples. For weekly changing documents, retrieval supports updating and evidence. Quality depends on the source, chunking, search, authorization, and instructions rather than just the model. Context budgeting must reserve output space and respect the concrete limits of the selected version. In classic Azure OpenAI contracts expecting a deployment name, the catalog identifier is not an automatic substitute. SDK, endpoint, and API version must belong to the same contract.

Guided application

In a fictional runbook assistant, compare a baseline against each prompt, retrieval, or model change. Use representative questions, unanswerable cases, and requests from users lacking rights to particular documents. Evaluate relevance, groundedness, completeness, safety, latency, and cost separately. A fluent answer can add an instruction absent from the source. Authorization must act before a document enters context; hiding its citation does not prevent disclosure. Version flows, prompts, environment connections, and evaluation results. In traces, retain search-to-generation correlation, minimize sensitive data, and restrict access. Do not log authentication tokens to facilitate reproduction: investigation should use controlled mechanisms and appropriate data.

IN PRACTICE

A rollback answer cites an old version. Trace the document, index version, and submitted context before attributing the error to the model.

Common pitfalls

RAG as automatic authorization; fine-tuning as fact updates; fluency as grounding; traces as an unrestricted archive.

Related topics: Services, identity, and operations · Agents, tools, and control · Vision, metrics, and lifecycle

Take this idea with you

Evaluate the answer with the sources, identity, and configuration that produced it.

Create account

Reference: Generative solution evaluation and observability · AI-102 historical skills measured 2025-12-23; exam retired 2026-06-30