Concept and mechanism
A generative model predicts outputs from context without guaranteeing every statement matches a current procedure. RAG retrieves material to support the answer; fine-tuning adapts behavioral patterns from examples. For weekly changing documents, retrieval supports updating and evidence. Quality depends on the source, chunking, search, authorization, and instructions rather than just the model. Context budgeting must reserve output space and respect the concrete limits of the selected version. In classic Azure OpenAI contracts expecting a deployment name, the catalog identifier is not an automatic substitute. SDK, endpoint, and API version must belong to the same contract.
Guided application
In a fictional runbook assistant, compare a baseline against each prompt, retrieval, or model change. Use representative questions, unanswerable cases, and requests from users lacking rights to particular documents. Evaluate relevance, groundedness, completeness, safety, latency, and cost separately. A fluent answer can add an instruction absent from the source. Authorization must act before a document enters context; hiding its citation does not prevent disclosure. Version flows, prompts, environment connections, and evaluation results. In traces, retain search-to-generation correlation, minimize sensitive data, and restrict access. Do not log authentication tokens to facilitate reproduction: investigation should use controlled mechanisms and appropriate data.
A rollback answer cites an old version. Trace the document, index version, and submitted context before attributing the error to the model.
Common pitfalls
RAG as automatic authorization; fine-tuning as fact updates; fluency as grounding; traces as an unrestricted archive.
Related topics: Services, identity, and operations · Agents, tools, and control · Vision, metrics, and lifecycle
Evaluate the answer with the sources, identity, and configuration that produced it.
Reference: Generative solution evaluation and observability · AI-102 historical skills measured 2025-12-23; exam retired 2026-06-30