Concept and mechanism
Generative models produce content from learned patterns and supplied context. Tokens can represent words, word fragments, or other elements; they are not a fixed unit of human language. Embeddings represent content as vectors useful for similarity. In a transformer, attention relates context elements without creating a truth guarantee. RAG retrieves relevant content and adds it to the request; it does not itself update model weights. The historical syllabus names Azure AI Foundry; current documentation uses Microsoft Foundry. The platform supports model exploration, solution development, and result evaluation, with capability and availability differences across models.
Guided application
In a fictional runbook assistant, retrieve only authorized, approved content. Check whether a citation points to the active version and actually supports the suggested action. Retrieved documents can contain malicious instructions: keep data and authority separate, with limited permissions for actions. Before release, evaluate anonymized real or synthetic questions, unsupported answers, omitted prerequisites, latency, and cost.. Record model version, prompt, index, and evaluation set. Define how to decline an unsupported recommendation and return to the previous version when quality worsens.
A citation to a retired runbook can explain a wrong answer; changing temperature does not update the source.
Common pitfalls
Fluency as truth; RAG as training; citation as freshness; document as authorized instruction.
Related topics: Workloads and operational responsibility · Machine learning and useful evaluation · Vision and documents with validation
Evaluate the complete system: sources, permissions, generation, and human decision.
Reference: Retrieval augmented generation · AI-900 historical skills measured 2025-05-02; exam retired 2026-06-30