Concept and mechanism
RAG retrieves relevant information and adds it to context before generation. It is useful for frequently changing runbooks without assuming retraining after every update. Split documents into meaningful chunks, retain versions, and check retrieval quality. A citation supports inspection but does not prove the cited text supports the conclusion. Access control must prevent unauthorized retrieval; ACL details vary by Knowledge Base type and connector. Fine-tuning changes parameters using examples and can improve task behavior. It does not automatically replace access to current facts. Distillation uses a teacher model to help train a smaller student; intended savings require renewed quality measurement. Compare data, training, evaluation, operating, and maintenance costs across approaches.
Guided application
In a fictional case, an assistant recommends an old procedure although a new runbook exists. Inspect retrieved documents, indexed versions, and the request before blaming the model. A clear prompt defines task, context, format, and boundaries; few-shot examples show the desired pattern. Test changes against a representative dataset and retain versions for comparison and rollback. Lower temperature can reduce variation but guarantees neither truth nor identical output across every system. Evaluate both model and application: relevant retrieval, supported answers, task completion, latency, and cost per interaction. ROUGE and BLEU measure aspects of textual overlap; BERTScore compares semantic representations. None alone proves safety or usefulness. Human evaluation and LLM-as-a-judge also need criteria, examples, and consistency checks. For tool workflows, include failures, refusals, and partially completed actions.
An answer linking to an old runbook still needs correction.
Common pitfalls
RAG treated as a truth guarantee; fine-tuning treated as continuous updating; generic benchmark treated as acceptance; unversioned prompt changes.
Related topics: AI and ML: problem, data, and metrics · Generative AI, context, and agents · Responsible AI and explainability
Evaluate the chain from retrieved data through task outcome.
Reference: Amazon Bedrock Knowledge Bases · AIF-C01