← AWS Machine Learning Engineer: prepare and operate models
04 / 8 · 40 MIN

Foundation models and RAG quality

Choose customization and measure the retrieval-generation chain.

Concept and mechanism

Foundation-model selection combines capabilities, task quality, latency, cost, and constraints. Prompt engineering changes instructions and context; fine-tuning changes parameters using examples; RAG supplies retrieved information; distillation transfers behavior from a teacher to a student. These mechanisms can be combined, but each brings its own data, maintenance, and evaluation. For changing operational information, check ingestion and retrieval before assuming retraining fixes failure. Version prompts, models, retrieval configuration, and embeddings. An index built with a representation incompatible with queries can degrade search. Reranking rearranges retrieved candidates; it does not necessarily recover a document that never entered the candidate set. Measure that stage separately when locating failures.

Guided application

In a fictional case, search finds the correct procedure for96 of120 requests with a known relevant document. Observed coverage is80%; the generator still needs evaluation for correct use of that material. A summary may overlap well with a reference while omitting a critical restriction. Combine factual criteria, completeness, behavior under insufficient evidence, and human evaluation. An LLM evaluator needs a rubric and consistency checks rather than being treated as infallible. Include documents without answers and malicious-instruction attempts. When evidence is missing, the system should follow a defined clarification or referral path. Product evaluation includes task completion, cost per useful result, and failure handling alongside individual model metrics.

IN PRACTICE

80% relevant retrieval does not equal80% correctly completed tasks.

Common pitfalls

Reranking confused with ingestion; benchmarks treated as acceptance; unversioned indexes; uncalibrated automated judges.

Related topics: Ingestion, storage, and quality · Features, splits, and leakage · Training, tuning, and reproducible experiments

Take this idea with you

Measure retrieval, generation, and operational outcome as distinct parts.

Create account

Reference: Amazon Bedrock evaluations · MLA-C02