Concept and mechanism
Foundation-model selection combines capabilities, task quality, latency, cost, and constraints. Prompt engineering changes instructions and context; fine-tuning changes parameters using examples; RAG supplies retrieved information; distillation transfers behavior from a teacher to a student. These mechanisms can be combined, but each brings its own data, maintenance, and evaluation. For changing operational information, check ingestion and retrieval before assuming retraining fixes failure. Version prompts, models, retrieval configuration, and embeddings. An index built with a representation incompatible with queries can degrade search. Reranking rearranges retrieved candidates; it does not necessarily recover a document that never entered the candidate set. Measure that stage separately when locating failures.
Guided application
In a fictional case, search finds the correct procedure for96 of120 requests with a known relevant document. Observed coverage is80%; the generator still needs evaluation for correct use of that material. A summary may overlap well with a reference while omitting a critical restriction. Combine factual criteria, completeness, behavior under insufficient evidence, and human evaluation. An LLM evaluator needs a rubric and consistency checks rather than being treated as infallible. Include documents without answers and malicious-instruction attempts. When evidence is missing, the system should follow a defined clarification or referral path. Product evaluation includes task completion, cost per useful result, and failure handling alongside individual model metrics.
80% relevant retrieval does not equal80% correctly completed tasks.
Common pitfalls
Reranking confused with ingestion; benchmarks treated as acceptance; unversioned indexes; uncalibrated automated judges.
Related topics: Ingestion, storage, and quality · Features, splits, and leakage · Training, tuning, and reproducible experiments
Measure retrieval, generation, and operational outcome as distinct parts.
Reference: Amazon Bedrock evaluations · MLA-C02