← AWS Machine Learning Engineer: prepare and operate models
07 / 8 · 40 MIN

Monitoring, drift, and useful cost

Distinguish infrastructure health, data change, and outcome quality.

Concept and mechanism

An endpoint can respond quickly while producing unsuitable predictions. Observe three layers: infrastructure, data quality, and model or application quality. Latency, errors, utilization, and queues help locate service issues. Input-distribution changes indicate data drift but do not alone establish lost accuracy. To measure predictive quality, correlate predictions with actual outcomes that may arrive days later. Retain identifiers, timestamps, evaluated windows, and label coverage. Without that coverage, a metric may reflect only easily confirmed cases. For RAG and agents, observe retrieval, factual support, tool failures, incomplete steps, and task completion. Text appearing in a response does not prove an action completed.

Guided application

In a fictional scenario, per-call cost fell while retries doubled and completion declined. Compare cost per useful outcome, including inference, embeddings, vector indexes, storage, and idle capacity. If240 units fund800 completed tasks, observed cost is0.30 per task even if more calls occurred. Keep service availability explicit: Model Monitor is closed to new customers according to inspected documentation. For new projects, assess supported monitoring pipelines, CloudWatch metrics, and adapted reference solutions; do not copy demonstration thresholds as bank policy. Existing customers can continue using the service. An alert needs ownership, context, and a diagnostic action; automatic retraining without investigating defective data can perpetuate failure.

IN PRACTICE

Input drift, infrastructure failure, and quality loss require different evidence.

Common pitfalls

Latency treated as accuracy; ignored delayed labels; per-call cost treated as useful cost; retraining without diagnosis; reference settings treated as policy.

Related topics: Ingestion, storage, and quality · Features, splits, and leakage · Training, tuning, and reproducible experiments

Take this idea with you

Connect metrics to outcomes and investigate causes before automating corrections.

Create account

Reference: Model Monitor availability change · MLA-C02