Concept and mechanism
Inference strategy depends on per-request deadlines, data size, demand pattern, and idle-capacity cost. A real-time endpoint supports low-latency interaction; serverless may suit intermittent demand that tolerates cold starts; asynchronous inference queues requests and handles lengthy operations. Batch serves datasets without interactive-response requirements. Check supported characteristics and limits before choosing, including GPU, networking, payload, and observability. Size using representative tests rather than average CPU alone. Queues, tail latency, memory, and concurrency may reveal the limiting resource. Autoscaling needs metrics, bounds, and reaction time; it does not replace suitable initial capacity for anticipated spikes.
Guided application
In a scoring exercise, the team wants to observe a new model without returning its predictions to users. A shadow variant receives request copies while the production variant continues answering. This differs from A/B, where production variants serve different requests. Shadow tests have exclusions, including serverless and asynchronous endpoints, that need checking. Deployment guardrails are a separate capability: they support blue/green or rolling changes with configured observation and rollback for real-time and asynchronous endpoints. A latency alarm does not necessarily detect declining predictive quality. Define operational and functional criteria, an observation window, and a recovery path. Endpoint rollback does not automatically undo decisions already consumed by downstream systems; planning must account for that effect.
An observed shadow response is not the response returned to the user.
Common pitfalls
Shadow confused with A/B; deployment guardrails confused with content filters; rollback ignoring downstream effects; assumed instant scaling.
Related topics: Ingestion, storage, and quality · Features, splits, and leakage · Training, tuning, and reproducible experiments
Combine inference mode, representative observation, and recovery of the complete outcome.
Reference: SageMaker inference options · MLA-C02