Concept and mechanism
Deployment should respect when the result is needed. A daily forecast before a batch can be computed in a batch; a decision requested during an interaction may require a synchronous response; a continuous event sequence can use incremental processing. Do not choose a permanent endpoint merely because a model exists. Define latency, freshness, volume, cost, and behavior when no valid prediction is available. Custom models in Model Serving use compatible artifacts and can be exposed through an API. A registered version, a served version, and a prediction actually consumed are different states. Include instructions for identifying the version and interpreting schema or access errors in the handover.
Guided application
An endpoint can route traffic percentages to different versions. That distribution does not guarantee an exact count per version in a small sample; directly invoking one specific model bypasses general percentages. Also check the identity recorded as creator: removing it from the workspace can prevent updates even when the API caller has permissions. Plan suitable service identities and controlled recovery. For streaming, relate the guide’s Delta Live Tables terminology to Lakeflow Spark Declarative Pipelines in current documentation. Processing state and checkpoints do not prove that an external side effect, such as opening a ticket, occurred only once.
A healthy API may respond using a version older than the intended promotion.
Common pitfalls
Alias treated as rollout; checkpoint treated as external-side-effect guarantee; caller permissions treated as the only control.
Related topics: Environment, AutoML, and reproducibility · Temporal features and consistency · MLflow, registry, and promotion
Verify version, identity, traffic, and the result observed by the consumer.
Reference: Create custom model serving endpoints and identity · 2025-03-01