← Monitoring and Observability: measure and investigate
07 / 8 · 60 MIN

PromQL: contracts and missing data

Build queries that preserve the measured population and interpret results with executable fixtures.

Preserve resets before aggregation

The lab represents two instances of the funds service. The first has values 0, 60, 120, 0, 60, and 120 at consecutive minutes; the second increases by 120 each minute. Their preaggregated sum never decreases despite the first reset. At five minutes, promtool calculates 2.75/s when summing per-series rates and 2.25/s over the preaggregated series. The difference shows information hidden by aggregation. Do not interpret it as an exact count of lost requests: rate estimates a rate from samples and handles window boundaries. In an APS dashboard, retain the contract that distinguishes resets from demand changes.

Choose a window with enough samples

A query can be syntactically correct yet lack enough data to produce a rate. In the fixture, a sample arrives at every exact minute. At five minutes, [1m] excludes the sample exactly on the left boundary and includes the right-hand sample. Only one sample remains, so rate does not produce the intended result. [2m] includes two and returns 1/s for the counter increasing by 60 per minute. This observation does not prescribe an ideal window for every service. In operations, consider scrape interval, failures, and detection delay. Test startup, resets, and gaps so a missing estimate does not become a claim of no activity.

Treat labels as a matching contract

The second fixture numerator has service=funds and region=eu; its denominator has only service=funds. Division without a modifier returns no series. With one series on each side and on(service), the result is 0.01 carrying service but not region. Before applying this solution in production, confirm population meaning. A global total can be the wrong denominator for a regional SLI even when the query starts returning numbers. If several series exist per service, examine cardinality and the aggregation or matching needed. group_left does not automatically distribute traffic across regions. The contract should also specify which labels the output needs to retain.

Distinguish observed zero, empty output, and fallback

In the fourth fixture, one present series remains at zero while another is never supplied. absent_over_time does not flag the first; it produces an absence signal for the second. Adding or vector(0) makes the missing series appear as zero on the panel but creates no service observation. It also does not turn 0/0 into a success measurement: a no-event case needs an explicit policy. Initializing expected series at zero can improve the instrumentation contract. Retain a separate way to observe collection health. During an incident, report unknown when evidence is missing; do not use a display convention to declare recovery.

Measure the approved latency boundary

The fifth fixture publishes a classic histogram with 300 ms, 500 ms, and infinity buckets. Rates are 90, 98, and 100 observations per second. The fraction within 300 ms is 90%; using the 500 ms bucket would produce 98% for a different question. Do not change the legend to repair the mismatch. If the approved boundary does not match available resolution, improve instrumentation or identify the result as an estimate. Population and window must match across numerator and denominator. The lab deliberately uses classic histograms; it does not establish query equivalence or precision for every native representation.

Separate sample freshness from result freshness

A fresh sample can carry old information. The batch fixture retains zero as the last-success time and supplies new samples through t=600 seconds. time minus the value gives an age of 600 seconds; time minus the metric timestamp gives zero. These answer different questions. In the real application, document what updates last_success: execution completion, output validation, or another agreed condition. The project manager should include that contract in handover. Healthy collection does not establish that the daily file is current. Correlate success age with business date, expected volume, and evidence of authorized processing. Keep collection recovery and application recovery separate in the incident report.

# Independent PromQL expressions for the matching synthetic fixtures
sum by (service) (rate(dr_requests_total[5m]))
dr_error_rate / on (service) dr_total_rate
absent_over_time(dr_missing{service="funds"}[3m])
time - dr_last_success_timestamp_seconds
IN PRACTICE

With fresh samples at t=600 and last_success=0, last-success age is 600 seconds. Sample age is zero. The panel must answer the agreed question.

Common pitfalls

Aggregating before reset handling; one-sample window; valid join with the wrong population; fallback as observation; 500 ms bucket as a 300 ms boundary.

Related topics: SLIs and SLOs · Telemetry collection · Batch operations

Take this idea with you

Check population, labels, type, window, and absence before turning a number into operational status.

Create account

Reference: Unit testing for rules · Observability 2026-09; selected OpenTelemetry, Prometheus and Dynatrace Classic concepts