In an ideal world (where your samples' timestamps are exactly on the second and your rule evaluation happens exactly on the second) rate(counter[1s]) would return exactly your ICH value and rate(counter[5s]) would return the average of that ICH and the previous 4. Except the ICH at second 1 is 0, not 1, because no one knows when your counter was zero: maybe it incremented right there, maybe it got incremented yesterday, and stayed at 1 since then. (This is the reason why you won't see an increase the first time a counter appears with a value of 1 -- because your code just created and incremented it.)
increase(counter[5s]) is exactly rate(counter[5s]) * 5 (and increase(counter[2s]) is exactly rate(counter[2s]) * 2).
Now what happens in the real world is that your samples are not collected exactly every second on the second and rule evaluation doesn't happen exactly on the second either. So if you have a bunch of samples that are (more or less) 1 second apart and you use Prometheus' rate(counter[1s]), you'll get no output. That's because what Prometheus does is it takes all the samples in the 1 second range [now() - 1s, now()] (which would be a single sample in the vast majority of cases), tries to compute a rate and fails.
If you query rate(counter[5s]) OTOH, Prometheus will pick all the samples in the range [now() - 5s, now] (5 samples, covering approximately 4 seconds on average, say [t1, v1], [t2, v2], [t3, v3], [t4, v4], [t5, v5]) and (assuming your counter doesn't reset within the interval) will return (v5 - v1) / (t5 - t1). I.e. it actually computes the rate of increase over ~4s rather than 5s.
increase(counter[5s]) will return (v5 - v1) / (t5 - t1) * 5, so the rate of increase over ~4 seconds, extrapolated to 5 seconds.
Due to the samples not being exactly spaced, both rate and increase will often return floating point values for integer counters (which makes obvious sense for rate, but not so much for increase).
In an ideal world (where your samples' timestamps are exactly on the second and your rule evaluation happens exactly on the second) rate(counter[1s]) would return exactly your ICH value and rate(counter[5s]) would return the average of that ICH and the previous 4. Except the ICH at second 1 is 0, not 1, because no one knows when your counter was zero: maybe it incremented right there, maybe it got incremented yesterday, and stayed at 1 since then. (This is the reason why you won't see an increase the first time a counter appears with a value of 1 -- because your code just created and incremented it.)
increase(counter[5s]) is exactly rate(counter[5s]) * 5 (and increase(counter[2s]) is exactly rate(counter[2s]) * 2).
Now what happens in the real world is that your samples are not collected exactly every second on the second and rule evaluation doesn't happen exactly on the second either. So if you have a bunch of samples that are (more or less) 1 second apart and you use Prometheus' rate(counter[1s]), you'll get no output. That's because what Prometheus does is it takes all the samples in the 1 second range [now() - 1s, now()] (which would be a single sample in the vast majority of cases), tries to compute a rate and fails.
If you query rate(counter[5s]) OTOH, Prometheus will pick all the samples in the range [now() - 5s, now] (5 samples, covering approximately 4 seconds on average, say [t1, v1], [t2, v2], [t3, v3], [t4, v4], [t5, v5]) and (assuming your counter doesn't reset within the interval) will return (v5 - v1) / (t5 - t1). I.e. it actually computes the rate of increase over ~4s rather than 5s.
increase(counter[5s]) will return (v5 - v1) / (t5 - t1) * 5, so the rate of increase over ~4 seconds, extrapolated to 5 seconds.
Due to the samples not being exactly spaced, both rate and increase will often return floating point values for integer counters (which makes obvious sense for rate, but not so much for increase).
Prometheus calculates rate(counter[d]) at timestamp t in the following way:
- It selects raw samples for the
countertime series on the time range(t-d ... t]. Note that thet-dtimestamp isn't included in the time range, whilettimestamp is included in the time range. If the selected time range contains less than two raw samples, then Prometheus returns an empty value (a gap) at the timestampt. - Then it calculates the increase of the selected raw samples. Usually it is calculated as the difference between the last selected sample and the first selected sample. Calculations become slightly complicated if the
counterwas reset to zero during the selected time range. Let's skip this for the sake of clarity. - Then the resulting increase can be extrapolated if timestamps for the first and/or the last raw samples are located too far from the bounds of the selected time range.
- Then the rate is calculated by dividing the extrapolated increase by
d.
Prometheus calculates increase(counter[d]) in the same way except the last step.
Let's look at a few examples applied to the original data:
second counter_value increase calculated by hand(call it ICH from now)
1 1 1
2 3 2
3 6 3
4 7 1
5 10 3
6 14 4
7 17 3
8 21 4
9 25 4
10 30 5
The
rate(counter[1s])will return nothing at any timestampt, since any time range(t-1s ... t]contains only a single raw sample, while Prometheus requires at least two samples for calculating bothrate()andincrease().The
rate(counter[2s])andincrease(counter[2])would return the following values per each timestamptwhen extrapolation isn't applied:
t counter_value rate(counter[2s]) increase(counter[2s])
1 1 - -
2 3 (3-1)/2=1.0 3-1=2
3 6 (6-3)/2=1.5 6-3=3
4 7 (7-6)/2=0.5 7-6=1
5 10 (10-7)/2=1.5 10-7=3
6 14 (14-10)/2=2 14-10=4
7 17 (17-14)/2=1.5 17-14=3
8 21 (21-17)/2=2 21-17=4
9 25 (25-21)/2=2 25-21=4
10 30 (30-25)/2=2.5 30-25=5
In reality Prometheus results for rate(counter[2s]) and increase(counter[2s]) may be slightly bigger because of extrapolation, since the first sample on the selected time range is located comparatively far from the start of the time range.
Such calculations have the following issues:
Prometheus can return fractional results from
increase()over time series, which contains only integer values. This is because of extrapolation. For example, Prometheus may return fractional results fromincrease(http_requests_total[5m]).Prometheus returns empty results (aka gaps) from
increase(counter[d])andrate(counter[d])when the lookbehind windowddoesn't cover at least two samples - seerate(counter[1s])andincrease(counter[1s])example above.Prometheus completely misses the increase between the raw sample just before the
(t-d ... t]interval and the first raw sample on this interval. This may result in inaccurate calculations. For example,increase(counter[1h])doesn't equal tosum_over_time(increase(counter[1m])[1h:1m]).
Prometheus developers are aware of these issues - see this link. These issues are addressed in the system I work on - VictoriaMetrics - more specifically, in MetricsQL query language - see this comment and this article for technical details.
This is known as aliasing and is a fundamental problem in signal processing. You can improve this a bit by increasing your sample rate, a 4m range is a bit short with a 2m range. Try a 10m range.
Here for example the query executed at 1515722220 only sees the [email protected] and [email protected] samples. That's an increase of 1 over 2 minutes, which extrapolated over 4 minutes is an increase of 2 - which is as expected.
Any metrics-based monitoring system will have similar artifacts, if you want 100% accuracy you need logs.
increase() will always (approximately) double the actual increase with your setup.
The reason is that (as currently implemented):
increase()is (as you observed) syntactic sugar forrate()i.e. it is the value that would be returned byrate()multiplied by the number of seconds in the range you specified. In your case, it israte() * 240.rate()uses extrapolation in its computation. In the vast majority of cases a 4 minute range will return exactly 2 data points, almost exactly 2 minutes apart. The rate is then computed as the difference between last and first (i.e. the 2 points in your case) divided by the time difference of the 2 points (around 120 seconds in 99.99% of cases) multiplied by the range you requested (exactly 240 seconds). So if the increase between the 2 points is zero, the rate is zero. If the increase between the 2 points is1.0, the computedrate()will be close to2.0 / 240and, as a result, theincrease()will be2.0.
This approach works mostly fine with counters that increase smoothly (e.g. if you have a more or less fixed number of signups every 2 minutes). But with a counter that rarely increases (as does your signups counter) or a spiky counter (like CPU usage) you get weird overestimates (like the increase of 2 you are seeing).
You can essentially reverse engineer Prometheus' implementation and get (something very close to) the actual increase by multiplying with (requested_range - scrape interval) and dividing by requested_range, essentially walking back the extrapolation that Prometheus does.
In your case, this would mean
increase(signups_count[4m]) * (240 - 120) / 240
or, more succinctly,
increase(signups_count[4m]) / 2
It requires you to be aware both of the length of the range and the scrape interval, but it will give you what you want: "ones for ones, and twos for twos, most of the time". Sometimes you'll get 1.01 instead of 1.0 because the scrapes were 119 seconds, not 120 seconds apart and sometimes, if your evaluation is closely aligned with the scrape some points right on the boundary might be included or not in a data point calculation, but it's still a better answer than 2.0.
Prometheus calculates increase(m[d]) at timestamp t in the following way:
- It fetches raw samples stored in the database for time series matching
mon a time range(t-d .. t]. Note that samples at timestampt-daren't included in the time range, while samples attare included. It is expected that every selected time series is a counter, sinceincrease()works only with counters. - It calculates the difference between the last and the first raw sample value on the selected time range individually per each time series matching
m. Note that Prometheus doesn't take into account the difference between the last raw sample just before the(t-d ... t]time range and the first raw samples at this time range. This may lead to lower than expected results in some cases. - It extrapolates results obtained at step 2 if the first and/or the last raw samples are located too far from time range boundaries
(t-d .. t]. This may lead to unexpected results. For example, fractional results for integer counters. See this issue for details.
Prometheus calculates rate(m[d]) as increase(m[d]) / d, so rate() results may be also unexpected sometimes. Prometheus developers are aware of these issues and are going to fix them eventually - see these design docs.
In the meantime you can use VictoriaMetrics - this is Prometheus-like monitoring solution I work on. It provides increase() and rate() functions, which are free from issues mentioned above.
Increase measure the increase in your window, and rate is the 'per second' rate.
So if you were to imagine that Prometheus write a database row each second of what the current value is. That helps reason through the metrics calculations.
If the total so far was 0 then the metric increased by 5 in the last 10 seconds, the
increaseis looking back 10 seconds and finding the value and reporting the increase.The same as
1.but it is looking back 30 seconds.The 0.5 comes from looking at the window size (10s), measuring the increase them normalizing it to a 'per second' rate. So the increase of 5/10 seconds: .5 calls per second
With the larger window, the 'per second' is spread over the larger window.
increase is easier to reason about, but rate standardizes on the 'per second' unit.
One problem with the 'per second' unit though is that it can be too small to talk through, but it makes it easy to multiple by 60 (to show a per minute) or 3600 (to show a per hour). The window is only used to calculate the window for how far to gather data for the calculation.