Every health endpoint and TLS certificate on a host, monitored from one Zabbix macro.
PUSH IT Webcheck is a Zabbix template for hosts that run web services. Describe the endpoints of a host as a
JSON list in the macro {$PUSHIT.WEBCHECK.CONFIG}, link the template, and low-level discovery creates the items and
triggers for HTTP status codes, for values inside JSON health responses and for TLS certificate validity and expiry.
All requests are made by the Zabbix agent on the monitored host itself: http://localhost:8080/health means port
8080 on that machine, not on the Zabbix server, so a service bound to localhost or to an internal port is as easy
to check as a public one.
- Features
- How it works
- Quick start
- Endpoint configuration
- Configuration examples
- Macros
- Items and triggers
- Tips and gotchas
- Development
- License
- π©Ί Three check types.
http_statuscompares the response status code,json_pathcompares one value inside a JSON body,certificatewatches TLS validity and expiry. - π§Ύ One macro per host. All endpoints live in
{$PUSHIT.WEBCHECK.CONFIG}; no item cloning, no template per service. - π Runs where the service runs. Checks are Zabbix agent items, so
http://localhost:8080/healthworks without exposing anything. - π§Ή Self-cleaning. Remove an entry from the macro and its items and triggers disappear on the next discovery run.
- π Three-stage certificate alerts. WARNING, AVERAGE and HIGH thresholds in days, chosen per endpoint.
- β³ Grace periods. A missing response becomes a problem only after a configurable grace period; a wrong value is reported immediately.
- ποΈ Tunable with macros. Intervals and grace periods are user macros you can override per host.
The template ships two discovery rules of type Script, HTTP and Certificate. Together they turn the
endpoint list in {$PUSHIT.WEBCHECK.CONFIG} into items and triggers:
flowchart LR
CFG(["Host macro<br/>{$PUSHIT.WEBCHECK.CONFIG}<br/>JSON array, one object per endpoint"])
subgraph LLD ["Discovery rules, every 10m"]
DH["HTTP<br/>TYPE http_status or json_path"]
DC["Certificate<br/>TYPE certificate"]
end
subgraph HI ["HTTP items, every HTTP.INTERVAL"]
RAW["Response {#NAME}<br/>web.page.get"]
ST["HTTP status {#NAME}<br/>status line regex"]
JP["JSON path {#NAME}<br/>body + JSONPath"]
end
subgraph CI ["Certificate items, raw data every CERTIFICATE.INTERVAL"]
CRAW["Certificate data {#NAME}<br/>web.certificate.get"]
NA["Certificate expiration {#NAME}<br/>notAfter timestamp"]
VAL["Certificate validity {#NAME}<br/>verdict"]
DAYS["Days until certificate expires ({#NAME})<br/>calculated every 6h"]
end
subgraph TRG ["Triggers"]
T1{{"Unexpected status code<br/>HIGH"}}
T2{{"Unexpected JSON value<br/>HIGH"}}
T3{{"Certificate data unavailable<br/>INFO"}}
T4{{"Certificate is invalid<br/>HIGH"}}
T5{{"Certificate will expire in N or less<br/>N = CERT_WARN / AVG / HIGH_DAYS<br/>WARNING / AVERAGE / HIGH"}}
end
CFG --> DH & DC
DH --> RAW
RAW -->|http_status| ST
RAW -->|json_path| JP
DC --> CRAW
CRAW --> NA & VAL
NA --> DAYS
ST --> T1
JP --> T2
CRAW --> T3
VAL --> T4
DAYS --> T5
Step by step:
- Configuration. Every endpoint is one JSON object in the host macro
{$PUSHIT.WEBCHECK.CONFIG}. Its keys are LLD macros such as{#NAME},{#URL}and{#TYPE}. - Discovery. Both rules run every 10 minutes and receive the macro through the script parameter
config. A filter on{#TYPE}decides which rule handles an entry: HTTP takeshttp_statusandjson_path, Certificate takescertificate. Anything else is ignored. - Raw items. Each entry gets one Zabbix agent item that talks to the endpoint:
web.page.get["{#URL}"]for HTTP entries,web.certificate.get["{#URL}"]for certificate entries. - Derived items. Dependent items extract the interesting part with preprocessing (regular expression and
JSONPath). Two overrides in the HTTP rule make sure that only the dependent item matching the entry's type is
created. The certificate rule adds a calculated item that turns the
notAftertimestamp into days. - Triggers. Trigger prototypes compare the values with your expectations (
{#EXPECT_STATUS},{#EXPECT_VALUE},{#CERT_*_DAYS}) and also raise a problem when no value has arrived for longer than the grace period. - Cleanup. Lost resources are deleted immediately: remove an entry from the macro and its items, triggers and history are gone after the next discovery run.
Changes to the macro take effect on the next discovery run, at most 10 minutes later, or right away with Execute now on the two discovery rules.
From download to the first problem in Zabbix in six steps.
| Requirement | Why |
|---|---|
| Zabbix server 7.4 or later | template.yaml is a Zabbix 7.4 export (zabbix_export.version: '7.4'); import it into Zabbix 7.4 or a later release. |
| A Zabbix agent on the monitored host | http_status and json_path use web.page.get, available in the classic Zabbix agent and in Zabbix agent 2. |
Zabbix agent 2 for certificate checks |
web.certificate.get exists in Zabbix agent 2 only. |
| An agent interface on the host | All collecting items are passive Zabbix agent items, so the server or proxy must be able to poll the agent. |
| A network path from the agent to the endpoints | Requests are made from the agent, not from the Zabbix server or proxy. localhost is the monitored host itself. |
-
Download
template.yamlfrom this repository. -
Import the template. Go to Data collection β Templates, click Import, select
template.yaml, keep the default import rules and click Import. PUSH IT Webcheck now appears in the template group Templates. -
Link it to the host that runs the services: open the host under Data collection β Hosts and, on the Host tab, type
PUSH IT Webcheckinto the Templates field, select it and click Update. -
Configure the endpoints. Open the host again, switch to the Macros tab, select Inherited and host macros, click Change next to
{$PUSHIT.WEBCHECK.CONFIG}and paste your endpoint list as one line of JSON. For a first test, a single HTTP status check is enough (adjust the port and path):[{"{#NAME}":"app","{#TYPE}":"http_status","{#URL}":"http://localhost:8080/health","{#EXPECT_STATUS}":"200"}]Click Update.
-
Run discovery. Wait for the next run (up to 10 minutes) or open Data collection β Hosts, click Discovery in the row of your host, select HTTP and Certificate and click Execute now. Give the server a few seconds after saving the macro so that its configuration cache has picked up the new value.
-
Check the result. Monitoring β Latest data, filtered by your host, shows Response app (the raw HTTP response) and HTTP status app with the value
200. If the endpoint answers with a different status code, or does not answer at all for more than 3 minutes, Monitoring β Problems shows Unexpected status code for app with severity HIGH.
From here, extend the list with the configuration examples. Every later change to the macro follows the same path: edit the value, then wait for the next discovery run or use Execute now.
{$PUSHIT.WEBCHECK.CONFIG} holds a JSON array in Zabbix LLD format. Each element describes one check as an object
whose keys are LLD macros. Three keys are common to all checks; the value of {#TYPE} decides which additional keys
are needed.
Write every value as a JSON string:
"200"rather than200,"true"rather thantrue. Strings are always safe for low-level discovery, and this matches how the values are compared: status codes and day thresholds numerically, JSON values as text.
| Key | Description |
|---|---|
{#NAME} |
Unique name of the check on this host. It is used unquoted inside item keys such as pushit.webcheck.http.status[{#NAME}] and appears in item and trigger names, so use only letters, digits, _, - and .; no spaces, commas, brackets or quotes. |
{#URL} |
The endpoint as scheme://host[:port][/path]. The default port of the scheme and the root path apply when omitted. Must be unique among the http_status and json_path entries of the host, and unique among its certificate entries; see example 5. |
{#TYPE} |
http_status, json_path or certificate. Any other value matches no discovery rule and the entry is silently skipped. |
Fetches the URL every {$PUSHIT.WEBCHECK.HTTP.INTERVAL} (default 1m) with the agent item
web.page.get and
compares the status code from the first line of the response. The body is not evaluated.
| Key | Description |
|---|---|
{#EXPECT_STATUS} |
Expected status code, compared numerically. "200" for a plain health endpoint, but any code works, for example "401" for an endpoint that is supposed to demand authentication. Any other code raises Unexpected status code for {#NAME} (HIGH). |
Fetches the URL every {$PUSHIT.WEBCHECK.HTTP.INTERVAL} (default 1m) with the same web.page.get item, drops
the response headers, treats the body as JSON and compares one value in it. The status code is not evaluated by this
type; add an http_status entry if you need both, see example 5.
| Key | Description |
|---|---|
{#JSON_PATH} |
Zabbix JSONPath expression selecting the value, for example $.status or $.components.db.status. Use a definite path: an expression that returns a list yields a JSON array as text, which is hard to compare. |
{#EXPECT_VALUE} |
Expected value, compared as a string. JSON strings arrive without quotes (UP), booleans and numbers as text (true, 200). A mismatch raises Unexpected JSON value for {#NAME} (HIGH). |
Opens a TLS connection every {$PUSHIT.WEBCHECK.CERTIFICATE.INTERVAL} (default 15m) with the agent 2 item
web.certificate.get,
validates the certificate chain for the host name in the URL and reads the expiry date.
The URL must use the https scheme; a port is optional, a path is ignored. This type needs Zabbix agent 2.
| Key | Description |
|---|---|
{#CERT_WARN_DAYS} |
Days before expiry at which a WARNING problem is raised. |
{#CERT_AVG_DAYS} |
Days before expiry at which an AVERAGE problem is raised. |
{#CERT_HIGH_DAYS} |
Days before expiry at which a HIGH problem is raised. |
Choose the thresholds so that WARN > AVG > HIGH, for example "30", "14" and "7".
Every example is a complete, valid value for {$PUSHIT.WEBCHECK.CONFIG} and follows the rules from
Endpoint configuration: every value is a JSON string, every {#NAME} is unique on the
host, and every {#URL} is unique within its group (see example 5).
The examples are pretty-printed for readability. In the macro value the same JSON is usually pasted as one line, as shown in the compact form at the end of this section. Whitespace is irrelevant to JSON but counts towards the 2048-character macro limit.
A single service on port 8080 whose health endpoint must answer 200.
[
{
"{#NAME}": "app",
"{#TYPE}": "http_status",
"{#URL}": "http://localhost:8080/health",
"{#EXPECT_STATUS}": "200"
}
]What you get for app
| Object | Name | Severity |
|---|---|---|
| Item | Response app | β |
| Item | HTTP status app | β |
| Trigger | Unexpected status code for app | HIGH |
The raw response is fetched every minute. The trigger fires when the code is not 200, or when no code has arrived
for 3 minutes, for example because the endpoint refuses connections.
Any status code works as expectation. An endpoint behind authentication, for example, should answer 401:
[
{
"{#NAME}": "admin-login",
"{#TYPE}": "http_status",
"{#URL}": "http://localhost:9000/admin",
"{#EXPECT_STATUS}": "401"
}
]Spring Boot Actuator answers GET /actuator/health with a body such as {"status":"UP"}. The first entry selects
$.status and expects UP. The second watches an Elasticsearch node whose cluster health must be green, the
third reads a boolean from a body like {"ready":true}, which is compared as the string "true". All three
endpoints must answer without authentication.
[
{
"{#NAME}": "shop-api",
"{#TYPE}": "json_path",
"{#URL}": "http://localhost:8080/actuator/health",
"{#JSON_PATH}": "$.status",
"{#EXPECT_VALUE}": "UP"
},
{
"{#NAME}": "search",
"{#TYPE}": "json_path",
"{#URL}": "http://localhost:9200/_cluster/health",
"{#JSON_PATH}": "$.status",
"{#EXPECT_VALUE}": "green"
},
{
"{#NAME}": "worker",
"{#TYPE}": "json_path",
"{#URL}": "http://localhost:3000/healthz",
"{#JSON_PATH}": "$.ready",
"{#EXPECT_VALUE}": "true"
}
]What you get for shop-api (and the same set for search and worker)
| Object | Name | Severity |
|---|---|---|
| Item | Response shop-api | β |
| Item | JSON path shop-api | β |
| Trigger | Unexpected JSON value for shop-api | HIGH |
The trigger fires when the selected value differs from UP, and also when no value has arrived for 3 minutes, for
example because the service is down, the body is not JSON or the path does not match.
Warn 30 days before the certificate expires, escalate to AVERAGE at 14 days and to HIGH at 7 days.
[
{
"{#NAME}": "www-cert",
"{#TYPE}": "certificate",
"{#URL}": "https://www.example.com",
"{#CERT_WARN_DAYS}": "30",
"{#CERT_AVG_DAYS}": "14",
"{#CERT_HIGH_DAYS}": "7"
}
]What you get for www-cert
| Object | Name | Severity |
|---|---|---|
| Item | Certificate data www-cert | β |
| Item | Certificate expiration www-cert | β |
| Item | Certificate validity www-cert | β |
| Item | Days until certificate expires (www-cert) | β |
| Trigger | Certificate data for www-cert is unavailable | INFO |
| Trigger | Certificate for www-cert is invalid | HIGH |
| Trigger | Certificate for www-cert will expire in 30 or less | WARNING |
| Trigger | Certificate for www-cert will expire in 14 or less | AVERAGE |
| Trigger | Certificate for www-cert will expire in 7 or less | HIGH |
The certificate data is fetched every 15 minutes, the days value is recalculated every 6 hours. Use the name the
certificate was issued for (here www.example.com), not localhost, otherwise the check reports invalid.
A shop server. The public storefront is checked through its own name, once for the status code and once for the
certificate: the same URL may appear once in the HTTP group and once in the certificate group. Two Spring Boot
services answer with JSON health, a protected admin interface is expected to answer 401 without credentials, and
an API on port 8443 has a certificate of its own. billing-db reads a component status, which Actuator only includes
when management.endpoint.health.show-components or show-details is set to always (when-authorized does not
help, the agent sends no credentials).
[
{
"{#NAME}": "shop-frontend",
"{#TYPE}": "http_status",
"{#URL}": "https://shop.example.com/",
"{#EXPECT_STATUS}": "200"
},
{
"{#NAME}": "shop-cert",
"{#TYPE}": "certificate",
"{#URL}": "https://shop.example.com/",
"{#CERT_WARN_DAYS}": "30",
"{#CERT_AVG_DAYS}": "14",
"{#CERT_HIGH_DAYS}": "7"
},
{
"{#NAME}": "orders-api",
"{#TYPE}": "json_path",
"{#URL}": "http://localhost:8080/actuator/health",
"{#JSON_PATH}": "$.status",
"{#EXPECT_VALUE}": "UP"
},
{
"{#NAME}": "billing-db",
"{#TYPE}": "json_path",
"{#URL}": "http://localhost:8081/actuator/health",
"{#JSON_PATH}": "$.components.db.status",
"{#EXPECT_VALUE}": "UP"
},
{
"{#NAME}": "admin-auth",
"{#TYPE}": "http_status",
"{#URL}": "http://localhost:9000/admin/",
"{#EXPECT_STATUS}": "401"
},
{
"{#NAME}": "api-cert",
"{#TYPE}": "certificate",
"{#URL}": "https://api.example.com:8443",
"{#CERT_WARN_DAYS}": "21",
"{#CERT_AVG_DAYS}": "10",
"{#CERT_HIGH_DAYS}": "3"
}
]Both http_status and json_path entries create the raw item web.page.get["{#URL}"], and item keys must be
unique per host. To check the status code and a JSON value of the same endpoint, make the two URLs differ. A
query string that the service ignores does the job; a trailing slash on one of them works as well when the server
treats both paths alike.
[
{
"{#NAME}": "payments-status",
"{#TYPE}": "http_status",
"{#URL}": "http://localhost:8080/actuator/health?check=status",
"{#EXPECT_STATUS}": "200"
},
{
"{#NAME}": "payments-health",
"{#TYPE}": "json_path",
"{#URL}": "http://localhost:8080/actuator/health?check=json",
"{#JSON_PATH}": "$.status",
"{#EXPECT_VALUE}": "UP"
}
]The endpoint is requested twice per interval, once per entry. The same rule applies among certificate entries
(web.certificate.get["{#URL}"]); a URL may however appear once in the HTTP group and once in the certificate
group, as shop-frontend and shop-cert do in example 4.
The macro value is the same JSON without line breaks and indentation. This is example 4 exactly as it goes into the macro field, about 850 of the 2048 characters a macro value can hold:
[{"{#NAME}":"shop-frontend","{#TYPE}":"http_status","{#URL}":"https://shop.example.com/","{#EXPECT_STATUS}":"200"},{"{#NAME}":"shop-cert","{#TYPE}":"certificate","{#URL}":"https://shop.example.com/","{#CERT_WARN_DAYS}":"30","{#CERT_AVG_DAYS}":"14","{#CERT_HIGH_DAYS}":"7"},{"{#NAME}":"orders-api","{#TYPE}":"json_path","{#URL}":"http://localhost:8080/actuator/health","{#JSON_PATH}":"$.status","{#EXPECT_VALUE}":"UP"},{"{#NAME}":"billing-db","{#TYPE}":"json_path","{#URL}":"http://localhost:8081/actuator/health","{#JSON_PATH}":"$.components.db.status","{#EXPECT_VALUE}":"UP"},{"{#NAME}":"admin-auth","{#TYPE}":"http_status","{#URL}":"http://localhost:9000/admin/","{#EXPECT_STATUS}":"401"},{"{#NAME}":"api-cert","{#TYPE}":"certificate","{#URL}":"https://api.example.com:8443","{#CERT_WARN_DAYS}":"21","{#CERT_AVG_DAYS}":"10","{#CERT_HIGH_DAYS}":"3"}]jq produces the compact form from a pretty-printed file, validates it and counts its characters in one go:
jq -c . webcheck.json # validate and print the compact form
jq -cj . webcheck.json | wc -m # count its characters: at most 2048 fit into a macro valueAll seven macros are defined on the template and can be overridden per host (Macros tab, Inherited and host
macros, Change). Time values are Zabbix time expressions such as 30s, 5m, 1h or 2d.
| Macro | Default | Used by | Purpose |
|---|---|---|---|
{$PUSHIT.WEBCHECK.CONFIG} |
[] |
both discovery rules | The endpoint list, a JSON array in LLD format. Set this on every host. |
{$PUSHIT.WEBCHECK.HTTP.INTERVAL} |
1m |
http_status, json_path |
Update interval of the raw item Response {#NAME}. The template also uses it as that item's history period, see Keep the timing consistent. |
{$PUSHIT.WEBCHECK.HTTP_STATUS.NO_STATUS_GRACE} |
3m |
http_status |
How long a missing status code is tolerated before Unexpected status code for {#NAME} fires. Must be larger than HTTP.INTERVAL, otherwise it provides no tolerance and can raise false no-data problems. |
{$PUSHIT.WEBCHECK.JSON_PATH.NO_RESPONSE_GRACE} |
3m |
json_path |
How long a missing response is tolerated before Unexpected JSON value for {#NAME} fires. Must be larger than HTTP.INTERVAL, otherwise it provides no tolerance and can raise false no-data problems. |
{$PUSHIT.WEBCHECK.CERTIFICATE.INTERVAL} |
15m |
certificate |
Update interval of the raw item Certificate data {#NAME}. |
{$PUSHIT.WEBCHECK.CERTIFICATE.NO_DATA_GRACE} |
30m |
certificate |
How long missing certificate data is tolerated before Certificate data for {#NAME} is unavailable (INFO) fires. Must be larger than CERTIFICATE.INTERVAL and smaller than CERTIFICATE.HISTORY. |
{$PUSHIT.WEBCHECK.CERTIFICATE.HISTORY} |
90m |
certificate |
History retention of the raw certificate JSON only. Must be larger than CERTIFICATE.NO_DATA_GRACE, should be larger than CERTIFICATE.INTERVAL and must be at least 1h. The derived items are not affected. |
Timing at a glance, with the default values:
| Check type | Raw item collected every | No-data grace before the trigger fires | Raw history kept for |
|---|---|---|---|
http_status |
{$PUSHIT.WEBCHECK.HTTP.INTERVAL} = 1m |
{$PUSHIT.WEBCHECK.HTTP_STATUS.NO_STATUS_GRACE} = 3m |
{$PUSHIT.WEBCHECK.HTTP.INTERVAL} = 1m, see note below |
json_path |
{$PUSHIT.WEBCHECK.HTTP.INTERVAL} = 1m |
{$PUSHIT.WEBCHECK.JSON_PATH.NO_RESPONSE_GRACE} = 3m |
{$PUSHIT.WEBCHECK.HTTP.INTERVAL} = 1m, see note below |
certificate |
{$PUSHIT.WEBCHECK.CERTIFICATE.INTERVAL} = 15m |
{$PUSHIT.WEBCHECK.CERTIFICATE.NO_DATA_GRACE} = 30m |
{$PUSHIT.WEBCHECK.CERTIFICATE.HISTORY} = 90m |
Two rules follow from this:
- Grace periods must be larger than the interval. The grace period is the argument of
nodata(). If it is equal to or shorter than the interval, the window runs out between two regular polls and the trigger can raise problems on a healthy endpoint. The defaults (1m interval, 3m grace) tolerate two missed polls. CERTIFICATE.INTERVAL < CERTIFICATE.NO_DATA_GRACE < CERTIFICATE.HISTORY. The certificatenodata()trigger watches the raw item Certificate data {#NAME}, whose history is set byCERTIFICATE.HISTORY, so the history has to outlast the grace period. When you raise one, raise the other with it. The HTTPnodata()triggers watch the dependent items, which keep the Zabbix default history of 31d, so no such rule applies to them.
A note on the raw HTTP history: Zabbix accepts history periods from 1h to 25y (or 0), and Response {#NAME}
reuses HTTP.INTERVAL as its history period. With the default 1m that period is out of range, so the housekeeper
logs invalid history storage period for the item and skips it, and the raw responses stay as long as the item
exists. If you want them cleaned up, set HTTP.INTERVAL to 1h or more, or give the item prototype a history
period of its own after import. The last value is always visible in Latest data.
Every discovered object carries the {#NAME} of its entry, so items and triggers of different services never
collide and are easy to tell apart in Latest data and Problems.
| Discovery rule | Key | Type | Runs | Accepts entries with | Lost resources |
|---|---|---|---|---|---|
| HTTP | pushit.webcheck.discovery.http |
Script | every 10m | {#TYPE} matching ^http_status$ or ^json_path$ |
Delete immediately |
| Certificate | pushit.webcheck.discovery.certificate |
Script | every 10m | {#TYPE} matching ^certificate$ |
Delete immediately |
Both rules receive {$PUSHIT.WEBCHECK.CONFIG} as the script parameter config and return the value of
JSON.parse(value).config.
The HTTP rule creates its two dependent prototypes with Discover set to No and switches on the matching one with an override, so each HTTP entry ends up with the raw item plus exactly one dependent item:
| Override | Condition | Effect |
|---|---|---|
| Enable http_status | {#TYPE} matches ^http_status$ |
Item prototypes whose name matches ^HTTP status.*$ are discovered |
| Enable json_path | {#TYPE} matches ^json_path$ |
Item prototypes whose name matches ^JSON path.*$ are discovered |
http_status: items and triggers
| Item | Key | Type | Value type | Interval | History |
|---|---|---|---|---|---|
| Response {#NAME} | web.page.get["{#URL}"] |
Zabbix agent | Text | {$PUSHIT.WEBCHECK.HTTP.INTERVAL} |
{$PUSHIT.WEBCHECK.HTTP.INTERVAL}, see Keep the timing consistent |
| HTTP status {#NAME} | pushit.webcheck.http.status[{#NAME}] |
Dependent on Response {#NAME} | Numeric (unsigned) | with the master item | 31d (Zabbix default) |
Response {#NAME} holds the raw response: status line, headers and body. HTTP status {#NAME} extracts the
three-digit status code from the status line with one Regular expression preprocessing step, pattern
\AHTTP/[0-9.]+[ \t]+([0-9]{3})(?:[ \t]|\r?\n) and output \1.
| Trigger | Severity | Condition |
|---|---|---|
| Unexpected status code for {#NAME} | HIGH | No status code for {$PUSHIT.WEBCHECK.HTTP_STATUS.NO_STATUS_GRACE}, or the last code differs from {#EXPECT_STATUS}. |
nodata(/PUSH IT Webcheck/pushit.webcheck.http.status[{#NAME}],{$PUSHIT.WEBCHECK.HTTP_STATUS.NO_STATUS_GRACE})=1
or
last(/PUSH IT Webcheck/pushit.webcheck.http.status[{#NAME}])<>{#EXPECT_STATUS}
json_path: items and triggers
| Item | Key | Type | Value type | Interval | History |
|---|---|---|---|---|---|
| Response {#NAME} | web.page.get["{#URL}"] |
Zabbix agent | Text | {$PUSHIT.WEBCHECK.HTTP.INTERVAL} |
{$PUSHIT.WEBCHECK.HTTP.INTERVAL}, see Keep the timing consistent |
| JSON path {#NAME} | pushit.webcheck.http.json_path[{#NAME}] |
Dependent on Response {#NAME} | Text | with the master item | 31d (Zabbix default) |
Response {#NAME} is the same raw item as for http_status. JSON path {#NAME} has two preprocessing steps: a
Regular expression step with the pattern \r?\n\r?\n([\s\S]*) and the output \1 drops the headers and keeps the
body, then a JSONPath step applies {#JSON_PATH}.
| Trigger | Severity | Condition |
|---|---|---|
| Unexpected JSON value for {#NAME} | HIGH | No value for {$PUSHIT.WEBCHECK.JSON_PATH.NO_RESPONSE_GRACE}, or the last value differs from {#EXPECT_VALUE}. |
nodata(/PUSH IT Webcheck/pushit.webcheck.http.json_path[{#NAME}],{$PUSHIT.WEBCHECK.JSON_PATH.NO_RESPONSE_GRACE})=1
or
last(/PUSH IT Webcheck/pushit.webcheck.http.json_path[{#NAME}])<>"{#EXPECT_VALUE}"
certificate: items and triggers
| Item | Key | Type | Value type | Interval | History |
|---|---|---|---|---|---|
| Certificate data {#NAME} | web.certificate.get["{#URL}"] |
Zabbix agent (agent 2) | Text | {$PUSHIT.WEBCHECK.CERTIFICATE.INTERVAL} |
{$PUSHIT.WEBCHECK.CERTIFICATE.HISTORY} |
| Certificate expiration {#NAME} | pushit.webcheck.certificate.not_after[{#NAME}] |
Dependent on Certificate data {#NAME} | Numeric (unsigned) | with the master item | 31d (Zabbix default) |
| Certificate validity {#NAME} | pushit.webcheck.certificate.is_valid[{#NAME}] |
Dependent on Certificate data {#NAME} | Text | with the master item | 31d (Zabbix default) |
| Days until certificate expires ({#NAME}) | pushit.webcheck.certificate.days_until_expiration[{#NAME}] |
Calculated | Numeric (float), unit days |
6h |
31d (Zabbix default) |
Certificate data {#NAME} holds the raw certificate JSON. Certificate expiration {#NAME} takes
$.x509.not_after.timestamp from it (the notAfter date as a Unix timestamp), Certificate validity {#NAME}
takes $.result.value (valid, invalid or valid-but-self-signed), and Days until certificate expires
({#NAME}) is calculated as round((last(//pushit.webcheck.certificate.not_after[{#NAME}]) - now()) / 86400, 1).
| Trigger | Severity | Condition |
|---|---|---|
| Certificate data for {#NAME} is unavailable | INFO | No certificate data for {$PUSHIT.WEBCHECK.CERTIFICATE.NO_DATA_GRACE}. |
| Certificate for {#NAME} is invalid | HIGH | The validity value is anything other than valid. |
| Certificate for {#NAME} will expire in {#CERT_WARN_DAYS} or less | WARNING | Days until expiry <= {#CERT_WARN_DAYS}. |
| Certificate for {#NAME} will expire in {#CERT_AVG_DAYS} or less | AVERAGE | Days until expiry <= {#CERT_AVG_DAYS}. |
| Certificate for {#NAME} will expire in {#CERT_HIGH_DAYS} or less | HIGH | Days until expiry <= {#CERT_HIGH_DAYS}. |
nodata(/PUSH IT Webcheck/web.certificate.get["{#URL}"],{$PUSHIT.WEBCHECK.CERTIFICATE.NO_DATA_GRACE})=1
last(/PUSH IT Webcheck/pushit.webcheck.certificate.is_valid[{#NAME}]) <> "valid"
last(/PUSH IT Webcheck/pushit.webcheck.certificate.days_until_expiration[{#NAME}]) <= {#CERT_WARN_DAYS}
last(/PUSH IT Webcheck/pushit.webcheck.certificate.days_until_expiration[{#NAME}]) <= {#CERT_AVG_DAYS}
last(/PUSH IT Webcheck/pushit.webcheck.certificate.days_until_expiration[{#NAME}]) <= {#CERT_HIGH_DAYS}
The three expiry triggers are independent of each other: once the days drop to the HIGH threshold, all three are in problem state.
- One URL per entry within a rule. Two
http_status/json_pathentries, or twocertificateentries, cannot share a URL because the raw item key contains it. Example 5 shows the query string trick. - Renaming an entry recreates it.
{#NAME}is part of the item keys, so a renamed entry gets new items and triggers and loses the history of the old ones. - Plain requests from the agent.
web.page.getandweb.certificate.getsend unauthenticated requests without custom headers or body, so point them at endpoints that answer anonymously, typically a health endpoint onlocalhost. To verify that a protected endpoint is up, usehttp_statuswith{#EXPECT_STATUS}"401". The requests originate on the monitored host: firewalls between the agent and the service matter, firewalls between the Zabbix server and the service do not. - Certificate checks need
httpsand Zabbix agent 2.web.certificate.getaccepts only thehttpsscheme and exists only in agent 2. On a host with the classic agent, Certificate data {#NAME} becomes Not supported and the only problem you will see is the INFO trigger after 30 minutes. httpsin HTTP checks needs cURL in the classic agent. The classic Zabbix agent must be built with cURL support to fetchhttpsURLs withweb.page.get, otherwise the item becomes Not supported. Zabbix agent 2 has no such requirement.- Point certificate checks at the certificate's name. The certificate is validated for the host name in the
URL, so
https://shop.example.comreportsvalidwherehttps://localhostreportsinvalidfor the same certificate. Self-signed certificates yieldvalid-but-self-signed, which also fires Certificate for {#NAME} is invalid. - Removed means gone. There is no way to pause a check from the configuration: an entry taken out of the macro loses its items, triggers and history on the next discovery run.
- Invalid JSON stops discovery. If the macro is not valid JSON, both discovery rules turn Not supported with the parse error shown in the rule status. Existing items stay as they are until the macro is fixed.
- Unknown
{#TYPE}values are ignored. A typo such ashttp-statusmatches neither rule and produces no items, no triggers and no error. Check Latest data after adding an entry. - The macro value is limited to 2048 characters. Keep the JSON compact and the names and URLs short. Depending on URL length, somewhere between a dozen and twenty entries fit into one macro. Count the characters as shown in compact form.
- Missing certificate data is only INFO. The nodata trigger of the raw certificate item has severity INFO. The alerts that matter come from the validity and expiry triggers. If you want it louder, raise the severity of the trigger prototype Certificate data for {#NAME} is unavailable after import.
- The days value moves every 6 hours. Days until certificate expires ({#NAME}) is a calculated item with a 6-hour interval. The first value appears shortly after discovery, afterwards it is refreshed only every 6 hours, so after a renewal the expiry problems resolve on the next run. Execute now on the item refreshes it at once.
- Slow endpoints and the item timeout. The item prototypes set no timeout of their own, so the global timeout for Zabbix agent items (default 3 seconds, Administration β General β Timeouts) or the proxy's timeout applies. Raise it if a health endpoint needs longer than that.
- Test a check by hand. Run the item key on the monitored host with the agent binary, for example
zabbix_agent2 -t 'web.page.get["http://localhost:8080/health"]'orzabbix_agent2 -t 'web.certificate.get["https://www.example.com"]'(zabbix_agentd -t ...for the classic agent), or from the server withzabbix_get -s <agent address> -k 'web.page.get["http://localhost:8080/health"]'. The returned value is exactly what Response {#NAME} or Certificate data {#NAME} will store.
The whole template is template.yaml, a Zabbix 7.4 YAML export. Keep it lint-clean with
yamllint:
pip install yamllint # or the package of your distribution
yamllint template.yamlThe rules live in .yamllint.yaml: yamllint's defaults plus a 140-character line limit and
single quotes wherever strings are quoted.
To verify a change functionally, import the file into a test Zabbix (importing again updates the existing
template), link it to a host with an agent, set a small {$PUSHIT.WEBCHECK.CONFIG} and run Execute now on both
discovery rules. Keep the uuid values in the file: they tie every object to its counterpart in an existing
installation, so re-importing updates instead of duplicating.
Apache License 2.0, see LICENSE.