WIP: an RFC for check-based monitoring semantics
Check-based monitoring has run production floors for twenty-five years and nobody ever wrote down what it means. A scheduler runs probes, each probe exits with a code, state machines turn streams of results into confirmed states, and a pager goes off on confirmed transitions. Every tool in the Nagios lineage implements this. No two of them agree at the edges, because there is no spec to agree with: the exit-code convention, the host/service split, UP versus OK, UNREACHABLE, soft and hard states, acknowledgements and downtimes are folklore, documented separately and inconsistently inside each tool that inherited them.
I’ve been building jaque, a monitoring engine in that lineage, and I kept hitting the same wall: every design question (“can a host be WARNING?”, “is UNREACHABLE a state or a rendering?”) had no authority to appeal to, only precedent, and the precedent disagreed with itself. The nearest prior art is ITU-T X.733’s alarm severities and the Monitoring Plugins development guidelines, and neither covers object kinds, reachability, or what a frontend may assume about a row it didn’t produce. So I started writing the spec I wanted to cite.
This is that draft, work in progress, published here because a spec nobody can read is a private opinion. jaque is the reference implementation; nothing below is jaque-specific. If you’ve shipped or operated anything in this lineage and a MUST below contradicts your scars, I’d rather hear about it than not.
1. Conventions
The key words MUST, MUST NOT, SHOULD, SHOULD NOT and MAY are to be interpreted as described in RFC 2119.
“Implementation” means the engine holding the state machines. “Consumer” means anything reading externalized state: a frontend, an exporter, an aggregator, another monitoring system.
2. Object model
2.1 Kinds
Every monitored object has exactly one kind, fixed at creation, part of its identity.
| Kind | Numeric | Definition |
|---|---|---|
| host | 0 | A thing with an address that can be reachable or unreachable; may own services; participates in the parents graph. |
| service | 1 | A named aspect of exactly one host, checked independently of it. |
| process | 2 | A derived object whose status is computed by a rule over other objects’ confirmed states; never scheduled, never executed. |
Kind MUST be carried explicitly on every externalized object. A consumer MUST NOT infer kind from name shape, from list membership, or from word choice. The numeric values are normative and MUST NOT be reassigned; extensions add kinds, they never renumber.
2.2 Identity
Identity is kind plus name(s): a host by name, a service by
host + service, a process by name. Textual form: name,
host/service, bp:name. The bp: prefix is reserved: no host name may
begin with it. A process name colliding with a host name is a
definition-time error.
Matching identities across systems MUST be exact on (kind, names). Treating two objects as the same because their names collide across kinds is a defect.
3. Status model
An object’s state is one severity enum plus three orthogonal qualifiers. The severity is the stored truth; every word an operator reads is derived from it at presentation time (section 8) and MUST NOT be stored.
3.1 Status
PENDING (0), OK (1), WARNING (2), CRITICAL (3), UNKNOWN (4).
PENDINGmeans no result has ever been received. It is not an error state and MUST NOT be produced by a check.- Severity order for rule evaluation and worst-first sorting is
OK < WARNING < UNKNOWN < CRITICAL. UNKNOWN sits below CRITICAL because “I could not determine the state” is worse than a warning and not as actionable as a confirmed failure.
One Status set serves every kind. An implementation MUST NOT define a separate host status enum: collapsing a host result into a binary up/down at check time destroys severity information that cannot be recovered, while the full severity can always be projected down at any edge that wants binary words (section 9).
3.2 State type
SOFT or HARD. SOFT means the state is still being confirmed by
retries; HARD means it is confirmed (section 5).
3.3 Reachable
A boolean, meaningful for hosts, inherited by their services, always true for processes. It is derived from the parents graph (section 6), never stored as a status, and MUST be recomputable at any time from the graph plus current hard states.
3.4 Flapping
A boolean: the object’s state is oscillating faster than the configured tolerance (section 7.1).
4. Check result protocol
A check result carries: the object identity, a Status (never PENDING), one line of human-readable output, optional additional output lines, optional performance data, and the execution timestamp.
4.1 Exit-code binding
For process-executed checks (the Monitoring Plugins convention, kept verbatim for compatibility with two decades of scripts):
| Exit code | Status |
|---|---|
| 0 | OK |
| 1 | WARNING |
| 2 | CRITICAL |
| 3, or any other value | UNKNOWN |
A timeout or a failure to execute the check at all MUST map to UNKNOWN, not CRITICAL: the check said nothing about the object.
4.2 Output
The first line of stdout is the status line. Text after a | on the
status line, and everything after a | on the last line of long output,
is performance data.
4.3 Performance data
Space-separated series, each:
'label'=value[UOM];[warn];[crit];[min];[max]
Quotes around the label are required only when it contains spaces.
value is a number or U (undetermined). UOM is one of: none, s
(seconds, with ms/us accepted as inputs), %, B (bytes, with
KB/MB/GB/TB accepted), or c (a monotonic counter). Threshold
fields follow the Monitoring Plugins range grammar and MAY be empty.
Implementations SHOULD normalize values to canonical units (seconds, bytes, percent) at storage or export boundaries, preserving the raw UOM as metadata; consumers MUST NOT be required to re-parse UOM suffixes.
4.4 Passive results
A result submitted by an external producer instead of a scheduled check MUST be indistinguishable to the state machine from an active result: same fields, same transitions, same gating.
5. State machine
- A non-OK result on an object in OK or PENDING enters SOFT at the result’s severity and schedules rechecks at the retry interval.
- After
max_check_attemptsconsecutive non-OK results, the state becomes HARD at the latest severity.max_check_attemptsof 1 makes every non-OK result immediately HARD. - An OK result at any point returns the object to OK. Recovery from a HARD problem is itself a HARD transition; recovery from SOFT is not.
- A severity change between non-OK states while HARD (e.g. WARNING to CRITICAL) is a HARD transition.
- Transitions are deterministic: replaying the same result stream from the same initial state MUST yield the same transition stream.
6. Reachability
Hosts form a dependency graph via parent declarations. When every path
from the monitoring vantage point to a host passes through at least one
host in HARD problem state, the host is unreachable: Reachable is
false.
- Reachability is derived. The stored Status of an unreachable host is whatever its last check said; an implementation MUST NOT overwrite it with a synthetic “unreachable” status.
- Services inherit their host’s reachability. Processes are always reachable: a computation cannot be behind a dead router.
- Problem notifications for unreachable objects MUST be suppressed; the root object (the nearest failed parent) keeps notifying. This is the point of the mechanism: one router down is one page instead of forty.
7. Suppression
Three independent mechanisms gate notifications. They compose: a notification fires only when none of them suppresses it.
7.1 Flap detection
An implementation MUST detect state oscillation and suppress problem and recovery notifications while it persists. Detection MUST use hysteresis (a higher threshold to enter flapping than to leave it) and MUST converge: when the input stabilizes, flapping ends. Entering and leaving flapping are themselves notifiable meta-events and MAY fire from SOFT states. The exact algorithm (windowed change rate, EWMA, weights) is implementation-defined.
7.2 Acknowledgements
An acknowledgement attaches to a problem state and silences further notifications for it. A sticky ack survives severity changes between non-OK states; a non-sticky ack clears on any state change. Every ack clears on recovery. An ack MUST record authorship.
7.3 Downtimes
A downtime is a scheduled window during which notifications for the object are suppressed and its transitions are marked as in-downtime. Fixed downtimes have a start and an end; flexible downtimes have a duration that starts counting at the first problem inside the window.
8. Presentation vocabulary
The word an operator reads is a pure function of (kind, Status, Reachable). The word alone identifies the kind; that is the cheapest kind signal a dense list can carry.
| Status | host | service | process |
|---|---|---|---|
| PENDING | PENDING | PENDING | PENDING |
| OK | UP | OK | OK |
| WARNING | WARNING | WARNING | WARNING |
| CRITICAL | DOWN | CRITICAL | CRITICAL |
| UNKNOWN | UNKNOWN | UNKNOWN | UNKNOWN |
| any, Reachable=false | UNREACHABLE | UNREACHABLE | n/a |
Departures from the Nagios lineage, both in the direction of not destroying information:
- A host in WARNING or UNKNOWN displays that word. The lineage collapses host results to UP/DOWN at check time; this specification keeps full severity and maps words only at the edge.
- UNREACHABLE is a presentation of
Reachable, never a stored status.
Rules for consumers:
- An aggregate over one kind labels its buckets in that kind’s vocabulary: a count of hosts says up, down, unreachable, never ok and critical.
- Mixed-kind lists keep kinds distinguishable per row; single-kind lists state their kind once, in the frame.
- A row for an object with no check of its own says NO CHECK, never PENDING. PENDING asserts a result is still awaited; that is a different claim.
- SOFT and FLAPPING are qualifiers appended to the status word, never substitutes for it.
- A derived object carries its own kind. Re-exporting a process as a synthetic host is out of spec.
9. Legacy projections
Compatibility with the lineage’s wire formats is a set of lossy projections applied at each edge, never a constraint on the core model.
- Livestatus host state: 2 when Reachable is false; otherwise 0 for OK
or PENDING, 1 for anything else. Service state: the Status numeric
minus one, PENDING reported via
has_been_checked = 0. - Nagios-style host words: project through section 8’s table.
A projection MUST be documented as lossy where it is (a DOWN reported to livestatus does not say whether the stored Status was WARNING, CRITICAL or UNKNOWN), and an implementation MUST NOT round-trip its own state through a lossy projection.
10. Event sourcing (informative)
This specification does not require event sourcing, but its determinism requirements (section 5) are most cheaply met by one: every domain fact as an appended event, current state as a replayable projection. Implementations using other persistence MUST still satisfy section 5’s replay determinism at the level of externally observable transitions.
11. Stability
While this document is a draft, anything may change. On acceptance: the kind numerics (2.1), the Status set and order (3.1), the exit-code binding (4.1) and the section 8 table freeze. Extensions add kinds and words; they never reassign numerics or change an existing mapping.
12. Prior art
- ITU-T X.733: alarm severities, no object kinds, no reachability.
- Monitoring Plugins Development Guidelines: exit codes, output and perfdata syntax for plugin authors; nothing on states, kinds or suppression.
- Livestatus (MK Livestatus / Checkmk): a query wire protocol over an implied model this document makes explicit.
- OpenTelemetry entities (OTEP 0264, formerly the entities data model of OTEP 0256): typed identity, no per-kind health-state vocabulary. The candidate upstream venue if this document stabilizes.