How alerting works

Where alerts come from, what they monitor, and how flood control groups them during a spike.

The short answer

HC watches every job and event and emails you when one fails. During a spike it groups failures into a single notification rather than sending one email per failure.

How it works

What alerts monitor

Alerts cover the three points where things can fail: a whole job failing, an event failing during transformation, and an event failing at the destination system.

What an alert contains

Each alert names the issue and error message, the entity affected, the stream name, the event ID and the job ID, and links directly to the event in the Control Panel. Alert emails are sent from alerts@highcohesion.com.

Flood control

To prevent alert fatigue during spikes, flood control groups alerts rather than sending one email per failure. Once the configured threshold is reached within the time window, for example 20 errors in 15 minutes, a single notification is sent covering all the grouped failures, with a note of how many occurred. Thresholds are configured per stream by HC.

Summary reports

Daily and weekly summary reports show alert counts per stream and type. The daily report is sent at 16:00 UTC and the weekly report on Fridays at 16:00 UTC, from insight@highcohesion.com with a DAILY REPORT or WEEKLY REPORT subject prefix. They are especially useful when flood control has been grouping alerts.

What this means for you

  • A high volume of alerts in a short period usually has a single root cause. A destination system outage, for example, can generate hundreds of individual failed events. Fix the root cause first, then ask HC to resend the affected events in bulk.

  • If your alert volume feels wrong in either direction, ask HC to adjust the flood-control thresholds for the stream.

  • Add alerts@highcohesion.com and insight@highcohesion.com to your safe senders list so alerts and reports are not filtered.

Related