Why Your Team Stops Trusting Safety AI by Week 3 (and How to Prevent It)

Computer-vision safety systems rarely fail on detection. They fail on trust. When an uncalibrated system floods EHS teams with false alarms, operators quietly stop responding, usually within three weeks of go-live. This guide draws on alarm-fatigue research from hospitals, security operations and physical security, shows the base-rate maths behind false positives at industrial scale, and sets out a five-part calibration discipline: silent baselining, a written false-positive target, per-camera tuning, weekly measured accuracy and an agreed alert budget. It also includes a field guide to the five most common industrial confusers, a 10-question vendor checklist, a 90-day pilot plan and an FAQ for EHS and operations leaders.
Wojciech Tubek
CEO @ Surveily
•
18 minutes
 read
Fabrication hall CCTV view with AI detection boxes: a real near miss between a worker and a moving forklift, and false positives on skylight glare (smoke), wall signs (person) and a red barrel (fire). People are blurred.

Written for EHS directors, site safety managers and operational excellence leaders, and for anyone who has watched a promising monitoring tool die in its first month.

The short version

Most computer-vision safety deployments don’t fail because the model misses hazards. They fail because it flags too many things that aren’t hazards, and the people on the receiving end stop looking. The failure is behavioural, it’s predictable, and it usually plays out within three weeks of go-live. Week one is enthusiasm. Week two is triage. Week three is the mental filter that files every alert under “probably nothing again.”

This isn’t speculation. It’s one of the most consistently replicated findings in every field that has industrialised alerting:

  • Hospitals: published studies report that 74–99% of physiological monitor alarms are non-actionable, with ICU monitors producing hundreds of alarms per patient per day. This is the research that gave “alarm fatigue” its name.
  • Security operations centres: a 2023 industry study found roughly 83% of daily alerts are false alarms and 67% eventually go ignored because analysts can’t keep up. One well-known qualitative study of SOC analysts is titled, without irony, “99% False Positives.”
  • Physical security: industry analyses put the share of camera-triggered alarms that are false at up to ~98%, driven by weather, animals, lighting and vegetation.

Workplace safety AI inherits exactly this dynamic, with one aggravating factor. The people receiving the alerts (safety managers, shift leaders, line supervisors) already have full-time jobs, and alert triage isn’t one of them. The margin for noise on a plant floor is thinner than in a SOC, not wider.

The fix isn’t a better launch-day accuracy number. It’s a calibration discipline: per-camera baselining, a bounded tuning period with a numeric false-positive target, weekly measured accuracy reporting, and an operating model that treats precision as something you manage every week, not a claim in a brochure. Below are the evidence, the arithmetic, and a checklist you can apply to any pilot, including ours.

Illustrative chart: without calibration, alerts stay near 100 per shift and the share opened falls from 95% to 12% by day 30; with calibration, alerts fall to about 16 and the opened share stays at 90–95%.
Figure 1. The week-3 curve: alert volume vs. operator response (illustrative)

1. Alert fatigue is a law of operations, not a software bug

Every monitoring discipline that scaled up alerting has rediscovered the same law: human attention is a finite resource, and false alarms drain it faster and faster. The psychology is old (the “cry-wolf effect” has been studied since the 1980s), and the operational data is unambiguous.

Healthcare ran the experiment first and at the largest scale. Reviews of clinical alarm studies consistently find that 74–99% of physiological monitor alarms are non-actionable, and intensive-care monitors can exceed 350 alarms per patient per day. The documented consequences include delayed responses, alarms switched off, and patient deaths. That’s why alarm management became a named patient-safety goal in US hospital accreditation. The core finding transfers directly to industry: after a handful of false alarms in a row, operators stop trusting the channel itself, not just the individual alert.

Cybersecurity repeated it. Industry surveys report that around 83% of everyday SOC alerts are false positives, 73% of security teams name false positives as their top detection challenge, and 25–30% of alerts are never investigated at all because of overload. Tiered triage, suppression rules and correlation engines now form a whole sub-industry, and its existence is proof of the problem.

Physical security completed the pattern. Analyses of camera-based alarm systems report false-alarm shares as high as 98%. Police forces in several countries now fine repeat false alarms or stop responding to them. That’s distrust written into public policy.

The lesson for safety AI buyers isn’t that alerting is hopeless. It’s that alert volume and alert precision are core product variables, as important as what the system can detect. If a vendor tells you what their system detects but not how its precision is measured, managed and reported, you’ve had half a conversation.

2. The arithmetic: why “95% accurate” can still bury a safety team

Here’s the calculation every buyer should run before a pilot. It explains week 3 better than any anecdote.

Take a mid-sized site with 40 cameras watching continuously. In effect, the system makes thousands of classification decisions per camera per day. Real hazard events, meanwhile, are rare, which is exactly what makes them worth catching. Rare targets are the worst case for precision. This is the base-rate problem: even a low per-decision false-positive rate, multiplied across 40 cameras and 24 hours, can produce false alerts that outnumber real events many times over.

With illustrative numbers: a site has 20 real hazard events a day and the system raises 100 alerts. Precision is 20%. Four out of five alerts waste an operator’s time, and within weeks the system has trained its users to ignore it. Cut the stream to 25 alerts a day with the same detections and precision jumps to 80%. That’s the difference between a channel people trust and a channel people mute. The model’s headline accuracy is identical in both scenarios. What changed is calibration to the site.

Three practical consequences:

  1. Headline accuracy is an average, and averages hide cameras. A deployment can run at a ~95% weighted average while camera 17, facing low sun every afternoon, produces half the site’s false alerts. Per-camera measurement is non-negotiable. Treat any accuracy figure, ours included, as a measured, continuously reported number, never a guarantee.
  2. Precision and recall trade off, so manage the trade-off explicitly. Tighten thresholds and you risk missing real events. Loosen them and you flood the channel. Mature operations set this per use case: a man-down detection can tolerate more false positives than a PPE check, because missing it costs far more.
  3. Routing is part of accuracy. A correct detection that reaches the wrong person at the wrong time is, operationally, a false positive. Severity tiers, zone-based routing and shift-aware delivery are calibration tools, not UX extras.
Illustrative icon array: 100 alerts with 20 real events gives 20% precision; 25 alerts with the same 20 events gives 80% precision.
Figure 2. The base-rate effect: same detector, two precision realities (illustrative)

3. Anatomy of the week-3 failure

The failure has a repeatable anatomy. It’s worth naming each stage, because each one is observable, and each one can be interrupted.

Days 1–5: the novelty surplus. Every alert gets opened. False positives are forgiven, even enjoyed, and screenshots get passed around, because the system is new and the real catches are genuinely impressive. Management sees the first dashboard and is delighted. This is the most dangerous phase: the deployment looks like a success while the alert economics are already unsustainable.

Days 6–12: silent triage. Operators build private rules of thumb: ignore that camera, ignore that rule after dark, check alerts twice a shift instead of live. Response times stretch. Nobody reports it, because it feels like efficiency, not failure.

Days 13–21: the trust collapse. The rules of thumb harden into a norm, and the norm spreads (“don’t bother with the walkway ones, it’s always the shadow”). The one real near miss that arrives on day 19 gets the same treatment as the noise. In hospitals, this is the documented mechanism by which alarm fatigue kills. On an industrial site, it’s how a safety system turns into shelfware with a subscription.

Day 22 onward: the quiet renegotiation. Alerts get rerouted to an inbox nobody owns. The vendor’s quarterly review shows detection counts going up. The site’s engagement numbers have collapsed. Six months later the renewal conversation is about price, because nobody can say what the value was.

The key point: every stage leaves measurable traces: alert volumes, open rates, time-to-acknowledge, false-positive share per camera. A deployment that tracks these from day one can see the collapse coming ten days before it happens. A deployment that only counts detections can’t. These are leading indicators for the safety system itself, and they deserve the same attention as the leading indicators it reports from the floor.

4. The calibration discipline that prevents it

Deployments that survive month three rarely differ from the ones that don’t in model architecture. They differ in whether calibration was run as a bounded, measured, openly reported process. That discipline has five parts.

4.1 Baseline before judgement

For the first days after connection, the system observes and the team measures. What does each camera see? What does the raw alert stream look like? Which rules fire, on which zones, at which hours? Nobody should be asked to act on alerts during baselining, and nobody should be shown a dashboard presented as the truth. The baseline is the denominator for everything that follows, including the before-and-after evidence the deployment will one day owe its budget sponsor.

4.2 A numeric target, on a clock

Calibration without a number and a deadline is just vibes. Good practice is an explicit target: within a fixed calibration window (measured in days, not quarters), false positives below a stated share of alerts, measured per use case and per camera. For example, a two-week window with a single-digit false-positive target. The exact numbers matter less than three things. The target is written down before the pilot starts. It’s measured by counting, not by impression. And it’s validated openly with the customer, alert by alert, in a shared review. A vendor’s willingness to be measured this way is one of the strongest selection signals you have. Even so, treat the figure as a target to be demonstrated, not a contractual guarantee to be assumed.

4.3 Per-camera, per-rule tuning by people who actually watch the footage

Industrial sites defeat generic models in mundane ways. Reflective vests that read as missing from certain angles. Steam plumes that read as smoke. A convex mirror that doubles every forklift. Fixing these isn’t retraining from scratch. It’s systematic per-camera work: zone masks, threshold adjustments, rule scheduling (the dock-door rule doesn’t need to run when the dock is closed), and model feedback on the site’s specific visual quirks. The question to ask any vendor: who does this work, how many hours are budgeted, and how do we see its effect week over week? A vendor with dedicated data specialists doing tuning as a service is describing a process. A vendor who hands you a sensitivity slider is describing your future weekends. Privacy matters here too. Tuning means reviewing event clips, so review should run on de-identified footage by default (faces and bodies blurred or shown as silhouettes). The job is to find the missing hard hat, never to identify who wasn’t wearing it.

4.4 Weekly measured accuracy: the weighted average and the per-camera view

After go-live, precision is a managed metric, permanently. The rhythm that works is a weekly accuracy report that gives a weighted average across active use cases and shows the per-camera distribution. That way a degrading camera (new lighting, moved racking, a seasonal change in sun angle) is caught by the report, not by operators quietly learning to ignore it. Accuracy drifts in production. Sites change, seasons change, camera housings get dirty. A ~95% weighted average that is measured and continuously reported is worth more than any launch-day claim, precisely because it can be proven wrong every week.

4.5 An alert economy the receiving team actually agreed to

Calibration ends with an explicit agreement with the people on the receiving end. Which alerts arrive in real time, and to whom? Which roll up into a daily digest? Which only show up as trends on a dashboard? One rule of thumb holds in every alerting discipline: if an alert type has no named owner and no defined response, it shouldn’t be a real-time alert. Severity tiers, zone routing and digest windows keep a safety team in the loop without drowning in the stream. And the loop always ends with a person: the system flags, people decide what happens next.

What calibrated looks like: measured outcomes

Deployments run this way produce a signature that is the exact opposite of the week-3 curve: alert volume falls, detection of real events holds, and engagement stays high. Published and customer-reported figures from industrial safety AI show the size of the effect. Each is specific to one site, scope and period, so read them as proof that it can be done, not as promises: a European refinery recorded an 85% reduction in alerts over six months while PPE compliance rose; a steel plant recorded 90% fewer total alerts and 92% fewer vehicle–pedestrian near-traverse alerts over six months; a combined heat-and-power operator recorded an 89% reduction in total alerts. Fewer alerts with detection maintained is what a system earning trust looks like in the data.

5. Field guide: the five industrial confusers

Every industrial site fools generic vision models in its own dialect, but the failure families repeat. Knowing them makes vendor conversations concrete. Instead of “how accurate is it?”, ask “how do you handle these five, and can I see the per-camera evidence?”

Five causes of false positives (optics and weather, steam and dust, reflections, PPE look-alikes, legitimate exceptions), each with its fix. The first three need tuning, the last two human judgement.
Figure 3. The five industrial confusers and how to fix them

1. Optics and weather. Low sun through dock doors, lens flare, rain and snow on camera domes, fog, condensation cycles in cold stores. Signature: false positives clustered at specific hours and seasons on specific cameras. Fixes: per-camera scheduling and thresholds, exposure-aware tuning, and sometimes the honest answer that a camera needs moving or a hood. A vendor willing to say “move the camera” has a healthy calibration culture.

2. Steam, dust and heat shimmer. The heavy-industry classics: steam read as smoke, dust clouds read as fire, shimmer distorting person detection near furnaces. Signature: false alarms on critical rules like smoke and fire. These are the most expensive false positives in the system, because they trigger the most disruptive responses. Fixes: zone-specific tuning, multi-frame confirmation, and severity-tier routing so unconfirmed detections arrive as items to review, not sirens.

3. Reflections and doubles. Polished floors, glass partitions, mirrors, wet surfaces after cleaning: each can double a forklift or put a phantom person inside an exclusion zone. Signature: physically impossible detections, like a person “inside” a wall or the same vehicle in two places. Fixes: masking reflective surfaces, geometry checks, and per-camera exclusion polygons drawn during calibration.

4. PPE look-alikes and awkward poses. High-vis rain covers over helmets, hoods over hard hats, gloves the same colour as sleeves, a worker bent over a machine with the helmet out of view. Signature: false negatives as often as false positives, which is the dangerous combination. Fixes: multi-angle coverage of critical zones, pose-aware detection windows (check PPE at the start of a task, not mid-crouch), and honest reporting of recall per use case, not just precision.

5. Legitimate exceptions. The maintenance crew authorised inside the fence. The forklift that has to cross the walkway during changeover. The zone whose rules change by shift. Signature: technically correct detections that operators experience as false, because the context says “allowed.” These do the most damage to trust, because the system doesn’t look wrong, it looks stupid. Fixes: rule scheduling, exception workflows with expiry dates, and a weekly review where operators nominate rules for adjustment. A calibration process without an operator feedback loop will win on the precision statistics and lose the room.

The field guide also shows why calibration can’t be fully automated. Families 1–3 yield to technical tuning. Families 4–5 need human judgement about how work is actually done on that site. That judgement is the “discipline” in calibration discipline.

6. The metrics that belong in your pilot contract

Precision language gets slippery in sales conversations, so pin the definitions down in writing before the pilot starts:

  • Precision (per use case, per camera): the share of raised alerts confirmed as real on human review. The headline hygiene metric.
  • Recall (per use case, where measurable): the share of real events in a sampled review window that the system caught. It’s harder to measure, because it means ground-truthing sampled footage, but a vendor should have a method for estimating it during validation. Precision alone can be gamed by detecting nothing.
  • Weighted average accuracy: the blended measure across active use cases, weighted by alert volume or criticality. Useful as a trend line; dangerous as a single quoted number without its per-camera distribution.
  • Alert volume per camera per shift: the operational load. Its trajectory over the first 90 days is the single best predictor of long-term engagement.
  • Time-to-acknowledge and open rate: the human half of the system. If these slip while precision holds, the problem is routing and volume, not detection.
  • False-positive share at calibration exit: the number the calibration window commits to. Written down, measured by counting, validated jointly, and understood by both sides as a target demonstrated during validation, not a permanent contractual guarantee. Sites change; that’s what the weekly report is for.
Six metrics for a pilot contract: precision, recall, weighted-average accuracy, alerts per camera per shift, open rate and time to acknowledge, and false positives at calibration exit.
Figure 4. Six metrics to write into the pilot contract

Two governance notes. First, insist that your own team adjudicates a sample of alerts during validation. Homework marked by the vendor converges on flattering numbers. Second, agree that the measurement rhythm survives go-live: weekly accuracy report, monthly per-camera review, quarterly recalibration check. Deployments keep their week-one discipline exactly as long as the reporting cadence forces them to.

7. The buyer’s verification checklist

Vendor answers compared: those showing a measured calibration discipline versus those that predict week-3 alert fatigue, such as quoting one accuracy number or offering a sensitivity slider.
Figure 5. How to read vendor answers to the checklist

8. A 90-day pilot plan built around precision

Here’s the discipline as a schedule you can adopt as it stands:

Days 1–10: connection and silent baseline. Cameras connected, detections running, nobody alerted. Deliverables: a coverage map, raw alert-stream statistics per camera and rule, and a list of visible confusers (see the field guide in section 5). Agree adjudication rules with the vendor: what counts as real, who decides, and how disputes are settled.

Days 11–25: calibration window. Per-camera tuning against the written false-positive target. A joint weekly review: your team adjudicates a random sample of alerts, and the vendor shows how per-camera precision moved week over week. Exit criterion: the agreed false-positive share, shown on counted data, or a documented plan for the cameras that miss it (reposition, mask, change the rule).

Days 26–40: supervised live operation. Alerts go to named owners under the agreed routing (real-time, digest or dashboard). Measure the human half from day one: open rate and time-to-acknowledge. Hold an operator forum at the midpoint. This is where the legitimate exceptions (family 5) surface, and every one you fix is trust in the bank.

Days 41–80: steady state under measurement. Weekly accuracy reports running. Alert volume flat or falling while detection holds. If you can, introduce one deliberate change (a layout change, a new shift pattern) to test recalibration while the vendor is still courting you.

Days 81–90: decision package. Final metrics against the written success criteria: precision per use case, estimated recall from sampled ground truth, alert-volume trajectory, engagement, and the operator forum’s verdict. The question that decides renewal isn’t “did it detect things?” It will have. It’s “does the team trust the channel more in week 12 than in week 2?” If yes, scale. If no, the pilot told you cheaply what a rollout would have told you expensively.

Two contract notes for the pilot. Put the calibration exit criterion and measurement method in the order form, not in a slide deck. And keep the baseline data whatever the outcome. It’s your site’s risk data, and it means your next evaluation, or your rollout business case, starts from evidence.

9. Frequently asked questions about false positives in safety AI

Short answers to the questions EHS and operations leaders ask most often about false positives, alert fatigue and calibration in AI safety monitoring.

What is a false positive in AI safety monitoring?

A false positive is an alert that an AI safety system raises for something that isn't a real hazard, such as steam flagged as smoke or a reflection flagged as a person inside an exclusion zone. In computer-vision safety monitoring, false positives matter as much as missed hazards. A high false-positive rate trains EHS teams and supervisors to ignore alerts, including the real ones.

Why do AI safety monitoring deployments fail?

Most AI safety deployments fail because of alert fatigue, not because the model misses hazards. When an uncalibrated system floods safety teams with false alarms, operators stop responding, usually within about three weeks of go-live. Other fields document the same pattern: clinical studies find that 74–99% of hospital monitor alarms are non-actionable, and security operations centres report similar overload.

What is alert fatigue in workplace safety?

Alert fatigue is the loss of trust and response that sets in when people receive too many alerts that turn out to be false or irrelevant. In workplace safety AI, it shows up as slower acknowledgement times, falling open rates and informal rules such as “ignore that camera.” Eventually a real near miss gets treated as noise.

Can AI safety systems reach zero false positives?

No. A vendor that claims zero false positives isn't measuring them. The realistic goal is precision high enough that every alert deserves attention. That precision should be measured per camera and per use case and reported weekly, so it holds as lighting, layouts and seasons change.

How accurate should AI safety monitoring be?

Ask for measurement before you ask for a number. A credible vendor reports a weighted-average accuracy across use cases, shows the per-camera distribution, and updates both every week. A continuously measured accuracy of around 95% is auditable. An unmeasured claim of 99% is marketing. Treat any accuracy figure as measured, not guaranteed.

How long should calibration take for AI safety cameras?

Calibration should take weeks, not quarters. A common structure is a short silent baseline, then a calibration window of about two weeks. That window has a written false-positive target, measured per camera and validated jointly with the customer. Be wary of a vendor with no calibration period, and equally of one whose calibration never ends.

How do you reduce false positives in computer-vision safety monitoring?

False positives come down through site calibration, not a new model. The main tools are per-camera zone masks, threshold tuning, rule schedules, multi-frame confirmation for smoke and fire, and alert routing by severity. The most common causes on industrial sites are low sun and weather, steam and dust, reflections, PPE look-alikes, and legitimate exceptions such as authorised maintenance work.

Should a safety AI pilot start with many use cases or a few?

Start with one or two use cases and calibrate them to trustworthy precision before adding more. Early precision builds the operator trust that later use cases inherit. Launching broad coverage on day one tends to reproduce the week-3 alert-fatigue curve at enterprise scale.

If our team ignores VMS motion alerts, will they ignore AI safety alerts too?

Not necessarily, but better detection alone won't fix it. Ignoring noisy motion alerts is rational behaviour, as alarm-fatigue research shows. AI safety alerts hold attention when the alert stream is designed with the team. Every real-time alert has a named owner and a defined response, and everything else goes to a daily digest or a dashboard.

What metrics should an AI safety pilot contract include?

An AI safety pilot contract should define six metrics: precision per use case and per camera, recall estimated from sampled footage, weighted-average accuracy, alert volume per camera per shift, open rate and time-to-acknowledge, and the false-positive share at calibration exit. Put these definitions and the measurement method in the order form, and have your own team adjudicate a sample of alerts.

Does AI safety monitoring identify individual workers?

It doesn't need to, and privacy-first systems don't. Surveily, for example, detects objects, equipment, vehicles and behaviours, such as a missing hard hat, without facial recognition or any other biometric identification. Footage is de-identified by default, with face and body blurring applied before anything is stored or shown. People, not the system, decide what happens after an alert.

How does Surveily handle false positives?

Surveily runs a bounded calibration period for every deployment, validated openly with the customer against a written false-positive target. Dedicated data specialists tune each camera. Customers get a weekly accuracy report with the per-camera view and a weighted average across use cases. That average is around 95% across deployments, measured and reported continuously, never guaranteed.

Sources

  1. Clinical alarm literature (PubMed Central): reviews reporting 74–99% non-actionable physiological monitor alarms and ICU alarm loads above 350 per patient per day. Example: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3752621/
  2. StrangeBee — Cybersecurity alert fatigue (2023 study: ~83% of alerts false; 67% ignored). https://strangebee.com/blog/what-is-cybersecurity-alert-fatigue-and-how-to-fight-back/
  3. SANS Detection and Response Survey 2025 — 73% of security teams cite false positives as top detection challenge (via Vectra AI summary). https://www.vectra.ai/topics/alert-fatigue
  4. AlAhmadi, B. et al. — “99% False Positives: A Qualitative Study of SOC Analysts’ Perspectives on Security Alarms” (USENIX Security).
  5. IntelliSee — 98% of Security Camera Alarms Are False (industry analysis). https://intellisee.com/98-of-security-camera-alarms-are-false-heres-what-thats-actually-costing-you/
  6. Splunk — Preventing Alert Fatigue in Cybersecurity (25–30% of alerts uninvestigated). https://www.splunk.com/en_us/blog/learn/alert-fatigue.html
  7. Industrial safety-AI deployment outcomes (refinery, steel plant, combined heat and power): scope-specific, customer-reported alert-reduction figures, 2024–2026.

Wojtek Tubek is the founder and CEO of Surveily, a Wrocław-based company that turns the CCTV a site already has into a proactive, privacy-first safety system.

How we run calibration at Surveily. Every deployment starts with a bounded calibration period, validated openly with the customer against a written false-positive target. Dedicated data specialists tune each camera, and customers get a weekly accuracy report: a weighted average across use cases (around 95% across our deployments, measured and continuously reported, never guaranteed) plus the per-camera view. Video is analysed on site and de-identified by default before anything is stored or shown, and raw footage isn’t kept. The system flags; your team decides.

Planning a pilot? Take the checklist in section 7 into every vendor meeting, ours included. This article is educational. Deployment outcomes cited are specific to their site, scope and timeframe.

‍

WHy to wait

Achieve complete risk visibility across  
your site’s operations.

Join companies worldwide that trust Surveily AI to elevate workplace safety and empower their teams. Discover how our AI-driven solutions proactively safeguard employees and optimize operational efficiency.
watch video