Sentinel Health RMM: AI Agents That Resolve Incidents

An RMM That Closes Tickets Before You Open Them

A 5-agent AI pipeline wrapped in a dark-mode interface built for continuous monitoring

"Monitoring incidents in the current dashboard is like staring at the Matrix." This is what an RMM specialist said during interviews. The diagnosis was clear: alert fatigue, raw tables, and an excessive density of undifferentiated alerts caused cognitive overload and delayed responses to critical incidents.

My decision was to attack the problem on two fronts at once. First, an agentic pipeline that monitors, diagnoses, debates, and remediates recurring incidents without human intervention. Second, an interface that makes machine state instantly legible, so the specialist only enters the loop for edge cases. Both were specified and built through an AI-orchestrated workflow that I directed.

90%
Incidents Self-Resolved
5
Agents in the Pipeline
500+
Assets Monitored
<6s
Interface Response Time

The Problem: Alert Fatigue, Not Ticket Volume

Why NOC specialists were drowning in data, not in work

What Interviews and Benchmarking Showed

  • Repetitive incidents: most tickets were recurring problems with known resolutions, yet every one still required a human to read, diagnose, and close it.
  • No prioritization signal: market tools present dense tables where a critical CPU failure looks the same as a routine warning.
  • Cognitive overload: excessive text density and low alert contrast slowed time-to-insight exactly when speed matters most.
  • Delayed response: the cost is not the ticket, it is the minutes lost finding the one critical row in a wall of data.
Competitor interface in light mode: dense table with many assets and alerts that are hard to read
Market standard: dense tables where everything looks equally important.
Sentinel dashboard in dark mode: status cards, heatmap, and critical asset detail
Sentinel Health: status surfaces the critical first, detail stays one click away.

The Pipeline: Five Agents, One Debate

An incident only gets closed after surviving an internal debate

Monitor Analyst Devil's Advocate Security + SRE Remediation Knowledge

What Each Agent Does

  • Monitor: watches telemetry across the fleet and flags anomalies before they become tickets.
  • Analyst: interprets the anomaly and proposes a diagnosis with a remediation plan.
  • Devil's Advocate: challenges the diagnosis, hunting for false positives and safer alternatives before anything executes.
  • Security + SRE: validates operational safety, permissions, and blast radius.
  • Remediation: executes the approved fix on the asset.
  • Knowledge: writes the outcome back into the knowledge base, so the next incident of the same class resolves faster.
Human-in-the-loop by design. Badges and cards explain exactly what the AI decided and why it decided it. The specialist stays in control of the operation, stepping in for edge cases rather than reading every step of every incident.

The Interface: Built for Continuous Monitoring

A dark surface for long shifts, with triple redundancy on every status

The dashboard runs in dark mode because NOC specialists stare at it for hours. Every status communicates through three channels at once: color, icon, and text. I am a colorblind designer, so accessibility here is not a compliance checkbox, it is a personal constraint I test against real color blindness cases.

Asset card in critical state
Critical: elevated border, badge, and hardware health.
Asset card in warning state
Warning: attention without alarm.
Asset card in healthy state
Healthy: quiet confirmation, zero noise.
Asset card in offline state
Offline: muted treatment, still readable.
Status menu listing four assets: healthy, warning, critical, and offline
Fleet status at a glance: four states, no table scanning required.
Product navigation structure
Navigation organized around the operator's mental model, not the system's.

Powered by the Sentinel Design System

The tokens the interface shows are the same ones the agents use to explain their decisions

Every card, badge, and status color in this product comes from a tokenized design system I specified through an AI-orchestrated workflow: 268 semantic tokens, WCAG 2.2 calibrated, and zero visual drift across 500+ workstations. The design system itself is a separate case study.

Explore the Design System case

Results and What I Would Do Differently

Autonomy with confidence: the specialist focuses only on edge cases

Results Achieved

  • 90% of incidents resolve through the pipeline without human intervention; the specialist reviews outcomes instead of executing every step.
  • The interface answers "does anything need me?" in under 6 seconds, replacing table scanning with status scanning.
  • 500+ simultaneous assets on the same visual surface, stress-tested without degradation.
  • Every AI decision is explainable on screen, which keeps operator trust high and onboarding fast.

What I Would Do Differently Today

  • Density modes: an ultra-compact view for NOC operations and a relaxed view for executive reports.
  • Deeper drill-downs: a direct path from an agent decision to the full debate that produced it, not just the summary.
  • Micro-interactions: intentional motion on status transitions (critical to healthy) to further reduce cognitive load.