Agent Ops Monitor

Operational cockpit for monitoring autonomous agent infrastructure health, service metrics, and incident history.

Features

  • Signal Feed: Real-time alerts with escalation + acknowledgement controls
  • Service Metrics: Keep thresholds and trends visible next to active incidents
  • Incident Timeline: Show latest actions for quick war-room context
  • Environment Aware: Swap context (production vs staging) via props

Examples

Production Overview

Production environment

Operations status

Uptime 99.97% · 3 active signals

Latency increased in US East

critical

P95 latency increased from 410 ms to 930 ms over five minutes.

Platform on-call2 min ago

Embedding quality drift

warning

Evaluation accuracy decreased by 4.8 points against the reference set.

Search team8 min ago

Backup pipeline

healthy

Nightly backup completed with no errors.

InfrastructureCompleted
Service metrics and incidents
P95 latency930 msTarget Below 550 ms
Error rate0.8%Target Below 1%
Availability99.7%Target Above 99.9%
Failover to EU West orchestratorMonitor
Customer impact flagged for priority accountsSupport notified

Incident Response

Production environment

Operations status

Uptime 99.92% · 3 active signals

Vector db replica lag

critical

Replica lag is 28s (> 10s threshold). Agents throttled to 60 req/min

Agent Atlas1m ago

API 500s

warning

Tool execution route returning 2.6% 500 errors for tier-1 customers

Agent Pulse3m ago

Backup pipeline

healthy

Nightly snapshot completed. Verifying restore checksums

Agent NovaCompleted
Service metrics and incidents
P95 latency780msTarget < 550ms
Error rate1.2%Target < 1%
Agent availability99.3%Target > 99.9%
Triggered read-only modeMonitor lag
Routed backups to secondary regionConfirm ETL status

Staging Sandbox

Staging environment

Operations status

Uptime 99.5% · 2 active signals

Staging data refresh

warning

Refresh pending security approval. Tokens expire in 4h

Agent RelayPending

Preview environment

healthy

Preview deployment succeeded · build #812

Agent ForgeNow
Service metrics and incidents
Deploy success97.1%Target > 95%
Preview latency420msTarget < 600ms
QA coverage84%Target > 80%

Installation

npx shadcn@latest add "https://agents-ui.github.io/agents-kit/c/agent-ops-monitor.json"

API Reference

AgentOpsMonitor

PropTypeDefaultDescription
environmentstring"Production"Environment name in header
uptimestring"99.97%"Uptime badge
signalsOpsSignal[]Sample signalsLive alerts feed
metricsOpsServiceMetric[]Sample metricsService metrics cards
incidentsIncidentEvent[]Sample incidentsRecent incident timeline
onAcknowledge(signal) => void-Called when acknowledging a signal
onEscalate(signal) => void-Called when escalating
onExportReport() => void-Fired by export button
classNamestring-Tailwind overrides

Types

interface OpsSignal {
  id: string
  title: string
  status: "healthy" | "warning" | "critical"
  detail: string
  owner?: string
  lastUpdated?: string
}

interface OpsServiceMetric {
  label: string
  value: string
  threshold: string
  trend: "up" | "down" | "stable"
}

interface IncidentEvent {
  id: string
  timestamp: string
  summary: string
  actionNeeded?: string
}

Design Notes

  • Signal list wraps in ScrollArea for consistent height in docs
  • Buttons follow h-7 sizing for secondary actions
  • Status badges reuse blue/emerald/amber/red palette established across the kit