Latency increased in US East
criticalP95 latency increased from 410 ms to 930 ms over five minutes.
Platform on-call2 min ago
Operational cockpit for monitoring autonomous agent infrastructure health, service metrics, and incident history.
Production environment
Uptime 99.97% · 3 active signals
P95 latency increased from 410 ms to 930 ms over five minutes.
Evaluation accuracy decreased by 4.8 points against the reference set.
Nightly backup completed with no errors.
Production environment
Uptime 99.92% · 3 active signals
Replica lag is 28s (> 10s threshold). Agents throttled to 60 req/min
Tool execution route returning 2.6% 500 errors for tier-1 customers
Nightly snapshot completed. Verifying restore checksums
Staging environment
Uptime 99.5% · 2 active signals
Refresh pending security approval. Tokens expire in 4h
Preview deployment succeeded · build #812
npx shadcn@latest add "https://agents-ui.github.io/agents-kit/c/agent-ops-monitor.json"| Prop | Type | Default | Description |
|---|---|---|---|
| environment | string | "Production" | Environment name in header |
| uptime | string | "99.97%" | Uptime badge |
| signals | OpsSignal[] | Sample signals | Live alerts feed |
| metrics | OpsServiceMetric[] | Sample metrics | Service metrics cards |
| incidents | IncidentEvent[] | Sample incidents | Recent incident timeline |
| onAcknowledge | (signal) => void | - | Called when acknowledging a signal |
| onEscalate | (signal) => void | - | Called when escalating |
| onExportReport | () => void | - | Fired by export button |
| className | string | - | Tailwind overrides |
interface OpsSignal {
id: string
title: string
status: "healthy" | "warning" | "critical"
detail: string
owner?: string
lastUpdated?: string
}
interface OpsServiceMetric {
label: string
value: string
threshold: string
trend: "up" | "down" | "stable"
}
interface IncidentEvent {
id: string
timestamp: string
summary: string
actionNeeded?: string
}
ScrollArea for consistent height in docsh-7 sizing for secondary actions