Observability MCP servers
Incident response is a search problem, and agents are fast searchers. These servers connect assistants to logs, metrics, error trackers, and traces, so 'why is prod slow' becomes a query rather than a scavenger hunt.
What engineers use them for
- Pull the stack traces and release health behind an alert
- Correlate a spike in errors with the deploy that caused it
- Summarize what changed in the last hour across services
- Draft the incident timeline from real telemetry
Highest-graded right now
Graded on maintenance, completeness, and reliability. The formula is public →
Grafana
An MCP server giving access to Grafana dashboards, data and more.
Superlog
Open-source agent that observes and fixes your application.

Prometheus MCP Server
MCP server providing Prometheus metrics access and PromQL query execution for AI assistants

HomeLab Monitor
Self-hosted homelab dashboard with a built-in read-only MCP server (hosts, Docker, GPU, services).
Auth0 MCP Server
Auth0 MCP Server: Manage Auth0 applications, APIs, actions, logs, and forms using natural language

Prometheus MCP Server
An API-complete MCP server to manage Prometheus-compatible backends.
Fly Io
MCP server for managing Fly.io machines, apps, logs, and Docker images.