Platform · Observability and reliability

From signal to root cause

Logs, metrics, traces and alerts in one module, with AI investigations that turn a detection event into a written RCA timeline.

Without XamOps

The alert fires, then three engineers open five tools and rebuild the same timeline by hand while the incident is still burning.

[ 01 ] 6 capabilities in detail
01
Beta

Observability

Logs, metrics, traces and alerts in one module with guided agent setup.

Logs, metrics, traces and alerts in one module, with guided setup for the collection agent so instrumentation is not a project in itself. Currently in beta and in use with customers.

  • Logs, metrics, traces and alerts together
  • Guided agent setup
  • Currently beta
02
AWS

Grafana embedding

Bring your existing Grafana dashboards inside XamOps, with Terraform setup help.

Teams that already have good Grafana dashboards should not be asked to rebuild them. Existing dashboards embed directly inside XamOps, with help for the Terraform setup.

  • Bring existing Grafana dashboards in as-is
  • Terraform setup assistance
  • Currently AWS
03
AWSGCPAzure

Performance insights

Utilization and bottleneck analysis per provider.

Utilization and bottleneck analysis per provider, so a slow service can be traced to the resource that is constraining it. This is the layer that turns a cost or latency symptom into a specific thing to change.

  • Utilization analysis per provider
  • Bottleneck identification
  • AWS, GCP and Azure
04
AWSGCPAzure

Alerts

Alarm creation and alert routing per provider.

Alarms are created and alert routing configured per provider from one place, so alerting rules do not have to be maintained separately in each console. Signals land with the people who can act on them.

  • Alarm creation per provider
  • Alert routing to the right owners
  • AWS, GCP and Azure
05
Pre-release

AI SRE investigations

Deep dive

Automated root-cause investigations with an RCA timeline, triggered by detection events or manually.

The expensive part of an incident is the manual reconstruction of what happened. Investigations run automatically from a detection event or on demand, and produce a root-cause analysis with a timeline instead of a raw pile of logs. Currently pre-release.

  • Triggered by detection events or started manually
  • Produces an RCA with an event timeline
  • Currently pre-release
06

AIOps

Anomaly detection and an AI advisor over your own telemetry.

Anomaly detection and an AI advisor that work over your own telemetry rather than a generic model of what a healthy system looks like. The value is catching the deviation that no static threshold was written for.

  • Anomaly detection on your own telemetry
  • AI advisor for what to do next
  • Complements static alert thresholds
[ 02 ] Questions

Observability and reliability, answered.

01
Do we have to replace our existing dashboards?

No. Existing Grafana dashboards embed directly inside XamOps, with help for the Terraform setup, so you can adopt the platform without rebuilding what already works.

02
What does an AI SRE investigation actually produce?

A root-cause analysis with an event timeline, triggered either by a detection event or manually, instead of leaving an engineer to reconstruct the sequence by hand. This module is currently pre-release.

03
How is AIOps different from normal alerting?

Alerts fire on thresholds someone wrote in advance. AIOps runs anomaly detection over your own telemetry and adds an AI advisor, which catches deviations nobody thought to write a rule for.

04
Is the observability module production-ready?

It is currently in beta and in customers' hands. Logs, metrics, traces and alerts are in one module with guided agent setup.

Ready to automate observability and reliability?

30-minute walkthrough. We connect to a sandbox and show this module running against real infrastructure.

Pricing