Engineering role

AI Incident Response Commander

An AI incident commander that pulls the errors, the dashboards and the recent changes into one timeline — then writes the updates and the blameless post-mortem.

When something breaks, it does the work that falls apart under pressure: reading the new errors in Sentry and the panels in Grafana, lining them up against what merged and deployed, proposing the next hypothesis to check, and drafting the status update for each audience. Afterwards it writes the blameless post-mortem with a timeline, contributing factors and action items that land as tickets — and between incidents it writes the severity levels and runbooks that make the next one shorter.

Things you could ask for

Written the way you would actually say them. The agent plans the steps itself.

  • “Error rates on checkout jumped at 14:20. Pull the new Sentry issues, the Grafana panels for the API, and everything merged since noon, and give me one timeline.”
  • “Draft three status updates for this incident — engineering channel, leadership, and customers — at the level of detail each one needs.”
  • “Write the blameless post-mortem for last night’s outage from this incident thread, with a timeline, contributing factors and five action items as Linear issues.”
  • “Write severity levels SEV1 to SEV4 for our team: criteria, response times, update cadence and who gets paged.”
  • “Turn our three most common incidents from the last quarter into runbooks with checks, fixes and when to escalate.”

What you get back

An incident timeline that cites the error, panel or commit behind each line; hypotheses ranked by evidence; status updates drafted per audience; a post-mortem document; and follow-up actions as issues in Linear or Jira. With Slack connected, it can post an update you approve to the incident channel.

Where it is the wrong tool

  • It does not fix production. The Sentry and Grafana connectors read issues and dashboards, and a remote agent has no shell or cloud credentials, so rollbacks, restarts and config changes stay with your on-call engineer.
  • It does not page anyone. There is no PagerDuty, Opsgenie or Statuspage connector, so alerting and the public status page stay in your existing tools.
  • It sees what is connected. Logs that live only in a cloud console, or a service with no Sentry or Grafana coverage, are invisible to it until someone pastes them in.
  • It works in runs, not in a channel all night. Each run has a time budget, so during a long incident ask it for the next timeline or update rather than expecting it to sit and watch.

Starting an agent with this role

  1. 1Create a Burrak account — any plan works, and your first week is $1.
  2. 2In the Marketplace, open the Roles tab, find Incident Response Commander and choose “Start an agent as Incident Response Commander”. That creates a remote agent with the role’s instructions in its system prompt.
  3. 3Connect the tools it needs — Sentry, Grafana, GitHub, GitLab, Linear, Atlassian (Jira + Confluence) — from Connectors.
  4. 4Give it a brief. It keeps working in the cloud after you close the tab, and you are only charged while it is actually working.

The Incident Response Commander role, in practice

Should an AI agent run our incidents?
It should do the paperwork, not hold the pager. The commander role needs a person who can make the call and take the emergency action. What an agent does well is the part that slips when everyone is busy: the timeline, the evidence, the updates and the write-up.
What does “blameless” mean in practice?
The post-mortem asks what let the failure happen — a missing alert, a test that did not exist, a deploy step that could be skipped — rather than who pressed the button. The agent writes in that frame by default, and will rewrite a draft that names a person as the cause.
How is this different from the SRE role?
The SRE works on reliability between incidents: SLOs, alerts, capacity and toil. The Incident Response Commander takes over when something is on fire and when it is time to learn from it. They share the dashboards and hand each other the action items.

Capability reference

Summarised from the role’s instructions, which are adapted from the agency-agents collection (AgentLand Contributors), used under the MIT licence.

  • Severity frameworks from SEV1 to SEV4 with escalation triggers and update cadence
  • Structured response roles: commander, communications, technical lead and scribe
  • Status updates written for engineers, executives and customers
  • Blameless post-mortems using five whys and contributing-factor analysis
  • Runbooks, on-call rotations and SLO-based paging rules
Your first 7 days for $1

Your AI assistant is waiting.
Put it to work today.

Put an agent on the work that never needed a human in the first place — the research, the reports, the follow-ups — and take the week back.

$1 for your first week · Cancel anytime · Works on Mac, Windows, iPhone, Android & Web