Engineering role
AI SRE
An AI site reliability engineer that reads your errors and dashboards and tells you what changed, and why.
Point it at Sentry, Grafana and your repositories and it does the reading that eats an on-call shift: which errors are new since the last deploy, what the dashboards were doing when they started, and which commit is the likeliest cause. It thinks in error budgets rather than alarms, so it tells you whether something is worth waking up for — and it only reads; it never silences an alert or touches production.
Things you could ask for
Written the way you would actually say them. The agent plans the steps itself.
- “Error rate on checkout doubled at 14:10. Pull the new Sentry issues from the last hour, check the Grafana latency panels for the same window, and tell me what changed.”
- “Every morning at 8, summarise yesterday’s new and regressed Sentry issues, grouped by the deploy that introduced them.”
- “Read the last 30 days of our availability metrics and tell me how much of this month’s error budget we have spent, and on what.”
- “Draft a blameless post-incident review for yesterday’s outage from the Sentry timeline, the dashboards and the Linear incident ticket.”
- “List the operational chores our team did by hand more than twice this month, from the Linear history, and which of them are worth automating first.”
What you get back
A short, time-stamped account: what changed, when, which errors and metrics moved with it, and the most likely cause with the evidence for it — plus what would confirm or rule it out. For recurring runs, a daily or weekly digest delivered as a notification. Post-incident reviews come back as a draft document for the team to edit.
Tools it works with
A role is a job description; connectors are what let it do the job on your real work instead of on whatever you paste into a chat.
Where it is the wrong tool
- It reads; it does not act. The Grafana and Sentry connectors cannot silence alerts, acknowledge incidents, roll back a deploy or change infrastructure — during an incident the agent assembles the picture and a person makes the call.
- A cause it offers is a hypothesis. Correlation between a deploy and an error spike is a lead, not a verdict; ask it for the evidence that would confirm or kill each explanation.
- It is not your paging system. Alerting has to be immediate and reliable; an agent that summarises is the layer on top, not a replacement.
- It only sees what the connectors reach. Logs or metrics that live in a tool without a connector are invisible to it — it will say what it could not check, but it cannot fill the gap.
Starting an agent with this role
- 1Create a Burrak account — any plan works, and your first week is $1.
- 2In the Marketplace, open the Roles tab, find SRE and choose “Start an agent as SRE”. That creates a remote agent with the role’s instructions in its system prompt.
- 3Connect the tools it needs — Sentry, Grafana, PostHog, GitHub, Linear — from Connectors.
- 4Give it a brief. It keeps working in the cloud after you close the tab, and you are only charged while it is actually working.
The SRE role, in practice
- Can I use it during an incident, or only afterwards?
- During, for the reading: pulling the new errors, lining them up against the dashboards and the last few merges takes it minutes, which is often the slowest part of the first half hour. Keep decisions — rollback, failover, who to page — with the person on call.
- Will it tell me when something breaks?
- Not in real time — that is your alerting’s job. What it does well is the scheduled look: a morning digest of what regressed overnight, or a weekly read of your error budget, delivered as a notification you can open on your phone.
- Does it need write access to anything?
- No. Everything in this role is reading, and that is deliberate. If you want it to open a ticket for what it finds, connect Linear and ask it to — that is the only kind of write it needs.
Capability reference
Summarised from the role’s instructions, which are adapted from the agency-agents collection (AgentLand Contributors), used under the MIT licence.
- Defines SLOs that reflect user experience, and acts on error budgets
- Builds observability that answers “why is this broken?” in minutes
- Automates repetitive operational work instead of doing it by hand
- Plans capacity from data, not guesses
- Blameless by default: systems fail, not people
Your AI assistant is waiting.
Put it to work today.
Put an agent on the work that never needed a human in the first place — the research, the reports, the follow-ups — and take the week back.
$1 for your first week · Cancel anytime · Works on Mac, Windows, iPhone, Android & Web




