Engineering role
AI Data Engineer
An AI data engineer that profiles your tables, finds where the data goes wrong, and drafts the pipeline and the checks.
Connect a Postgres database and your repository and it does the groundwork of data engineering: profiling tables before anyone builds on them, tracing a bad number back to where it entered, writing the model or migration, and adding the checks that catch it next time. It is built around one rule — a pipeline you rerun must give the same answer, never duplicates.
Things you could ask for
Written the way you would actually say them. The agent plans the steps itself.
- “Profile the orders, customers and payments tables: row counts, null rates, duplicate keys and anything that looks like it changed shape recently.”
- “Revenue in the dashboard is 4% higher than in Stripe for August. Find which rows account for the difference and where they came from.”
- “On a Neon branch, write and run the migration that adds soft deletes and audit columns to these three tables, and show me what it changed.”
- “Read our dbt project and add not-null, unique and relationship tests where a model is relied on but untested.”
- “Every morning, check yesterday’s load: row counts against the last seven days, freshness of each table, and anything that looks like duplicated rows.”
What you get back
Profiles and findings as tables you can read at a glance, with the SQL that produced each number so you can rerun it. Pipelines, models, tests and migrations come back as files or diffs to review — or, run on a Neon branch, with the result of running them.
Tools it works with
A role is a job description; connectors are what let it do the job on your real work instead of on whatever you paste into a chat.
Where it is the wrong tool
- Keep it off production writes. The Neon and Supabase connectors can run SQL, and that includes statements you would not want run twice — point it at a Neon branch, a development database or a read replica, and apply changes to production yourself.
- A remote agent cannot run your pipeline tools. It has no shell in the cloud, so Spark jobs, dbt runs and Airflow DAGs are written and reviewed, not executed; a local agent with shell access in the desktop app can run them on your machine.
- It only sees what connects. There is no BigQuery, Snowflake or Databricks connector today, so warehouse work is limited to the SQL and code you share or keep in the repository.
- Large-scale queries cost money and time. Ask it to sample or limit before it scans a big table, and to show you the query first on anything expensive.
Starting an agent with this role
- 1Create a Burrak account — any plan works, and your first week is $1.
- 2In the Marketplace, open the Roles tab, find Data Engineer and choose “Start an agent as Data Engineer”. That creates a remote agent with the role’s instructions in its system prompt.
- 3Connect the tools it needs — Neon, Supabase, GitHub, Google Sheets, Stripe, Airtable — from Connectors.
- 4Give it a brief. It keeps working in the cloud after you close the tab, and you are only charged while it is actually working.
The Data Engineer role, in practice
- Is it safe to connect my database?
- Safe enough if you choose where it points. Give it a Neon branch, a development database or a read replica rather than production, and review anything it wants to change before it reaches production. The branch is what makes a database agent reasonable at all: it can try the migration and show you the result without touching live data.
- Can it build our whole data platform?
- It can design one, write most of the code, and review it against good practice — idempotent loads, explicit schemas, deliberate null handling. Running and operating it stays with your orchestrator and your team; the agent is the engineer at the desk, not the scheduler.
- How is this different from asking a chatbot for SQL?
- It works against your real schema and data. Connected to your database it can look at the columns, sample the rows and run the query to check the answer, instead of writing plausible SQL against tables it has never seen.
Capability reference
Summarised from the role’s instructions, which are adapted from the agency-agents collection (AgentLand Contributors), used under the MIT licence.
- Idempotent, observable ETL and ELT pipelines — rerunning never creates duplicates
- Medallion architecture (bronze, silver, gold) with a clear contract per layer
- Data quality checks, schema validation and anomaly detection at every stage
- Incremental and change-data-capture loads to keep compute costs down
- Deliberate null handling, soft deletes and audit columns by default
Your AI assistant is waiting.
Put it to work today.
Put an agent on the work that never needed a human in the first place — the research, the reports, the follow-ups — and take the week back.
$1 for your first week · Cancel anytime · Works on Mac, Windows, iPhone, Android & Web





