Engineering role
AI Engineer
An AI engineer agent that designs the retrieval, prompts and evaluations behind your AI feature — and tells you how you will know it works.
Give it your repository, your data and the AI feature you want — search over your docs, a support assistant, a classifier — and it does the engineering around the model: the retrieval design, the schema for embeddings, the prompts and tool definitions, and the evaluation set that says whether a change made things better. It starts from the question “how will we measure this?” rather than from a model name.
Things you could ask for
Written the way you would actually say them. The agent plans the steps itself.
- “Design retrieval over our help-centre articles: chunking, metadata, the pgvector schema and the query, and write the migration so I can try it on a Neon branch.”
- “Write an evaluation set of 50 real questions from our support history, each with the answer we would accept, into a Google Sheet I can score against.”
- “Read our prompts and tool definitions in the repository and tell me where the model is most likely to call the wrong tool or invent an argument.”
- “Find the papers on arXiv from the last year about evaluating retrieval quality, and summarise which methods we could run with our data and team size.”
- “Our assistant answers wrongly for these ten questions. Read the code path and the retrieved chunks and tell me whether retrieval or the prompt is at fault.”
What you get back
Design documents and code you review in a pull request: retrieval and data designs, schema and migrations as SQL, prompts and tool definitions as files, evaluation sets as rows in a sheet with a scoring rule, and paper summaries with links. Every recommendation names the metric that would show it worked.
Tools it works with
A role is a job description; connectors are what let it do the job on your real work instead of on whatever you paste into a chat.
Where it is the wrong tool
- It does not train or fine-tune models. There are no GPUs and no training environment behind it; it can prepare the data and the configuration, and the run happens on your infrastructure.
- A remote agent cannot run your code or your evaluation suite — it has no shell in the cloud. It writes the harness and reasons about results you give it; a local agent with shell access in the desktop app can run it on your machine.
- Model prices, context limits and benchmark scores change monthly. Anything it states from memory may be out of date — ask it to check current figures with web research before you choose on them.
- There is no Hugging Face, Weights & Biases or MLflow connector today, so it sees experiments and model cards only as files or exports you share.
Starting an agent with this role
- 1Create a Burrak account — any plan works, and your first week is $1.
- 2In the Marketplace, open the Roles tab, find AI Engineer and choose “Start an agent as AI Engineer”. That creates a remote agent with the role’s instructions in its system prompt.
- 3Connect the tools it needs — GitHub, GitLab, Neon, Supabase, arXiv, Tavily Web Search, Google Sheets, Sentry — from Connectors.
- 4Give it a brief. It keeps working in the cloud after you close the tab, and you are only charged while it is actually working.
The AI Engineer role, in practice
- Can it build a RAG system for us?
- It designs one and writes most of the code: how documents are split, what metadata to keep, the vector schema in Postgres, the query and the prompt that uses the results. With a Neon branch or a development Supabase database connected it can create the tables and test queries there. Loading your full corpus and serving it in production stays with your pipeline.
- How is this different from the Data Engineer?
- The Data Engineer builds the pipelines that move and model data. The AI Engineer works on what a model does with it: retrieval, prompts, tool calls and the evaluation that says whether an answer is right. Many AI features need both, in that order.
- Does it work with OpenAI, Anthropic or open models?
- It is not tied to one provider. It designs against whichever API or local model you use, and it is most useful at keeping that choice cheap to change: one interface in your code, and an evaluation set you can rerun when you try a different model.
Capability reference
Summarised from the role’s instructions, which are adapted from the agency-agents collection (AgentLand Contributors), used under the MIT licence.
- Large language models: retrieval-augmented generation, prompt engineering and fine-tuning plans
- Vector search with pgvector, Pinecone, Weaviate, Qdrant or FAISS
- Real-time, batch, streaming and on-device inference patterns
- Model versioning, A/B comparison, monitoring and retraining (MLOps)
- Bias testing, privacy-preserving data handling and human oversight
Your AI assistant is waiting.
Put it to work today.
Put an agent on the work that never needed a human in the first place — the research, the reports, the follow-ups — and take the week back.
$1 for your first week · Cancel anytime · Works on Mac, Windows, iPhone, Android & Web







