Lesson 1 of 5

Start with the job, not the agent

Many first agent projects go wrong because they begin with the technology. Begin with the job instead:

  • What is the task? Write it down in one or two sentences, as you would for a new colleague.
  • What does success look like? A drafted reply, a filled form, a correct classification: something you can check.
  • What is out of scope? The actions the agent must never take, and the cases it must hand to a person.
  • Who is accountable? A named person who owns the agent's results.

Then ask whether you need an agent at all. Sometimes the steps are always the same. Then a fixed workflow is usually cheaper, faster and easier to test: ordinary code that calls a language model at set points. An agent, where the model decides the next step, earns its place when the path really varies from case to case.

Also estimate the running cost early. Each step is usually a separate model call, paid for by the amount of text in and out, so a task of twenty steps costs far more than one of two.

A good first agent has a narrow job, a clear finish line and a person reviewing its output.

Lesson 2 of 5

Designing the tools

An agent can only act through the tools you give it: functions such as "look up an order", "search the policy documents" or "draft an email". The model asks for a tool and supplies the inputs; your code runs it and returns the result. Much of the real engineering work is in these tools.

Good tools are:

  • Narrow. "Get the status of order N" is safer, and easier for a model to use correctly, than "run any database query".
  • Clearly described. The model chooses tools from their names and descriptions. Write them as you would for a new colleague: what the tool does, when to use it, what each input means.
  • Read-only where possible. Keep tools that only look at data apart from tools that change it, and have few of the second kind.
  • Checked in code. Validate every input before acting, and enforce permissions in the tool itself, not in the prompt.
  • Run with the agent's own limited access. Give the agent its own account or key with only the rights its tools need. Never give it a shared administrator login.
  • Helpful when they fail. A clear error such as "order not found; check the number" lets the agent recover instead of guessing.

Open protocols such as the Model Context Protocol (MCP), now run as an open project under the Linux Foundation, let a tool be described once and used by any compatible agent. Treat a tool server written by someone else like any third-party software: check who maintains it and what access it gets before you connect it.

Lesson 3 of 5

The loop, the instructions and the limits

At its core, an agent is a loop. The model reads the goal and what has happened so far, then chooses an action. Your code carries it out, and the result is added to what the model sees next. The loop repeats until the model says it is done.

Around that loop you set:

  • Instructions (often called the system prompt): the role, the goal, the rules to follow, and when to stop and ask a person.
  • Stop conditions: a maximum number of steps, a time limit and a cost limit, so a confused agent cannot run forever.
  • Structured output: ask for the final result in a fixed shape, such as named fields, so your code can check it before using it.
  • Approval points: before any action that is hard to undo, the loop pauses and a person approves.

Keep the rules that really matter outside the model. The instructions can say "never refund more than 5,000 shillings". Only a check in the refund tool's code makes sure it never happens.

Lesson 4 of 5

Memory and context

A language model works only with the text it is given on each step, called its context. Everything the agent "remembers" has to be put there by your code.

  • Working memory is the history of the current task: the goal, the steps taken and their results. Long tasks can outgrow the space available, so older steps are often summarised.
  • Knowledge comes from documents and systems. Rather than pasting everything in, agents usually retrieve the few relevant passages for each step. This technique is called retrieval-augmented generation (RAG).
  • Long-term memory, such as notes kept between sessions, is useful but needs care. Decide what may be stored, for how long, and who can see it. If it holds personal data, Kenya's Data Protection Act, 2019 applies to it.

Two rules help. First, give the model only what the step needs: too much context costs more and can distract it. Second, treat anything retrieved from emails, web pages or uploaded files as untrusted. Hidden text in them can try to redirect the agent (prompt injection), and a model cannot be relied on to ignore it. So limit what the agent can do, and keep approvals for consequential actions.

Lesson 5 of 5

Testing before anyone relies on it

An agent that worked once in a demonstration has been tested once. Before real use:

  • Build a test set of real tasks: easy ones, hard ones, unusual ones, and ones the agent should refuse or hand over.
  • Run each test more than once. The same input can give different results on different runs.
  • Check the path, not only the answer. Did it use the right tools, in a sensible order, without unnecessary steps?
  • Try to break it. Feed it misleading documents, hidden instructions and missing data, and confirm it fails safely.
  • Log everything: every input, tool call, result and approval, so any outcome can be explained later. Logs often contain personal data, so protect them and decide how long to keep them.
  • Start small. Run it alongside people first, with a person approving every action. Widen its freedom only as the evidence builds.
  • Re-test after every change, including a new model version, new tools or edited instructions.

Keep a way to pause or switch the agent off at once, and a named person who watches how it performs.

Knowledge check

Ten questions

Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.

Sources

The official documents this course relies on. Laws and guidance change, so check the current version.

  1. OWASP Top 10 for Agentic Applications for 2026 · OWASP Gen AI Security Project
  2. LLM01: Prompt Injection · OWASP Gen AI Security Project
  3. Guidelines for secure AI system development · UK National Cyber Security Centre, with international partners
  4. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020) · arXiv
  5. Linux Foundation announces the formation of the Agentic AI Foundation (includes the Model Context Protocol) · The Linux Foundation
  6. Data Protection Act, 2019 (No. 24 of 2019) · Kenya Law