How attackers use AI
Attackers adopt useful technology quickly, and AI is no exception. In January 2024 the UK National Cyber Security Centre (NCSC) assessed that AI will almost certainly increase the volume and heighten the impact of cyber attacks. It expected this to come mainly from improving existing techniques, not from entirely new ones. Its May 2025 update reached similar conclusions. It noted that the time between a vulnerability being disclosed and being exploited has already shrunk to days, and that AI will almost certainly shorten it further.
Current and expected uses include:
- Phishing and social engineering: fluent, personalised messages in any language, and cloned voices or deepfake video for impersonation.
- Reconnaissance: gathering and summarising information about a target organisation and its people from public sources.
- Speed: helping less-skilled attackers write scripts, understand stolen data or adapt existing tools.
- Finding weaknesses: analysing code and systems for exploitable flaws, a task defenders also use AI for.
- Evasion: attempting to vary malware or messages so they slip past filters that look for known patterns.
The main effect so far is uplift: attacks that were already possible become cheaper, faster and more convincing, and more people can carry them out. That makes the basics of defence, such as patching, strong authentication, backups and trained staff, more important, not less. Expect less time between a flaw being published and being exploited, so patch internet-facing systems quickly. Counter impersonation with process: confirm payment or credential requests through a separate, known channel, and require two people to approve high-value transfers.
How defenders use AI
Security teams face more data and alerts than people can review. AI helps in several ways:
- Detection: models learn normal behaviour for users, devices and networks, and flag departures that may indicate an attack, such as unusual logins, data movement or process activity.
- Triage: grouping related alerts, scoring their severity and enriching them with context, so analysts start with the most important.
- Phishing and malware analysis: classifying suspicious emails, files and links at scale.
- Analyst assistants: language models can summarise incidents, explain unfamiliar code or log entries, draft queries and reports, and search threat intelligence, saving analysts time.
- Secure development: helping find vulnerabilities in code before it is released.
The benefit is time. Used well, AI lets a small team cover more ground and respond faster. It does not replace security fundamentals or the judgement of experienced analysts; it amplifies both.
Attacks on AI systems
As organisations deploy AI, the AI systems themselves become targets. They face ordinary security risks, such as stolen credentials and vulnerable software, plus some of their own:
- Prompt injection: input that tries to make an AI system ignore its rules, leak data or take unintended actions. It can be typed directly by a user, or hidden in content the system reads, such as a web page, email or document (indirect prompt injection). There is no complete fix; the main defence is limiting what the system can do and what harm a manipulated system could cause.
- Data and model poisoning: tampering with training data (including data used for fine-tuning, the extra training that adapts a model to a task), or with the model itself, so it learns wrong or harmful behaviour or carries a hidden backdoor.
- Evasion: crafting inputs that a model misclassifies, for example making malicious files look harmless to an AI detector.
- Model extraction and data leakage: copying a model's behaviour by querying it repeatedly, or getting it to reveal sensitive data it was trained on or has access to.
- Supply chain: third-party models, datasets, plugins and libraries that may be tampered with or vulnerable; check where they come from before use.
- Excessive agency: AI agents given more tools and permissions than they need, so that a manipulated agent can do real damage.
Several public resources map these risks:
- the OWASP Top 10 for LLM Applications, a list of the most serious risks in applications built on large language models (LLMs);
- MITRE ATLAS, a knowledge base of adversary tactics and techniques against AI systems, modelled on MITRE ATT&CK and drawn from real-world attacks and red-team demonstrations;
- NIST's taxonomy of adversarial machine learning attacks and mitigations (NIST AI 100-2).
Protect AI systems like any critical system, and more: least privilege for AI agents (only the tools and access each task needs), human approval for consequential actions, input and output monitoring, testing against these attacks before release, and a clear owner for each system. Log the prompts, outputs and tool actions of AI systems so incidents can be investigated, and include AI systems in the incident response plan.
Where AI helps, and where it misleads
AI is powerful but imperfect, and in security the cost of mistakes is high.
- False positives (alerts about harmless activity) waste analyst time and, if there are too many, teach people to ignore alerts.
- False negatives (real attacks the tool does not flag) are often more costly: a missed attack that everyone assumed the AI would catch.
- Confident errors: language models can give plausible but wrong answers, for example misreading a log or inventing a detail about a threat. In an incident, a wrong answer delivered confidently can send a team in the wrong direction.
- Adversaries adapt: attackers study defensive tools and shape their behaviour to avoid them.
- Opacity: it can be hard to know why a model raised or ignored something, which matters for investigation and accountability.
Good practice:
- Measure AI tools on your own environment, with your own data, before trusting them.
- Keep humans in the loop for decisions that matter, such as isolating systems, blocking users or declaring an incident.
- Ask for evidence: an AI assistant's conclusion should point to the logs, files or events behind it, so an analyst can check.
- Keep the basics strong: AI detection is a layer on top of patching, authentication, logging and backups, not a replacement for them.
Governing AI in a security team
Using AI safely in security needs the same discipline as any other powerful tool.
- Know what you use: an inventory of AI tools, what data they can see and what they can do.
- Protect the data: security data contains passwords, personal data and details of your weaknesses. Decide which AI services may process it, where that data goes, and what the contracts say. In Kenya, the Data Protection Act, 2019 applies to personal data in security logs and investigations as anywhere else. Tell staff which AI tools are approved for security data, and block or monitor the rest; pasting logs, code or customer data into unapproved AI tools is a common real leak.
- Limit permissions: AI assistants that can take actions, such as blocking accounts or changing firewall rules, need tight limits and human approval.
- Test before trusting, including attempts to manipulate the tool, and re-test after changes.
- Keep records of what AI tools recommended and what people decided, for audit and learning.
- Train the team to use AI well: what it is good at, where it fails, and how to check its work.
In Kenya, the National Computer Incident Response Team – Coordination Centre (National KE-CIRT/CC), based at the Communications Authority of Kenya, coordinates the national response to cyber incidents, and the Kenya Artificial Intelligence Strategy 2025–2030 names AI-enabled cyber attacks, such as automated phishing, among the risks it aims to address.
Frameworks help structure this work. The NIST AI Risk Management Framework covers AI risk in general, and the NIST Cybersecurity Framework 2.0 covers the wider security programme that AI tools sit within.
AI will not decide the contest between attackers and defenders on its own. Organisations with strong fundamentals, clear processes and skilled people will use it to get stronger; those without them will find it amplifies their weaknesses.
Ten questions
Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.
Your answers
Your certificate of completion
Enter your name as you want it to appear, then save the certificate as a PDF. In the print window, choose Save as PDF. A certificate is issued once per completion of the course.
Saolix does not record who takes this course, so it cannot verify these certificates. The certificate confirms completion of a free self-paced course and is not an accredited qualification.
Sources
The official documents this course relies on. Laws and guidance change, so check the current version.
- The near-term impact of AI on the cyber threat (2024) · UK National Cyber Security Centre
- Impact of AI on cyber threat from now to 2027 (May 2025) · UK National Cyber Security Centre
- Prompt injection is not SQL injection (it may be worse) · UK National Cyber Security Centre
- AI Risk Management Framework (AI RMF 1.0) · US National Institute of Standards and Technology
- Kenya Artificial Intelligence Strategy 2025–2030 · Ministry of Information, Communications and the Digital Economy, Kenya
- National KE-CIRT/CC · Communications Authority of Kenya
- MITRE ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems) · MITRE
- NIST AI 100-2 E2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations · US National Institute of Standards and Technology
- OWASP Top 10 for LLM Applications 2025 · OWASP Gen AI Security Project (OWASP Foundation)
- The NIST Cybersecurity Framework (CSF) 2.0 · US National Institute of Standards and Technology