Cloud & devices

AI & LLM security testing

A test of the AI features you are shipping — assistants, agents, retrieval and anything that hands model output to another system — focused on prompt injection, data leakage and what your tools will do when the model is persuaded.

8
areas covered
5
stages, scoping to retest
In-house
testers, never subcontracted
portal.cyberlysecure.com/acme-health/coverage
Acme Health — Test coverage4 practices · 15 services · one team
In-house
Web app
API
SaaS
Mobile
Client-side
External
Internal
Segment test
Wireless
Cloud config
Hardware & IoT
AI & LLM
Social eng.
Physical
Red team
ApplicationsNetworkCloud & devicesPeople & premises
6 in scope · one team · one report
In-house testers 0 subcontracted
Every surface one team
What this is

What the test covers.

An AI feature turns text into instructions. If untrusted text reaches the model — from a document, a web page, a ticket or another user — that text is trying to give orders, and the model has your tools and your data behind it.

We test the whole feature, not just the model: what reaches the context, what the model is allowed to do, and what happens downstream when it returns something hostile.

OWASP Top 10 for LLM Applications NIST AI Risk Management Framework MITRE ATLAS

What we test

  • Direct prompt injection and jailbreak resistance against your guardrails
  • Indirect prompt injection through documents, pages, emails and other users’ content
  • System prompt and context leakage, including other users’ data in shared context
  • Excessive agency: which tools, functions and actions the model can trigger on its own
  • Retrieval-augmented generation — access control on retrieved documents and source poisoning
  • Insecure output handling where model output reaches a browser, a shell, a query or an API
  • Cost and availability abuse, including denial-of-wallet
  • Model, plugin and dependency supply chain
How it runs

From scoping call to retest, here’s what happens.

Typically one to two weeks, depending on the number of tools and data sources.

1

Scope and authorize

We map the feature: models, tools, data sources and everywhere untrusted text can enter the context.

2

Attack the prompt path

Direct and indirect injection is tested against your real guardrails, not a generic benchmark.

3

Test the agency

We establish exactly what the model can cause to happen, and whether a persuaded model can cause it.

4

Follow the output

Wherever output is consumed by another system, we test what hostile output does there.

5

Report, readout and retest

Findings are reproducible with the exact inputs used, followed by a readout and a retest.

What it surfaces

The kind of thing this test tends to find.

Real examples of what this engagement uncovers — anonymized, and never every time. What matters is that you find out before somebody else does.

Instructions hidden in an uploaded document that the assistant obediently follows

Assistants that will summarize data the signed-in user was never allowed to see

Tool calls triggered on the model’s own initiative with no confirmation step

Model output rendered as HTML, executed as a query or passed to a shell

Unbounded generation loops that cost you money on demand

Getting started

What you get, and what we need from you.

What you get

  • Technical report — Every finding, the evidence behind it and clear guidance your engineers can act on.
  • Executive summary — Your risk explained in plain language for leadership and the board.
  • Attestation letter — Signed confirmation of testing to hand your auditors.
  • Readout call — A walkthrough with the people who tested your systems.
  • Retest — Confirmation your fixes worked, documented for whoever needs to see it.

What we need from you

  • Access to the feature in a test environment, with accounts at each permission level
  • A description of the tools and data sources the model can reach
  • Any guardrail or moderation layer you want tested as deployed
  • Confirmation of usage limits so testing does not break your budget

Missing something on this list? Bring it to the call — we scope around what you have.

Free attack-surface snapshot

Give us a domain. See what an attacker sees.

Not sure where to start? One of our testers reviews your internet-facing footprint and sends you a short summary of what an attacker would see — free. Nothing you don’t own is ever touched, and there’s no sales sequence.

  • Internet-facing hosts
  • Exposed services
  • Leaked credentials
  • TLS certificate hygiene

Request your snapshot

Free. No obligation.

We only ever test assets you own, with your written authorization.