Viktar Patotski Viktar Patotski · · security  · 15 min read

Threat Modeling Example: STRIDE on a Real AWS Stack

A worked STRIDE threat modeling example on a real AWS stack: the filled-in matrix, 29 findings, and the two categories that came back completely empty.

A worked STRIDE threat modeling example on a real AWS stack: the filled-in matrix, 29 findings, and the two categories that came back completely empty.

Sooner or later an auditor asks your team for a threat model, because you are going for SOC 2 (the US trust-services audit report) or ISO 27001 (the international security-management certification). Search for a STRIDE threat modeling example and the 2026 answer is everywhere: paste your architecture into an LLM (Large Language Model, ChatGPT and its peers) and get a STRIDE table back in ten seconds.

That produces a confident, well-formatted document that falls apart the moment the auditor asks one question: “how do you know that link is unauthenticated?” The LLM guessed. It cannot show its work.

The version that holds up is not much slower. The rule: the tool owns completeness, never posture. Here is that version, worked end to end on a small Spring Boot plus AWS stack, with a real tool finding real risks. Everything is reproducible: the model, the tool output and the LLM experiment are in a public GitHub repo, xp-vit/stride-threagile-example.

What the STRIDE model is

STRIDE is a threat modeling framework that sorts threats into six categories, each the inverse of a security property: Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, and Elevation of privilege. Microsoft engineers Loren Kohnfelder and Praerit Garg created it in 1999.

  • Spoofing breaks authentication. Can someone pretend to be another user or service?
  • Tampering breaks integrity. Can data or code be modified in transit or at rest?
  • Repudiation breaks non-repudiation. Can an actor deny an action with no trace?
  • Information disclosure breaks confidentiality. Can data leak to someone who should not see it?
  • Denial of service breaks availability. Can the system be made unavailable?
  • Elevation of privilege breaks authorization. Can someone gain rights they should not have?

Threat modeling is the structured process of finding what can go wrong in a system before an attacker does, and deciding what to do about it. With STRIDE you apply the six categories to each element of a data flow diagram (DFD): every process, data store, data flow, and external entity, and every trust boundary they cross (the line where one side trusts the other less, such as the internet edge or the wall between your network and a third party). The Threat Modeling Manifesto compresses the whole discipline to four questions: what are we working on, what can go wrong, what are we going to do about it, and did we do a good enough job.

STRIDE is one of several threat modeling frameworks. PASTA is the risk-centric one that starts from business objectives, attack trees model how a specific attacker goal is reached step by step, and DREAD is a scoring scheme (Damage, Reproducibility, Exploitability, Affected users, Discoverability) that Microsoft itself stopped using for being too subjective. I use the STRIDE methodology because it maps one to one onto a data flow diagram, which is what a tool can check.

The threat modeling example: a Spring Boot and AWS stack

The system under test is a fictional field-service SaaS, a typical vertical product: dispatchers book jobs, technicians get notified, customers pay.

Threat modeling diagram (data flow diagram) of the example Spring Boot and AWS stack: browser to a load balancer, into the API, then a booking service, Postgres, Redis, S3, SQS, Stripe and an email provider, across a public-facing zone and a private application network inside one AWS account.

Nothing exotic: an ALB (Application Load Balancer, AWS’s front door for web traffic), a Spring Boot API behind it, an internal booking service, Postgres for the system of record, Redis for sessions, S3 for documents, SQS for events, and two third parties (Stripe for payments, a provider for email). PII (Personally Identifiable Information: names, addresses, phone numbers) lives in Postgres. Session tokens pass through Redis. This is AWS threat modeling at the level it makes sense to do it: the account as the outer trust boundary, each managed service as its own asset.

This is the kind of architecture an LLM will describe plausibly and get subtly wrong. So I did not let it describe the architecture. I wrote the model.

How to run a threat model, before any tool

You can do all of this in a spreadsheet. The tool later automates steps four and five, nothing else.

  1. List what you have. Every process, data store, data flow, and external entity. For each data asset, write down what it is and how bad a leak would be. Ours: customer PII, payment tokens, booking data, session tokens.
  2. Draw the data flow diagram. Boxes for assets, arrows for calls, exactly as they run in production. The diagram above took ten minutes.
  3. Mark the trust boundaries. The lines where one side trusts the other less: the internet edge, the wall between your network and a third party. Threats cluster on the arrows that cross them.
  4. Sweep STRIDE across every element. Six questions per box and per arrow. Most cells come back empty. That is fine and it is informative.
  5. Record the evidence for every answer. For each security property you claim, write down how you know. “TLS on, checked in Terraform” is evidence. “Probably encrypted” is not, and it should be logged as an assumption, not a finding.
  6. Write the register. Each surviving threat with a severity, an owner, and a decision: fix, accept, or transfer.

Step five is the one everybody skips, and it is the one that decides whether the model survives contact with an auditor.

Threat model as code with Threagile

Threagile is an open-source, MIT-licensed tool that treats a threat model as code. You describe the system in one YAML file (a plain-text configuration format): the technical assets, the data assets, the trust boundaries, and the communication links between assets. It runs a library of deterministic risk rules against that model and produces a data flow diagram, a risk report, and machine-readable JSON. It is written in Go, runs from Docker, and the current release is v0.9.1. No AI anywhere in it: the risk logic is code you can read.

Two details in the YAML matter more than the rest.

First, every communication link carries an explicit authentication, authorization, and protocol. Threagile refuses to run if one is missing; delete a single authentication: line and it exits with unknown 'authentication' value and writes nothing. That forces a decision on every edge in the graph: is this call authenticated, or not?

Second: never derive a security attribute from topology. Knowing that the API calls Redis tells you nothing about whether that call is encrypted or authenticated. If you fill those fields in from the diagram, you invent findings that trace to no evidence, and that is how a threat model stops being believed. So the one link I could verify in the example (the API to Postgres, TLS on and credentials confirmed in config; TLS is Transport Layer Security, the encryption behind https) is marked as verified. The two I could not (the cache and the email provider) get the weaker posture on purpose: assume no encryption, assume no auth, and mark them as assumptions to resolve. Over-report, never under-report. An over-reported risk is visible and gets triaged. An under-reported one is silent.

The full model, the report it produced and all 29 findings are in a public repo: xp-vit/stride-threagile-example. Clone it and running it is one command:

docker run --rm -v "$(pwd)":/app/work threagile/threagile \
  -model /app/work/threagile-model.yaml -output /app/work/out

One trap cost me twenty minutes: keep the model’s title: at 31 characters or fewer. Threagile names the Excel sheet after it, Excel rejects longer sheet names, and the run dies before the PDF is written. (The fontconfig warnings it prints about a cache directory are harmless.)

What the STRIDE threat model found

Threagile flagged 29 risks on this small stack, on its five-level scale from low to critical.

Threagile STRIDE findings: bar chart of the 29 risks by severity, 9 elevated, 17 medium, 3 low, and 0 critical or high.

Nine are elevated: five server-side request forgery risks (a service tricked into making calls to internal addresses on an attacker’s behalf), one for each outbound call the two services make, to the booking service, S3, Stripe, SQS, and the email provider; two injection risks on the database and cache links; one missing-authentication risk (the unverified Redis link, exactly as designed); and one unguarded-access-from-the-internet risk on the hop from the load balancer into the API. Seventeen are medium, mostly hardening: containers and datastores without hardening, backdoorable base images, no secrets vault, cloud-hardening gaps, no second factor, unencrypted data at rest, and the unencrypted cache link. Three are low: no WAF (Web Application Firewall, the filter in front of the load balancer), flat network segmentation, and one internal call that drops the end user’s identity along the way.

Here is the STRIDE view of the ones that matter (STRIDE category as Threagile’s own rules classify each finding), with the column an auditor cares about most:

ElementSTRIDEFindingSeverityEvidence
API to Redis (cache link)Elevation of privilegeNo authentication on the linkElevatedAssumed (unverified)
API to RedisTamperingInjection via the cache protocolElevatedAssumed
API to RedisInformation disclosureUnencrypted linkMediumAssumed
API to PostgresTamperingSQL injection on the database linkElevatedLink verified; injection is about the code, not the link
Five outbound calls (both services)Information disclosureServer-side request forgeryElevatedStructural
Load balancer to APIElevation of privilegeUnguarded access from the internetElevatedStructural
Containers and datastoresTamperingMissing hardeningMediumStructural
Public request pathTamperingNo WAFLowStructural

The verified Postgres link produced no unencrypted-communication finding. The unverified Redis link produced three findings across three STRIDE categories. That contrast is the point: you can see which findings rest on evidence and which rest on an assumption.

The STRIDE matrix, filled in

This is the artifact a threat model is supposed to produce: six categories against one element of each kind. Categories are assigned by Threagile’s own rules, which is why the missing WAF lands under tampering rather than where you might file it by hand. Empty cells are results, not gaps in the work.

Element (type)SpoofingTamperingRepudiationInformation disclosureDenial of serviceElevation of privilege
Customer browser (external entity)nonenonenonenonenonenone
API service (process)no identity storeinjection, missing hardening, backdoorable base image, no build pipeline, no WAFnoneno secrets vault, unencrypted asset, SSRF on 3 outbound calls, unencrypted linknoneopen from the internet, no second factor
Postgres (data store)nonemissing hardeningnonenonenoneflat network segmentation
API to Redis (data flow)noneinjection over the cache protocolnoneunencrypted linknoneno authentication

Two whole columns came back empty on every element. Repudiation: zero findings. Denial of service: zero findings. Not because the system handles them well. Because a rules engine reading a topology file cannot see them. Whether anyone can prove who deleted a booking depends on audit logging that does not appear in the model at all, and whether the API falls over depends on load behavior no YAML describes.

So the empty cells are the most useful part of the table. They are a precise list of what you still owe the system after the tool has run.

What it did not find, and never will

Every one of those 29 findings is mechanical: missing encryption, an unauthenticated link, a container without hardening. Real, worth fixing, and all spottable by a rule that reads structure.

The abuse cases I wrote into the model include payment fraud (manipulating the charge or refund flow) and a customer reading another customer’s data. Threagile found neither, because no structural rule can. Whether one tenant can reach another tenant’s records is a question about your authorization code, not your architecture diagram. Whether the refund path can be gamed is a question about your domain.

This is the case for automated threat modeling: it clears the boring 80 percent fast and consistently, so the threat-modeling session you run with actual engineers spends its hour on the 20 percent a rules engine will never reach. The session still happens. It just starts at the hard part.

What to do about the nine elevated findings

A threat model that stops at a list is half a threat model. Each fix below is the standard countermeasure for that finding class, not advice specific to your system.

FindingFix
No authentication on the cache linkTurn on ElastiCache in-transit encryption and an AUTH token, then put the credential in a secrets manager rather than config
Injection over the cache protocolNever build cache keys or commands from raw user input; use the client library’s parameterized calls
SQL injection on the database linkParameterized statements only; the ORM does this unless someone hand-builds a query string
Server-side request forgery on five outbound callsAllow-list the destinations each service may reach, block link-local metadata addresses, and resolve hostnames before connecting
Open from the internet at the load balancerTerminate TLS at the ALB, put a WAF in front, and let only the ALB security group reach the API
Missing hardening on containers and datastoresA minimal base image, a non-root user, read-only filesystem, and a rebuild cadence that ships patches
No secrets vaultMove credentials out of environment config into Secrets Manager or Parameter Store with rotation
Backdoorable base imagePin digests, scan images in CI, and fail the build on a known-vulnerable layer
No second factor for privileged accessRequire a second factor on anything that reaches production data

Two of these are worth doing this week regardless of your architecture: the unauthenticated cache link, because a session store readable by anything on the network is an account-takeover path, and the secrets vault, because it is the difference between one leaked config file and one leaked credential.

For the categories the tool left empty, the fix is not a control but a conversation: decide what actions must be provable (that is your audit log), and what happens when a dependency is slow rather than down (that is your timeout and rate-limit story).

Where AI threat modeling helps, measured on the same stack

If Threagile is deterministic, where does an LLM fit? I tested it on this exact stack instead of guessing. I gave gpt-5.5 (the 2026-04-23 snapshot) the same architecture as topology-only prose, no security attributes anywhere, three ways. The raw prompts and unedited outputs are in the repo’s llm-comparison/ folder so you can re-score them.

PromptWhat came backThe problem
”Write the Threagile YAML for this system”A model that failed validation on the first run, and a security posture asserted for every link it could not know: the cache link came back encrypted and credential-authenticated, the email link token-authenticatedThreagile then found 37 risks, more than my 29, but the two posture findings on the cache link (missing authentication, unencrypted) were gone. The guess deleted the findings that mattered and added noise
”Create a STRIDE threat model” (the naive prompt)165 threats, 54 of them rated Critical, 14 about components that are not in the descriptionNothing anchored to a fact about this system. Every row is a maybe, so the auditor’s question has no answer
Same, plus “label each row KNOWN, ASSUMED, or LOGIC”100 threats: 43 known, 42 honestly labeled as assumptions, 15 business-logicThis is the useful one. The 15 LOGIC rows are cross-tenant access, payment-token binding, forged events: exactly what Threagile cannot see

The split is a division of labor. The rules engine owns the evidenced structural findings. The LLM, asked carefully, is the fastest route to a list of business-logic candidates for the human session, and it will label its own assumptions if you tell it to. What it must never do is write the security attributes into the model. Given the chance, it filled every unknown with the secure value, and a wrong “secure” is silent.

Does this satisfy SOC 2 or ISO 27001?

If an audit is why you are reading this, be precise about what the frameworks say. Neither SOC 2 nor ISO 27001 mandates threat modeling by name. The phrase appears zero times in the Trust Services Criteria published by the AICPA (American Institute of CPAs, the body that owns SOC 2). What they mandate is a risk-assessment process:

  • SOC 2 puts risk assessment in the CC3 criteria: identify and analyze risks (CC3.2), consider fraud (CC3.3), reassess on change (CC3.4). A documented threat model is evidence for CC3.2.
  • ISO 27001:2022 mandates a risk assessment process in clause 6.1.2, treatment and a Statement of Applicability (the list of which controls you apply and why) in 6.1.3, and reassessment in clause 8.2. The Annex A controls a threat model supports directly are A.8.27 (secure system architecture and engineering principles) and A.8.25 (secure development life cycle).
  • NIST (National Institute of Standards and Technology, the US standards body) is the one that names it. Task PW.1.1 of its Secure Software Development Framework (SP 800-218) calls for forms of risk modeling, and lists threat modeling, attack modeling, and attack surface mapping as the examples.

So a threat model is one input to the risk assessment the frameworks do mandate. It counts as evidence only with human sign-off and a record of each decision.

Other threat modeling tools I will test the same way

Nobody writing about threat modeling tools seems to run them, so that is the follow-up: the same stack through STRIDE GPT, OWASP pytm, OWASP Threat Dragon and the Microsoft Threat Modeling Tool, and what each one really finds.

Where to start

Everything from this post is free on GitHub: the Threagile model, the report and all 29 findings, and the raw LLM experiment, in xp-vit/stride-threagile-example. Swap in your own assets and links, run it on your own stack first, then spend the engineering hour on what it could not see.


Threat modeling is one of the technical controls behind a clean audit, and I build and evidence those as the engineer, not the auditor. That is Security & Compliance work. Related on this site: AI Wrote Your Code. It Wrote the Vulnerabilities Too., Using AI for Secure Code Review, and SOC 2 for SaaS Startups: What It Costs.

Back to Blog

Related Posts

View All Posts »