A data flow diagram in threat modeling is a drawing of how data moves through one system, made so a team can ask what could go wrong with it. The same three words name an older thing, the data flow diagram of classic software engineering, which models business processes and carries no security meaning. This guide is about the security one: much the same shapes put to a different job, plus one element type that does most of the work, the trust boundary. DFD is the abbreviation, and it means this one.
OWASP's Threat Modeling Cheat Sheet says data flow diagrams "are arguably the most common approach" to modelling a system for threat modeling, and describes what a good one gives you: "a clear view of trust boundaries, data flows, data stores, processes, and the external entities" that interact with the system. Five kinds of thing, and that is the whole notation.
Why a data flow diagram matters
The Threat Modeling Manifesto frames the practice as four questions: "What are we working on? What can go wrong? What are we going to do about it? Did we do a good enough job?" OWASP's cheat sheet repeats them word for word. The diagram is the answer to the first question, and each of the other three needs something the one before it produced.
Question 01
What are we working on?
The drawing: the elements, the flows between them, and the line where the level of trust changes.
Question 02
What can go wrong?
The same questions asked of every element, and every answer names the element it lands on.
Question 03
What are we going to do about it?
Every threat leaves the session with one of three decisions against it, and a name on the decision.
Question 04
Did we do a good enough job?
Which parts of the system were looked at, and which were not, counted against the element list.
Asked
Not yet
Question 01
What are we working on?
The drawing: the elements, the flows between them, and the line where the level of trust changes.
All three start here
Question 02
What can go wrong?
The same questions asked of every element, and every answer names the element it lands on.
Question 03
What are we going to do about it?
Every threat leaves the session with one of three decisions against it, and a name on the decision.
Question 04
Did we do a good enough job?
Which parts of the system were looked at, and which were not, counted against the element list.
Asked
Not yet
The fourth question is answered by counting. Every element on the drawing either had the questions asked of it or did not, and the answer is that list. Without the drawing there is nothing to count, which is why the first question comes first.
Without a diagram the second question has nothing to ask of, so a threat modeling session is a brainstorm, and a brainstorm finds what the loudest person in the room remembers. With one, the unit of work changes. Every threat is anchored to an element, every element gets the same questions, and at the end of the session you can say which parts of the system were looked at and which were not. That last part is the difference between a meeting and a method.
There is an external reason as well. CISA, the NSA, the FBI and fifteen international partners wrote in Shifting the Balance of Cybersecurity Risk (April 2023, revised 25 October 2023) that "secure by design products start with written threat models that describe what the creators are trying to protect and from whom", and that effective models cover the enterprise and development environments as well as the way the software is meant to be used by customers. That is guidance for software makers, not a rule anyone enforces, but it has changed what customers ask for, and it turns a diagram drawn in a room into an artefact other people read.
What earns a model is a change, not a date on a calendar. Four triggers cover most of it: a new system or a materially new service, a new external integration, a change to authentication or authorisation, and a new category of data. Our step-by-step guide to threat modeling sets out the method those triggers start; this page stays on the diagram.
The last reason is what happens afterwards. Number the elements and a threat can name the thing it targets, a control can name the threat it answers, and an assessor's question can be walked backwards from the evidence to the design decision that produced it. An unnumbered picture carries none of that, and gets redrawn from scratch every time somebody asks.
How to draw a data flow diagram, step by step
The notation is small enough to hold in your head, which is the point: anyone in the room can read the drawing, argue with it and correct it. Five element types, and each attracts a different set of questions later.
External entity
A rectangle with sharp corners
Anyone or anything outside your control that touches the system. You cannot change how it behaves, only what you accept from it.
On the label: Who or what it is, in the words your team uses.
EE1 · Customer browser
Process
A circle, or a rounded box in most tools
Code that handles data: it reads something, decides something, writes something. One process per responsibility you would defend separately.
On the label: What the code does, not the server it runs on.
C1 · API gateway, C2 · Order service
Data store
Two parallel lines with the name between them
Anywhere data rests. Databases, object storage, caches, queues, log files and backups all count.
On the label: What is kept there, and the most sensitive thing in it.
DS1 · Orders database
Data flow
An arrow, one per direction
Data in motion between two elements. The arrow points the way the data moves, and the label is where most threats are found.
On the label: What moves, over what protocol, authenticated how.
DF1 · Order request, HTTPS, session cookie
Trust boundary
A dashed line, drawn around or between elements
The line where the level of trust changes. Flows that cross one are the flows an attacker reaches, and the controls belong at the crossing.
On the label: Which two levels of trust meet, and who controls each side.
TB1 · Internet edge, TB2 · Database tier
Tools draw the shapes differently: Microsoft's tutorial uses a square for the outside entity, a circle for the web server and two parallel lines for the database, and most modern canvases draw every process as a rounded box. The shapes are a convention. The five element types are the part that matters, because each one attracts a different set of questions.
Ten steps take a blank page to a diagram that has been checked, questioned and decided on. The first eight produce the drawing; the last two are what the drawing was for.
- 01
Set the scope and the question
Name the one system you are modelling and write down what you are trying to protect. Model that system and its immediate neighbours: everything past the first hop is an external entity with an arrow pointing at it.
Output: One sentence of scope, agreed out loud
- 02
Place the external entities
Anything outside your control that touches the system: customers, partner systems, a payment provider, an identity provider, a scheduled job somebody else owns.
Output: Every actor on the page, none of them yours
- 03
Draw the processes
The code that handles data, one per responsibility you would defend separately. Two components under the same owner, holding the same data at the same level of trust, are one process.
Output: The parts that act on data
- 04
Add the data stores
Anywhere data rests, including the quiet ones: caches, queues, object storage, log files and backups. The quiet stores hold the same data as the loud ones and are guarded less.
Output: Every resting place, named by what it holds
- 05
Draw every flow with what it carries
One arrow per direction, each labelled with what moves, over what protocol, authenticated how. The label is where most of the threats come from, so an unlabelled arrow is an unasked question.
Output: Labelled arrows, direction shown
- 06
Mark the trust boundaries
Draw them last, once the elements are on the page, around groups of things under one level of control. Ask who owns each side and what each side can enforce.
Output: Dashed lines, and a list of what crosses them
- 07
Number every element
EE1, C1, DS1, DF1, TB1. Numbering is what makes the diagram addressable: a threat, a control, a test and an audit finding can all name the same thing without describing it again.
Output: An element list a stranger can use
- 08
Check it against the real system
Read the code, the configuration or the infrastructure definition, and correct the drawing. Two engineers who draw the same system differently have found something before a single threat is written down.
Output: A diagram that matches what runs
- 09
Ask the six STRIDE questions, element by element
Spoofing, tampering, repudiation, information disclosure, denial of service and elevation of privilege, asked of each element and of every flow that crosses a boundary. Phrase each answer as a sentence with an actor, a verb and an asset.
Output: Threats, each anchored to an element
- 10
Decide every threat, and record the decision
Change the design, add a control, or accept it with a named owner and a date. A threat with no decision against it is the most expensive kind, because the work looks done.
Output: A decision per threat, with a name on it
1. Set the scope and the question
Scope is the step teams skip and then pay for. Model one system and its immediate neighbours: everything past the first hop is an external entity with an arrow pointing at it, which is what keeps a diagram from turning into a map of the company. Write the protection question down in one sentence before drawing anything, because every later argument about a control is really an argument about whether a particular loss matters.
2. Place the external entities
Start outside. An external entity is anything that interacts with the application from outside it, and the defining property is that you cannot change how it behaves: you decide only what you accept from it and what you send it. Each one is a place where your assumptions meet somebody else's decisions.
3. Draw the processes
A process is code that handles data: it reads something, decides something, writes something. One per responsibility you would defend separately, and the test is whether splitting the element would change an answer. Three copies of the same service behind a load balancer are one process: same code, same credentials, same level of trust. A gateway that authenticates and a service that acts on the authenticated request are two.
4. Add the data stores
A data store is anywhere data rests, and OWASP's definition is the one to draw by: data stores do not modify the data, they only store it. Databases are the obvious ones. The stores that get missed are the quiet ones: caches, message queues, object storage, log files, exports on a shared drive and backups. They hold the same data as the database, in much the same shape, and they are guarded less carefully.
5. Draw every flow with what it carries
Draw one arrow per direction, and label each one with three things: what data moves, over what protocol, and what authenticates it. The label to aim for reads "Order request, HTTPS, session cookie", with all three slots filled; an arrow labelled "API call" leaves every one of them open. The labels matter more than they look, because most threats are found by reading them: an arrow carrying personal data with nothing in the authentication slot has just written its own threat.
6. Mark the trust boundaries
Boundaries go on last, once there is something to draw them around. A trust boundary is used, in OWASP's words, "to represent the change of trust levels as the data flows through the application", and boundaries "show any location where the level of trust changes". Microsoft's tutorial draws them as red dotted lines "to show where different entities are in control", which is the practical test: if the two sides of the line have the same owner, the same credentials and the same exposure, there is no boundary there.
Three questions settle it, and in practice the boundaries you draw fall in four places.
Start here
A line between two elements on the page. Ask all three questions of it.
01
Same owner?
Ask who controls each side of the line. Two teams, two companies or two contracts is two owners.
02
Same credentials?
Ask what gets you in on each side. A separate account, key or session means the sides are not the same.
03
Same exposure?
Ask who can reach each side. One side open to the internet and the other reachable only from inside is two levels of exposure.
One of two outcomes
Yes to all three
No boundary here
Same owner, same credentials, same exposure: the line is decoration, and the walk gains nothing from it.
No to any one
A trust boundary
Draw it, list what crosses it, and put the controls on the crossings. Those flows are the ones the walk starts with.
Where they fall in practice
Outside your control
Where the public internet meets anything you run
Outside your control
Where a third party you cannot inspect takes over
Both sides yours
Where an application tier meets its data tier
Both sides yours
Where two of your own systems sit under different owners or different access rules
Two of the four have somebody else on the far side. With the boundaries on, the drawing has been built in five passes, one kind of thing at a time.
- External entity
- Process
- Data store
- Boundary and crossing
External entities
Everything outside your control goes down first.
Processes
One per responsibility you would defend on its own.
Data stores
Every place data rests, including the quiet ones.
Data flows
One arrow per direction, each one labelled.
Trust boundaries
Drawn last, where the level of trust changes.
What you have now
Four elements, five flows and two boundaries, every one of them numbered. Nothing has been analysed yet: the drawing is the thing the six questions then get asked of, one element at a time.
7. Number every element
Give every element a code and keep it: EE for external entities, C for components or processes, DS for data stores, DF for data flows, TB for trust boundaries. A numbered element can be named by a threat, a control, a test, a ticket and an audit finding, and all five will be talking about the same thing a year later.
8. Check it against the real system
Now go and look. Read the code, the configuration, the infrastructure definition, and correct the drawing where it is wrong. Most diagrams drawn from memory are wrong somewhere, and the wrong places are interesting: a flow nobody expected, or a credential shared between components meant to be separate.
9. Ask the six STRIDE questions, element by element
STRIDE is Microsoft's taxonomy of six threat categories, and the letters are the mnemonic: spoofing (using someone else's authentication information), tampering (malicious modification of data, at rest or in transit), repudiation (a user denies an action and nothing can prove otherwise), information disclosure (exposure to people not meant to have access), denial of service (denying service to valid users) and elevation of privilege (an unprivileged user gains privileged access). Ask all six of every process. For the other element types some questions have nothing to bite on, and the STRIDE walkthrough on our blog sets out which letters attach to which element type. Work the flows that cross a boundary first.
10. Decide every threat, and record the decision
Every threat leaves the session in one of three states: the design changes so the threat has nowhere to land, a control is added and written as something an engineer builds, or the risk is accepted by a named person with a date against it. Analysis with no decision is the habit the Threat Modeling Manifesto names as the anti-pattern "Admiration for the Problem".
How much detail is enough
Two diagrams of the same system can both be right at different depths, and teams talk about depth in levels: a context diagram that fixes the scope, a level 1 diagram that opens the system up into the parts holding or moving data, and a level 2 diagram that opens up one of those parts when a question needs it.
Level 0 · Context
The whole system as one process
EE1
Customer browser
Order request
Order history
TB1 · Internet edge
Process 0
Customer portal
EE1
Customer browser
Order request, down.
Internet edge. Both flows cross it.
Order history, up.
Process 0
Customer portal
One external entity, one process, one boundary. A context diagram this small is a finding rather than a failure: the portal has one way in, and anything that fills a page here means the scope is too wide to model in one sitting.
Open the one process up. The parts inside it that hold or move data become elements of their own.
Level 1 · The diagram this guide draws
The portal opened up, every element numbered
TB2 exists only at this level. At level 0 the database sat inside the single process, so the line between the application tier and the data tier had nowhere to be drawn. Opening the process up is what made it visible, and two of the five flows cross it.
Go one level further only when a question needs it, and only for the element the question is about.
Level 2 · Only when a question is asked
Inside C2, if anyone needs to see it
These are drawn the day somebody asks what the export job sends and where it sends it. Until then C2 stays one circle, because a level 2 diagram of every process is a week of drawing that answers nothing.
The level numbering is a working convention, not a standard: OWASP describes data flow diagrams as hierarchical, a high-level diagram for scope with lower-level diagrams that decompose individual parts, without fixing the numbers. Teams that say level 0, level 1 and level 2 mean context, decomposition and detail.
A worked example: a customer portal with an API and a database
An online retailer lets customers sign in to a portal and look at their past orders. The browser calls an API gateway, the gateway calls an order service, and the order service reads an orders database that holds line items, delivery addresses and the last four digits of the card used. A new release is about to add order cancellation, which is a change to what the service can do on a customer's behalf, so the team threat models the design before it is built.
Step 1, scope and question. One system and its immediate neighbours: the portal's read and cancel paths, and the customer at the other end of them. The protection question is written in a sentence: no customer should ever see, or change, another customer's order.
Steps 2 to 4, the elements. One external entity, EE1, the customer browser, which is a device the retailer does not control and cannot inspect. Two processes: C1, the API gateway, which ends the connection and authenticates the session, and C2, the order service, which builds and runs the query. One data store, DS1, the orders database. The team argued about whether the gateway deserved its own element and decided it did, because C1 holds the session and C2 does not: a compromise of one is not a compromise of the other.
Steps 5 and 6, flows and boundaries. Five flows and two boundaries, drawn in that order.
- External entity
- Process
- Data store
- Boundary and crossing
EE1
Customer browser
The signed-in customer, on a device you do not control.
DF1 down: the order request, over HTTPS on a session cookie.
Internet edge. Both flows cross it.
DF5 up: the order history, carrying addresses and the last four digits back.
C1
API gateway
Public entry point. Authenticates the session and forwards the account identifier.
DF2 down: the request forwarded on a service credential.
DF2 up: the reply, on the same flow inside one boundary.
C2
Order service
Builds every query from the account identifier C1 authenticated.
DF3 down: the order query and the cancellation update, scoped to one account.
Database tier. Both flows cross it.
DF4 up: the order rows, addresses and last four digits included.
DS1
Orders database
Where the orders rest: line items, delivery addresses, last four digits.
What each flow carries
- DF1EE1 to C1
- Order request: the session cookie the customer signed in with, and the identifier of the order they asked for, over HTTPS. Crosses TB1.
- DF2C1 to C2, both ways
- The same request forwarded with the account identifier C1 authenticated, and the reply, over the internal network on a service credential. Inside TB1.
- DF3C2 to DS1
- The order query and the cancellation update, both scoped to one account, on the database account C2 holds. Crosses TB2.
- DF4DS1 to C2
- Order rows: line items, delivery addresses and the last four digits of the card used. Crosses TB2.
- DF5C1 to EE1
- Order history rendered in the browser, carrying the same addresses and last four digits back out over HTTPS. Crosses TB1.
DF1 and DF5 are drawn separately because they carry different data and fail differently: a session cookie and an order identifier going in, delivery addresses and the last four digits of a card coming back out into a browser. DF3 carries two movements on one arrow, the order query and the cancellation update, because the release puts the write on the same path to the same database account; if the update ever moves to its own credential, it becomes its own flow. DF2 is the other judgement call. The request from C1 to C2 and its reply travel the same authenticated channel inside one boundary, so the team drew one flow with an arrowhead at each end and wrote down why: if the reply ever starts carrying data the request does not, it gets its own number.
TB1 is the internet edge: everything inside it is under the retailer's control, and EE1 is not. TB2 is the line between the application tier and the data tier, and it exists because the database account C2 holds can read and change every order in the system, while C2 itself is only ever meant to act for one customer at a time. Four crossings fall out of those two lines: DF1 and DF5 across TB1, DF3 and DF4 across TB2. Those four are where the session gets shown, where the query gets trusted and where the personal data comes back, and they are where the walk starts.
Steps 7 and 8, numbers and the check. Every element got a code. The check against the running system found one disagreement worth having: a second engineer had drawn C1 reading DS1 directly for a status widget on the account page. Reading the code settled it. The widget calls C2 like everything else, and the drawing above is what runs.
Step 9, the walk. Six questions of each element, and the flows that cross a boundary first.
S Spoofing
T Tampering
R Repudiation
I Information disclosure
D Denial of service
E Elevation of privilege
EE1Customer browser
External entity
C1API gateway
Process
C2Order service
Process
DS1Orders database
Data store
DF1Order requestCrosses TB1
Data flow
DF3Query and updateCrosses TB2
Data flow
DF4Order rowsCrosses TB2
Data flow
DF5Order historyCrosses TB1
Data flow
DF2Request and reply, inside TB1
Data flow
What the walk found
T1 · Spoofing
on EE1, landing on DF1 at TB1. A session cookie copied off a customer's machine is replayed from somewhere else, and C1 cannot tell the two apart.
Mitigated · CTL-01
T2 · Information disclosure
on C2. C2 looks an order up by the identifier in the request without checking who owns it, so changing a digit returns another customer's order.
Design change · CTL-02
T3 · Denial of service
on DF1, at TB1. Requests arrive faster than C1 can answer them, from an address nobody has signed in from, and real customers queue behind them.
Mitigated · CTL-03
T4 · Repudiation
on C2. A customer says they never cancelled the order. Nothing records which account asked for the cancellation, or when.
Mitigated · CTL-04
T5 · Information disclosure
on DS1. Delivery addresses and the last four digits sit in plain columns, so anything holding the database account reads every order at once.
Accepted · RISK-207
Four of the five threats left with a control written against them. The fifth left as a risk somebody owns, which is a decision with a name and a date on it rather than a gap.
Five threats came out of it, T1 to T5, and two of them are worth following here. T2 is the one that paid for the session. C2 looked orders up by the identifier in the request, and nothing in the query said which account owned them, so changing a digit in the request would have returned somebody else's order. T4 arrived with the cancellation feature and would not have existed a release earlier: a customer says they never cancelled the order, and nothing records which account asked.
C1 is the one process the walk finished empty, and the drawing explains why: it authenticates and forwards, it stores nothing, and the two questions it would have failed are asked of DF1, which reaches it, and of C2, which trusts what it sends. Four of the five flows came back empty as well. An element with nothing against it is a result, not a gap, and each one was recorded as such so the next person does not assume it was skipped.
Step 10, the decisions. T2 changed the design instead of adding a control: C2 now builds every query and every update from the account identifier C1 authenticated, and the identifier in the request only narrows a set that is already restricted to that account. The rule is written as CTL-02 and a test goes with it. T1 became CTL-01, which binds the session to the account and re-authenticates before an address change or a cancellation. T3 became CTL-03, a rate limit at C1 per account and per source address, written as a build requirement with a number in it. T4 became CTL-04: C2 writes an event for every order change, carrying the account, the time and the request identifier, into a store C2's own account cannot rewrite. That store is not on the diagram above, because it does not exist yet. It arrives on the next pass as a second data store with a flow into it, and the coverage figure drops the day it is drawn, because a new element turns up carrying six unanswered questions.
T5 was accepted. Encrypting the address columns meant a migration the release could not carry, so it became RISK-207 on the register, owned by the engineering lead, with an expiry at the end of March 2027 and two compensating steps recorded against it: DS1 accepts connections only from C2, and reads on it are logged. Accepting a risk this way is a decision with a name and a date on it, which is a different thing from not having noticed.
What the numbers did afterwards. CTL-01 to CTL-04 left the review as build requirements written against the elements they protect, so each one can be tested without reading the whole design: the tester checks CTL-04 by cancelling an order and reading the event back. When a nightly export job is added inside C2 two releases later, the level 1 diagram does not change, because the job sits inside an element already drawn. A level 2 diagram of C2 answers what the export sends and where it goes, and the new flow out of the retailer's network gets its own crossing and its own questions. An assessor asking later how the portal keeps one customer out of another customer's orders gets CTL-02, the threat that produced it, the element it sits on and the test that proves it.
Common mistakes
Eight failure modes account for most of the data flow diagrams that produce nothing. They share one root: the drawing gets treated as the deliverable instead of as the thing the questions are asked of.
- Drawing the deployment instead of the data. A network diagram answers where the software runs, which is a real question and a different one. The fix is to translate it: a box becomes an element only if it holds or handles data, and the copies collapse into one.
- No trust boundaries, so every element on the page looks equally exposed. The fix is to draw the boundaries before the walk starts, and to start the walk on the flows that cross them.
- Unlabelled arrows. An arrow with no label is an unasked question, because the question is in the label. The fix is three slots on every flow: what moves, over what protocol, authenticated how.
- One diagram for the whole estate. It takes a month, it is wrong by the time it is finished, and nobody opens it. The fix is OWASP's: one high-level diagram for scope, plus focused ones for the sub-systems that need them.
- Chasing a perfect picture. The Threat Modeling Manifesto names this anti-pattern "Perfect Representation". A diagram that is good enough to argue with beats a diagram that is still being drawn.
- Forgetting the quiet stores. Caches, queues, log files, exports and backups hold the same data as the database and are guarded less. The fix is to ask where else this data exists before the boundaries go on.
- Analysing without deciding, which the same manifesto calls "Admiration for the Problem". The fix is a rule that no threat leaves the session without one of three decisions and a name against it.
- A diagram that goes stale because nobody owns it. The fix is to re-open the model on the four triggers above, not on a calendar.
The first mistake is worth showing, because the translation out of it is mechanical. The portal has a perfectly good infrastructure drawing: an internet gateway, a load balancer called alb-01 in a public subnet, three application servers called app-01 to app-03 in another, and db-primary and db-replica in a third. It is accurate, it is useful for other work, and almost none of it survives the translation.
What the team brought
A deployment drawing, in the grey it is usually drawn in. It answers where the software runs, which is a real question and a different one.
What the threat model needed
The same system asked a different question: what data moves, and where does the level of trust change. Seven boxes and four frames become four elements, two flows and two boundaries.
- External entity
- Process
- Data store
- Boundary and crossing
The internet gateway at the top
EE1 · Customer browser, and TB1 · Internet edge
The network drawing shows the pipe. The data flow diagram shows who is on the other end of it and what they send.
alb-01, in the public subnet
C1 · API gateway
It ends the connection and authenticates the session, so it handles data and earns a process of its own.
app-01, app-02 and app-03
C2 · Order service
Three copies of the same code at the same level of trust are one process. Copies change how much load it takes, not what can go wrong.
db-primary and db-replica
DS1 · Orders database
One data store, for the same reason: the replica holds the same rows, reachable on the same account.
The lines from the app subnet to the db subnet
DF3 and DF4, crossing TB2
Two flows, each labelled with what it carries and what authenticates it. The boundary is drawn because the level of trust changes there, not because a subnet does.
Seven boxes and four frames become four elements, and two things appear that the infrastructure drawing had no way to show: the customer at the other end of the connection, and the line where trust changes.
A data flow diagram compared with its neighbours
Two questions separate them: how much of the system each drawing covers, and whether it is drawn to explain the system or to find what could go wrong.
Threat model
A whole estate
A whole system
One system
One scenario
Explaining how it is built
Finding what could go wrong
What separates each one from a data flow diagram
Network diagram
Where things run. A data flow diagram asks what moves between them.
Architecture diagram
How the parts fit together. A data flow diagram asks what passes between them.
Data-centric model
Starts from the data to be protected rather than from the system drawing.
Data flow diagram
What moves between the parts, in which direction, and where the level of trust changes.
Sequence diagram
One scenario in time order rather than the whole system at once.
Attack tree
Starts from the outcome you fear rather than from a survey of the system.
The threat model is not one of the drawings. It is the whole analysis, and it sits around whichever of them you used: the model, the threats found, the decision taken on each one, and who accepted what.
Everything in the right-hand half is a way of finding what could go wrong. The rest clears up in one line each.
| Term | What it is | Reach for it when |
|---|---|---|
| Data flow diagram | What data moves between the parts of a system, in which direction, and where the level of trust changes. | You want to know what could go wrong with a design, and you want every threat anchored to something specific. |
| Architecture diagram | The parts of a system and how they fit together, usually including hosting and technology choices. | You are explaining the shape of a system to someone who has to build it, buy it or approve it. |
| Network diagram | Segments, addresses and devices: where things run and what can reach what. | You are designing or debugging connectivity, or showing an assessor how segmentation works. |
| Sequence diagram | The order of calls between parts over time, for one scenario at a time. | One interaction is subtle and the order matters, such as a sign-in, a payment or a token exchange. |
| Attack tree | One attacker goal at the top, broken down into the ways of reaching it. | You already know the outcome you fear most and want the paths to it rather than a survey of the system. |
| Data-centric model | The data to be protected at the centre, with the attack and defence sides modelled around it. NIST's draft SP 800-154 is the published example. | The thing you must protect is obvious and the system around it is large, shared or not yet drawn. |
| Threat model | The whole analysis: the model, the threats found, the decision taken on each one, and who accepted what. | Someone asks what you did about security in this design. The diagram is one part of the answer. |
Microsoft makes a point in its own tutorial that stops the comparison turning into an argument: thinking about assets, thinking about attackers and thinking about software design are all workable entry points into threat modeling. Microsoft teaches the software-design entry and says why plainly, that many software engineers understand their software better than they understand the concept of assets. Pick the entry your team can start from.
Standards and references
OWASP Cheat Sheet Series, Threat Modeling Cheat Sheet (read September 2026). The current OWASP guidance: the four questions, the "arguably the most common approach" line, the advice to keep one high-level diagram plus focused ones, and the tooling list. The page carries no publication date.
OWASP, Threat Modeling Process (read September 2026). The older OWASP page, which OWASP itself marks historical. Its element definitions are still the clearest short ones in print, including the privilege boundary as the change of trust levels as data flows through an application.
Microsoft Threat Modeling Tool documentation (articles dated 17 August 2017, getting started updated 10 December 2025, threats updated 4 March 2026) and the Microsoft Learn module Create a threat model using data-flow diagram elements (27 March 2025, updated 3 April 2025). The module teaches the five element types one unit each, and of the pages read for this guide it is the clearest short teaching of the notation. The threats article carries the STRIDE definitions quoted above, and the tool it documents generates threats per interaction from them.
The Threat Modeling Manifesto (2020). Written by fifteen practitioners including Adam Shostack. It holds the four questions, a set of values and principles, and the two anti-patterns named in this guide.
Adam Shostack, Threat Modeling: Designing for Security, Wiley (February 2014). Book-length coverage of the practice, from one of the manifesto's authors. A second edition, Threat Modeling: Designing for Security in an AI World, is announced for 2 February 2027.
NIST SP 800-154, Guide to Data-Centric System Threat Modeling (initial public draft, March 2016). The published alternative entry point, starting from the data to be protected instead of the system drawing. It has been an initial public draft for a decade, with a NIST note in January 2025 stating an intention to finalise it, so cite it as a draft.
CISA, NSA, FBI and fifteen international partners, Shifting the Balance of Cybersecurity Risk (April 2023, revised 25 October 2023). Guidance for software manufacturers, and the source of the position that secure by design products start with written threat models.
How Alvor helps
With Alvor, the diagram is the record rather than a picture of one. You model the system on a canvas where components, datastores, data flows, external entities and trust boundaries are all first-class elements. Each one is numbered, C1 for a component or DS1 for a datastore, and each can link to the asset it represents in your inventory, so the element on the drawing and the asset in the register are the same thing rather than two descriptions of it.
Threats hang off those elements. A threat is anchored to the element it targets and carries a STRIDE category, a severity and a status, and it can reference a MITRE ATT&CK technique by ID, sit on a Cyber Kill Chain phase or cite an OWASP or CAPEC source. Mapping a threat to a control assigns that control to the project and flags it as threat-driven. Those are the same controls the Compliance module cross-maps across frameworks such as ISO 27001, SOC 2 and NIST CSF, so a mitigation decided at the whiteboard becomes evidence for every framework that control satisfies. A threat you cannot mitigate is accepted as a risk with an owner and a rationale on the register, which is where an entry like the worked example's RISK-207 lives.
Coverage recounts as the diagram changes, so an element added in a redesign shows up as unanalysed, not quietly unasked, and the question "which parts of this have we looked at" keeps a current answer. For the drawing itself, the AI Design Studio draws the architecture as editable shapes from a written description or a photograph of a whiteboard, which needs a vision-capable model, and the AI Threat Modeling Studio reads the diagram, registers elements and proposes threats from your library, with each batch pausing on an approval card showing the exact payload for a person to approve or reject.
Every module runs in every deployment: a dedicated single-tenant instance in the region you choose, the same containerised platform on your own servers, or fully air-gapped with a self-hosted model such as vLLM or Ollama inside the boundary. Threat Modeling in Alvor covers the canvas, the STRIDE walk and the coverage figures in more detail.