← CISSP: security, risk, and operations
16 / 17 · 65 MIN

Classification, exports and indirect identification

Define the purpose of a release and assess what recipients can infer from its data.

Start with authorized use

A data request needs a concrete purpose. In a fictional funds project, analysts want to measure delays by region; that does not mean they need names, account numbers or support notes. The information owner defines classification and use requirements under policy; the custodian implements rules and retains evidence. Distinguish permission to view an application from permission to create a reusable export. The new copy has its own recipients, location and retention period. Record necessary fields and the justification for each. Classification helps select controls, but a label on a file does not itself prevent reading, copying or unauthorized transfer.

An explicit output schema

A SELECT * query can start including a confidential field added after export approval. Explicit output columns prevent that silent expansion. Code inspection alone is insufficient: inspect the generated file, its header, population and transformed values. An approved field can acquire a new meaning or contain free text with unexpected data. The lab adds internal_rating to the table and compares an explicit projection with a wildcard selection. The former retains three columns; the latter exposes the new field. This establishes behavior of that export contract without claiming that every combination of those three columns is suitable for public release.

Removing a name does not remove every link

Indirect identifiers can link a record to outside information. In the exercise, the released file omits names but retains birth year and a zone code. A synthetic auxiliary table contains those combinations and fictional identifiers. The join finds one match per row. This result was deliberately constructed to teach the failure and does not estimate identification rates in a real population. Before sharing data, identify auxiliary information a recipient might possess and attributes that could support linkage. Stable pseudonyms can also enable joins across datasets. Separately protecting a mapping table reduces one exposure route but does not automatically eliminate every other route.

Larger groups and attribute disclosure

Generalization can increase the number of records indistinguishable under selected identifiers. In the lab, birth decade and region create two groups of four rows. For those two fields, the smallest group has four members. However, every row in one group has the same fictional credit status. Someone who knows a person belongs to that group can infer the attribute without identifying which of the four rows is theirs. Group size therefore does not provide a complete privacy guarantee. Also review sensitive attributes, auxiliary information, rare values, output utility and access context. Do not invent diversity merely to obtain a favorable metric.

Accepting a release

A useful assessment compares options: reduce precision, remove fields, aggregate, restrict access or create entirely synthetic data. Each can affect the intended analysis. For regional delays, check that transformation preserves the ability to answer the project question. Record schema version, assessed population, transformation and decision owner. Reassess when recipient, purpose or dataset changes; earlier approval does not necessarily cover a new combination. The lab demonstrates SQL, CSV and linkage of fictional data. It does not certify anonymity, legal compliance or permission to release actual data. Production decisions require the applicable policies and accountable owners.

-- Reading exercise; synthetic fixture schema, not a production query.
SELECT birth_year, postal, credit_status
FROM people
ORDER BY id;
-- The explicit projection prevents new columns from entering automatically.
-- It does not prove these values are safe to publish.
IN PRACTICE

The nameless file still allows every row to link to the auxiliary table. After generalization, one group of four still reveals a shared sensitive attribute.

Common pitfalls

Nameless as anonymous; pseudonym as authorization; SELECT * as a stable contract; k as proof of no disclosure.

Related topics: Governance and risk · Key custody and recovery · Evidence and assessment

Take this idea with you

Purpose determines necessary fields; linkage analysis determines the risk that remains.

Create account

Reference: De-Identification of Personal Information · CISSP outline effective April 15, 2024; current AI guidance consulted 2026-09-29

CISSP® is a registered trademark of ISC2, Inc. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by ISC2. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.