Privacy and Data Access#
A tool you point at a production database has to be clear about what it reads. Test Data Workbench is metadata-first by design, and can be made metadata-only by a flag. This chapter states exactly what touches your rows and what does not.
What the analyzer reads#
Analysis has two layers:
Structure comes entirely from schema reflection: table names, columns, types, primary keys, and foreign keys, read through SQLAlchemy’s inspector. This layer never reads a single row of your data.
Value patterns (a column’s value ranges, whether a text column holds email-shaped values) require looking at data. By default the analyzer samples a bounded number of rows per table to inform this layer.
Metadata-only mode#
Pass --metadata-only to tdw analyze or tdw deploy and the second
layer is switched off completely. In this mode the analyzer issues zero SQL
statements that read table data: no SELECT, no COUNT, no sampling. It
works from reflected structure alone.
tdw analyze "postgresql://user:pass@host/db" --metadata-only
This is the mode to use when you are not permitted to read the data itself, only its shape. The guarantee is verified by a test that captures every SQL statement executed during a metadata-only analysis and asserts that none of them read a user table. The trade-off is that value-pattern detection is unavailable, so classification leans entirely on names and types.
PII flagging#
The analyzer flags columns that are likely to hold personal data, so that a
schema review surfaces them explicitly. Flagging is based on column-name
heuristics (an email column, an ssn column, an address column), which
means it works from metadata alone and therefore functions in metadata-only mode
as well.
The flag is a prompt for human judgment, not a classifier with authority. It is reported as a name-based heuristic and named as such. It will miss PII hidden behind an unhelpful column name, and it may flag a column whose name suggests PII but whose contents do not. Treat it as a checklist aid.
What this is not#
The privacy features here are about access discipline, controlling what the tool reads and surfacing what is sensitive. They are deliberately not the following, each of which is a substantial field in its own right:
Differential privacy. The tool adds no calibrated noise and makes no formal privacy guarantee about a generated dataset’s relationship to a real one.
k-anonymity, l-diversity, t-closeness. These are properties of a release of real (masked) data. Workbench generates synthetic values from structure; it does not transform and release your rows.
Production-data masking. Preserving the statistical shape of real data while anonymizing it is a different product with a different risk profile.
Naming these explicitly is the point of the chapter. A tool that gestures at “privacy” without saying which guarantees it does and does not make is harder to trust than one that draws the line clearly.