Architecture#
Test Data Workbench is a pipeline: a schema goes in one end, a directory of generator code comes out the other. The stages are separable and each has a single responsibility.
The pipeline#
graph TB
subgraph Interface
CLI[tdw CLI]
API[FastAPI REST API]
end
subgraph "Adaptation Engine"
SA[Schema Analyzer]
GF[Generator Factory]
CB[Config Builder]
RD[Rapid Deployment]
end
subgraph Safety
DM[Defensive Manager]
TG[Template Generator]
V[Validators]
end
CLI --> SA
CLI --> RD
API --> RD
SA --> GF
GF --> CB
RD --> SA
DM --> GF
TG --> CB
V --> DM
Stages#
Schema Analyzer#
The entry point. Given a SQLAlchemy connection string, it reflects the database
structure through SQLAlchemy’s Inspector: table names, columns, types,
primary keys, and foreign keys. It then does two inference passes on top of the
raw structure:
Entity classification. Each table is scored against name patterns to assign an entity type (user, product, order, and so on). The scoring is heuristic and deliberately transparent; its limits are documented in Schema Adaptation.
Relationship and rule detection. Foreign keys become explicit relationships; recurring shapes (a non-null unique
email, a positiveprice) become candidate business rules.
Structural analysis never requires reading row data. Value-pattern detection can sample a bounded number of rows, and can be switched off entirely; see Privacy and Data Access.
Generator Factory#
Turns the analysis into code. For each table it selects a generation strategy (the fallback ladder in The Fallback Ladder) and emits a generator class. Tables are processed in dependency order via a topological sort of the foreign- key graph, so a table’s parents are generated before it. The sort includes cycle detection: a cyclic foreign-key graph degrades to a defined order instead of looping forever.
Defensive Manager#
Owns the fallback ladder at runtime and isolates failures. If a specialized generator raises, the manager drops to the next ladder level for that entity rather than aborting the run. Errors are tracked, not swallowed silently.
Template Generator and Config Builder#
The Template Generator produces the human-facing scaffolding that ships alongside the generators (contribution templates at graded complexity levels). The Config Builder emits the YAML configuration and volume-scenario presets (small store, growing business, enterprise, testing) that let a generated set be re-run at different scales.
Validators#
Gate generated and contributed code before it is trusted. Checks are
AST-level, not string matching: syntactic validity, absence of unsafe imports,
and presence of the required generator interface. This is what lets the tdw
validate command give a meaningful verdict on a hand-edited generator.
Rapid Deployment#
The orchestrator behind tdw deploy. It runs analyze, factory, config, and
templating end to end and writes the output tree, reporting timing per phase.
Constraint check#
After generation, a lightweight check runs the generated code over a small
sample and verifies that the result satisfies the constraints it can discover:
primary-key uniqueness, populated not-null columns, and foreign-key values that
reference generated parent rows. The outcome is reported as a one-line summary
(for example, Constraint check: 25 checks passed). It is a best-effort
verification of the output, not a gate that can block a deploy.
Interfaces#
The same engine is reachable two ways:
CLI (
tdw): the primary interface, documented in CLI Reference.FastAPI REST API: the same operations over HTTP with automatic OpenAPI documentation, for integration into other services.
Package layout#
src/test_data_workbench/
├── core/ # models, defensive manager, validators,
│ # template generator, feature flags, demo coordinator
├── adaptation/ # schema analyzer, generator factory,
│ # config builder, rapid deployment
├── api/ # FastAPI app and endpoints
└── cli/ # tdw entry point, one module per subcommand