Python#
Python is the implementation language for the whole system: the schema
analyzer, the generator factory, the CLI, and the optional API. Every
dependency in pyproject.toml, SQLAlchemy for schema reflection, Faker for
values, pandas for pattern detection, Jinja2 for code generation, FastAPI and
Pydantic for the optional HTTP layer, aiofiles, PyYAML, Rich, lives in the
same ecosystem. There is no compiled extension, no second language, and no
separate build toolchain anywhere in the stack; a contributor who can read
one file can read all of them. This page covers the language features the
codebase actually leans on, not a general Python pitch.
The idea underneath#
Python’s type hints are an instance of gradual typing: a type system that
lets a program mix statically checked and dynamically checked code, so a
codebase can be annotated incrementally rather than all at once or not at
all. PEP 483, “The Theory of Type Hints,” lays out that theory, and PEP 484,
“Type Hints,” turned it into the concrete syntax and standard-library
typing module Python has used since 3.5.
The detail that matters most for reading this codebase is what gradual
typing deliberately does not do: CPython erases annotations at runtime. A
function declared def f(x: int) -> str will happily run f("not an
int") and return whatever it returns; nothing in the interpreter consults
the annotation to stop it. PEP 484 says so explicitly: the hints exist for
static analysis tools, IDEs, and human readers, not for the runtime to
enforce. That single fact is why this project also depends on Pydantic (see
Pydantic): Pydantic is the library that reads the same annotations
and, at one specific boundary, actually checks them when an object is built.
Everywhere else in this codebase, an annotation is a claim a static checker
or a reader could verify, not a promise the interpreter keeps.
How it fits this project#
Every public function and dataclass field in core and adaptation is
type-hinted. Those annotations feed two other things directly:
sphinx.ext.autodoc, configured in docs/source/conf.py and applied to
test_data_workbench.core.models and
test_data_workbench.adaptation.schema_analyzer in
API Reference, renders them into the API reference without a
separate parameter-documentation step, and Pydantic’s BaseModel classes
in api/models.py build their request and response validation from
annotated fields the same way. Nothing else in the toolchain reads them; see
Sharp edges and limits.
The src layout#
pyproject.toml declares [tool.hatch.build.targets.wheel] packages =
["src/test_data_workbench"], and the actual package lives at
src/test_data_workbench/, not at the repository root. docs/source/conf.py
has to compensate for it directly:
# Make the package importable for autodoc (src layout).
sys.path.insert(0, os.path.abspath("../../src"))
That extra line is the src layout’s whole tradeoff in one place. A flat
layout, with the package directory sitting next to pyproject.toml, lets
Python import test_data_workbench straight out of the working directory,
because the interpreter always puts the current directory on sys.path.
That is convenient and also exactly the failure mode the src layout closes
off: an import that only succeeds because of where a command happens to be
run, which can hide the fact that the package was never actually installed,
that an entry point is missing, or that a dependency did not get pulled in.
Under src/, import test_data_workbench fails from the repository
root unless the package has been installed, editable or not, exactly as it
would be for anyone else using it. The CI workflows exercise that same path
on every run: pip install -e ".[dev]" before pytest, pip install -e
. before the Sphinx build.
Type hints without enforcement#
Type hints appear throughout core and adaptation, favoring the
explicit typing module spellings, Optional[int], Dict[str,
List[Any]], over the newer | union syntax. Most of the codebase
sticks to this consistently, but not uniformly: cli/analyze.py declares
a local variable as entity_counts: dict[str, int] = {}, the lowercase
built-in generic rather than typing.Dict. Several cli modules
(main.py, analyze.py, deploy.py, demo.py, team.py,
validate.py, _seeding.py) also open with from __future__ import
annotations, which postpones evaluation of every annotation in that file
to a string; core and adaptation do not use that import. Neither
choice changes what runs, since annotations are not evaluated for
enforcement either way, but it means the CLI layer and the core layer do not
share one typing style.
No tool in this repository checks whether any of these annotations are
actually true. .github/workflows/ci.yml runs three jobs, a pytest matrix
across Python 3.10 to 3.12 on Linux and Windows, a coverage run, and a
Postgres integration test, and none of them invoke mypy, pyright, or any
other static checker; there is no [tool.mypy] section in
pyproject.toml either. That means two things follow directly. First, the
annotations in core and adaptation are documentation and
autodoc input, nothing more; they carry the same weight whether they are
right or wrong, because nothing runs to tell the difference. Second, when a
signature and its default disagree, nothing catches it. DefensiveGeneratorManager.register_generator
in core/defensive_manager.py is typed:
def register_generator(self, entity_name: str, generator_class: Type,
fallback_class: Optional[Type] = None,
entity_type: EntityType = EntityType.UNKNOWN,
dependencies: List[str] = None) -> None:
dependencies is annotated List[str] with a default of None, a
value no List[str] can ever hold. A type checker running in strict mode
would reject this outright; none runs, so it ships. create_scenario_template
in core/template_generator.py has the identical shape:
entities: List[str] = None. Both work at runtime, because Python never
looks at the annotation to begin with, and both rely on a caller either
passing a real list or reading the docstring to know that None is
actually the accepted default.
Dataclasses versus Pydantic#
core/models.py defines the system’s internal vocabulary, twelve
@dataclass types: Column, Relationship, Table,
BusinessRule, PIIColumn, SchemaInfo, GeneratorConfig,
ScenarioConfig, AdaptationResult, ValidationResult,
DemoScenario, and DemoStatus. core/feature_flags.py adds
FeatureFlags the same way, and core/defensive_manager.py adds
GeneratorResult and GeneratorInfo. None of these are Pydantic
models; Pydantic is confined to api/models.py, the one file where
untrusted JSON arrives over HTTP (see Pydantic for that side in
detail). Internally, between the analyzer, the factory, the defensive
manager, and the CLI, data never crosses a trust boundary, so a plain
dataclass is the lighter-weight choice for the same amount of type
information.
Mutable defaults follow the documented dataclass idiom, field(default_factory=...),
consistently. Column.constraints and Column.sample_values both
default to an empty list this way; FeatureFlags.enabled_features
defaults to a set literal through a default_factory=lambda: {...}. None
of these use a bare mutable literal as the default value itself, which is
the classic version of this mistake (a default list built once at
def-time and then shared, and silently mutated, across every call that
does not pass its own).
A few of these dataclasses add an alternate constructor as a
@classmethod instead of overriding __init__: BusinessRule.create
builds a rule_id from a truncated uuid.uuid4() and fills in the
rest, and FeatureFlags.safe_mode returns an instance with only the
safest feature set enabled:
@classmethod
def create(cls, description: str, rule_type: str, columns: List[str],
confidence: float = 0.8) -> 'BusinessRule':
return cls(
rule_id=str(uuid.uuid4())[:8],
description=description,
rule_type=rule_type,
affected_columns=columns,
confidence=confidence
)
This keeps the generated __init__ as the one place field defaults live,
while giving call sites a named, more specific way to build a common case.
Enums for closed vocabularies#
core/models.py defines two Enum classes for values that should never
be arbitrary strings: EntityType (eleven members: USER,
CUSTOMER, PRODUCT, and so on down to UNKNOWN) and
ConstraintType (PRIMARY_KEY, FOREIGN_KEY, UNIQUE,
NOT_NULL, CHECK, INDEX). They are used for real membership
checks, not just labels; Column.is_primary_key is exactly
ConstraintType.PRIMARY_KEY in self.constraints, which only works, and
only means what it says, because constraints holds ConstraintType
members rather than the strings "primary_key" or "pk" that a
careless caller could otherwise substitute.
The same file does not apply this consistently. BusinessRule.rule_type
is a bare str with a comment naming its actual closed set,
# "format", "range", "dependency", "pattern"; Relationship.relationship_type
is the same, # "one_to_one", "one_to_many", "many_to_many"; and
PIIColumn.category is a str populated from the ten keys of
PII_NAME_PATTERNS in adaptation/schema_analyzer.py
("email", "name", "phone", and so on). Every one of these is a
closed, known, comment-documented vocabulary in the same file that defines
EntityType for exactly that purpose, and none of them is an Enum. A
value of rule_type="pattern-ish" or category="unknown" is
accepted the same as a correct one; only the two constraint- and
entity-facing fields get the type-safety Enum provides. The parallel
case on the API side, plain str fields in api/models.py where a
worked example implies a fixed set of choices, is covered in
Pydantic.
Async: real concurrency in one place, none in another#
asyncio.gather appears in two places in this codebase, and only one of
them overlaps real I/O.
SchemaAnalyzer.analyze_production_schema in
adaptation/schema_analyzer.py is async def. It builds one
_analyze_table coroutine per discovered table, capped at twenty
(table_names[:20] # Limit for rapid analysis), and runs them together:
tasks = [
self._analyze_table(engine, inspector, table_name, metadata_only)
for table_name in table_names[:20] # Limit for rapid analysis
]
tables = await asyncio.gather(*tasks)
_analyze_table is also async def, and it awaits two helpers,
_get_sample_values and _estimate_row_count, both declared async
def in turn. Read their bodies and the async stops being real: both
connect through the plain synchronous SQLAlchemy Engine created earlier
with sa.create_engine(connection_string, pool_pre_ping=True), not
sqlalchemy.ext.asyncio.create_async_engine, and both call
conn.execute(query) directly.
async def _get_sample_values(self, engine: sa.Engine, table_name: str,
column_name: str, limit: int = 10) -> List:
"""Get sample column values for pattern detection."""
try:
with engine.connect() as conn:
query = text(f'SELECT DISTINCT "{column_name}" FROM "{table_name}" LIMIT {limit}')
result = conn.execute(query)
return [row[0] for row in result if row[0] is not None]
except Exception:
return []
There is no await anywhere inside that function, or inside
_estimate_row_count beside it. Both are async only by signature: marked
async def so they can be called with await, but nothing in their own
bodies ever reaches a genuine suspension point. The Python documentation is
direct about the consequence of calling blocking code from inside a
coroutine this way, in the “Running Blocking Code” section of Developing
with asyncio: a blocking call delays every other concurrent task sharing
that event loop, the same way a one-second CPU-bound calculation would.
Fanning _analyze_table calls out through asyncio.gather cannot
overlap their database queries the way it would if those queries were
issued through an async driver and awaited properly; each table’s blocking
work runs on the same single thread as everything else the event loop is
doing, with no run_in_executor handoff to release it early.
Contrast that with RapidDeployment._write_file_async in
adaptation/rapid_deployment.py, which does get real overlap:
try:
import aiofiles
async with aiofiles.open(file_path, 'w', encoding='utf-8') as f:
await f.write(content)
except ImportError:
with open(file_path, 'w', encoding='utf-8') as f:
f.write(content)
aiofiles runs the actual file write on a background thread and awaits
its completion, the same executor pattern the asyncio documentation
recommends for blocking calls, so await f.write(content) is a real
suspension point. _write_output_files fans several of these out through
await asyncio.gather(*write_tasks, return_exceptions=True), and because
each write genuinely yields control while its thread finishes, the writes
do overlap. (The except ImportError fallback would itself be another
async-only-by-signature call if it ever ran; it should not, since
aiofiles>=23.2.1 is a pinned dependency.)
Either way, what asyncio buys here is concurrency for I/O-bound work
sharing one event loop and one thread, never parallelism across CPU cores;
Python’s own GIL rules that out for ordinary async code regardless of how
many tasks gather schedules. The CLI is the synchronous/async boundary
in both directions: cli/analyze.py, cli/deploy.py, and
cli/demo.py each call into this async code with a single
asyncio.run(...), so the rest of the CLI, argument parsing, Rich output,
file saving, stays ordinary synchronous Python.
The standard library first#
The CLI entry point, cli/main.py, is built on argparse from the
standard library. typer is not a dependency of this project; every
subcommand is registered by hand through argparse.ArgumentParser.add_subparsers,
one add_subparser(subparsers) function per module
(analyze, deploy, validate, team, demo), each returning
the argparse.ArgumentParser it configured:
def build_parser() -> argparse.ArgumentParser:
# Imported lazily so that `tdw --version` stays fast and side-effect free.
from . import analyze, deploy, demo, team, validate
parser = argparse.ArgumentParser(
prog=PROG_NAME,
description="Test Data Workbench - schema-adaptive test data generator. "
"Point it at a database, get realistic relational test data back.",
)
...
subparsers = parser.add_subparsers(dest="command", metavar="command")
analyze.add_subparser(subparsers)
...
Each subcommand module wires its own arguments the same way;
cli/analyze.py is representative, adding a positional connection, a
--format choice among table/json/summary, an --output
path, and a --metadata-only flag, then binding the subcommand to a
handler with parser.set_defaults(func=_run). Subcommand modules are
imported inside build_parser() rather than at module load time, so that
tdw --version stays fast and does not pay the import cost, and pull in
the dependencies, of every subcommand it is not running. Rich (see
Rich) sits on top of this for output, tables, panels, and progress
spinners, but it plays no part in defining or parsing arguments; that is
argparse end to end. It mirrors the project’s broader dependency
discipline described in Stack overview: reach for the standard library before
adding a dependency, and a CLI parser is squarely inside what argparse
already does well.
Version floor: 3.10, and what depends on it#
pyproject.toml declares requires-python = ">=3.10", with classifiers
for 3.10, 3.11, and 3.12, and the CI matrix in Testing and Quality
runs all three, on both Ubuntu and Windows.
Nothing in this codebase’s own source actually requires 3.10. Two syntax
features shipped in that release: the match/case statement (PEP
634) and the X | Y union syntax in annotations (PEP 604). Neither
appears anywhere under src/; there is no match statement in the
project, and every union type is written the older way, Optional[int],
Union[str, int]. Nothing else in 3.10 changes what syntax is legal
either; the release’s other headline items, better SyntaxError and
NameError messages (“did you forget parentheses around the
comprehension target?”, “did you mean: …”), and parenthesized with
statements spanning multiple lines, are interpreter and grammar
conveniences a codebase benefits from just by running on 3.10, not features
its own code has to opt into. So the floor is a deployment-target choice
more than a language-feature requirement: 3.10 is old enough to be widely
available and new enough to still receive attention, and the project simply
has not needed anything past it.
Learn more resources#
Official documentation
PEP 483: “The Theory of Type Hints,” the gradual-typing theory this page’s “idea underneath” section names.
PEP 484: “Type Hints,” the specification that made annotations erased-at-runtime by design, and the same PEP Pydantic cites for the boundary Pydantic enforces instead.
src layout vs flat layout: the Python Packaging User Guide’s own explanation of the failure mode the src layout closes off.
dataclasses: the standard-library module behind every internal model in
core/models.py.enum: the module behind
EntityTypeandConstraintType.argparse: the module
cli/main.pyis built on.asyncio - Coroutines and Tasks: documents
asyncio.gather, used in bothschema_analyzer.pyandrapid_deployment.py.Developing with asyncio: the “Running Blocking Code” section documents exactly the pitfall
_get_sample_valuesand_estimate_row_countrun into, and therun_in_executorescape hatch they do not use.PEP 604: the
X | Yunion syntax that shipped in 3.10 and that this codebase does not use anywhere.PEP 634: the
match/casestatement specification, the other 3.10 syntax feature absent from this codebase.What’s New In Python 3.10: the full list of what the version floor actually adds, syntax and otherwise.
Tutorials and blogs
Using mypy with an existing codebase: mypy’s own guide to the path this project would take if it ever added the static checker described in Sharp edges and limits.
Speeding Up Python with Concurrency, Parallelism, and asyncio: a walkthrough of the concurrency-versus-parallelism distinction this page’s async section relies on.
Python 3.10: What’s New: an accessible tour of the match statement and the other 3.10 features this codebase does not use.
Videos
Łukasz Langa - Thinking In Coroutines - PyCon 2016 (PyCon 2016): a talk on how coroutines and the event loop actually behave underneath
async/await, the model this page’s async section depends on.