Python#

Python is the implementation language for the whole system: the schema analyzer, the generator factory, the CLI, and the optional API. Every dependency in pyproject.toml, SQLAlchemy for schema reflection, Faker for values, pandas for pattern detection, Jinja2 for code generation, FastAPI and Pydantic for the optional HTTP layer, aiofiles, PyYAML, Rich, lives in the same ecosystem. There is no compiled extension, no second language, and no separate build toolchain anywhere in the stack; a contributor who can read one file can read all of them. This page covers the language features the codebase actually leans on, not a general Python pitch.

The idea underneath#

Python’s type hints are an instance of gradual typing: a type system that lets a program mix statically checked and dynamically checked code, so a codebase can be annotated incrementally rather than all at once or not at all. PEP 483, “The Theory of Type Hints,” lays out that theory, and PEP 484, “Type Hints,” turned it into the concrete syntax and standard-library typing module Python has used since 3.5.

The detail that matters most for reading this codebase is what gradual typing deliberately does not do: CPython erases annotations at runtime. A function declared def f(x: int) -> str will happily run f("not an int") and return whatever it returns; nothing in the interpreter consults the annotation to stop it. PEP 484 says so explicitly: the hints exist for static analysis tools, IDEs, and human readers, not for the runtime to enforce. That single fact is why this project also depends on Pydantic (see Pydantic): Pydantic is the library that reads the same annotations and, at one specific boundary, actually checks them when an object is built. Everywhere else in this codebase, an annotation is a claim a static checker or a reader could verify, not a promise the interpreter keeps.

How it fits this project#

Every public function and dataclass field in core and adaptation is type-hinted. Those annotations feed two other things directly: sphinx.ext.autodoc, configured in docs/source/conf.py and applied to test_data_workbench.core.models and test_data_workbench.adaptation.schema_analyzer in API Reference, renders them into the API reference without a separate parameter-documentation step, and Pydantic’s BaseModel classes in api/models.py build their request and response validation from annotated fields the same way. Nothing else in the toolchain reads them; see Sharp edges and limits.

The src layout#

pyproject.toml declares [tool.hatch.build.targets.wheel] packages = ["src/test_data_workbench"], and the actual package lives at src/test_data_workbench/, not at the repository root. docs/source/conf.py has to compensate for it directly:

# Make the package importable for autodoc (src layout).
sys.path.insert(0, os.path.abspath("../../src"))

That extra line is the src layout’s whole tradeoff in one place. A flat layout, with the package directory sitting next to pyproject.toml, lets Python import test_data_workbench straight out of the working directory, because the interpreter always puts the current directory on sys.path. That is convenient and also exactly the failure mode the src layout closes off: an import that only succeeds because of where a command happens to be run, which can hide the fact that the package was never actually installed, that an entry point is missing, or that a dependency did not get pulled in. Under src/, import test_data_workbench fails from the repository root unless the package has been installed, editable or not, exactly as it would be for anyone else using it. The CI workflows exercise that same path on every run: pip install -e ".[dev]" before pytest, pip install -e . before the Sphinx build.

Type hints without enforcement#

Type hints appear throughout core and adaptation, favoring the explicit typing module spellings, Optional[int], Dict[str, List[Any]], over the newer | union syntax. Most of the codebase sticks to this consistently, but not uniformly: cli/analyze.py declares a local variable as entity_counts: dict[str, int] = {}, the lowercase built-in generic rather than typing.Dict. Several cli modules (main.py, analyze.py, deploy.py, demo.py, team.py, validate.py, _seeding.py) also open with from __future__ import annotations, which postpones evaluation of every annotation in that file to a string; core and adaptation do not use that import. Neither choice changes what runs, since annotations are not evaluated for enforcement either way, but it means the CLI layer and the core layer do not share one typing style.

No tool in this repository checks whether any of these annotations are actually true. .github/workflows/ci.yml runs three jobs, a pytest matrix across Python 3.10 to 3.12 on Linux and Windows, a coverage run, and a Postgres integration test, and none of them invoke mypy, pyright, or any other static checker; there is no [tool.mypy] section in pyproject.toml either. That means two things follow directly. First, the annotations in core and adaptation are documentation and autodoc input, nothing more; they carry the same weight whether they are right or wrong, because nothing runs to tell the difference. Second, when a signature and its default disagree, nothing catches it. DefensiveGeneratorManager.register_generator in core/defensive_manager.py is typed:

def register_generator(self, entity_name: str, generator_class: Type,
                      fallback_class: Optional[Type] = None,
                      entity_type: EntityType = EntityType.UNKNOWN,
                      dependencies: List[str] = None) -> None:

dependencies is annotated List[str] with a default of None, a value no List[str] can ever hold. A type checker running in strict mode would reject this outright; none runs, so it ships. create_scenario_template in core/template_generator.py has the identical shape: entities: List[str] = None. Both work at runtime, because Python never looks at the annotation to begin with, and both rely on a caller either passing a real list or reading the docstring to know that None is actually the accepted default.

Dataclasses versus Pydantic#

core/models.py defines the system’s internal vocabulary, twelve @dataclass types: Column, Relationship, Table, BusinessRule, PIIColumn, SchemaInfo, GeneratorConfig, ScenarioConfig, AdaptationResult, ValidationResult, DemoScenario, and DemoStatus. core/feature_flags.py adds FeatureFlags the same way, and core/defensive_manager.py adds GeneratorResult and GeneratorInfo. None of these are Pydantic models; Pydantic is confined to api/models.py, the one file where untrusted JSON arrives over HTTP (see Pydantic for that side in detail). Internally, between the analyzer, the factory, the defensive manager, and the CLI, data never crosses a trust boundary, so a plain dataclass is the lighter-weight choice for the same amount of type information.

Mutable defaults follow the documented dataclass idiom, field(default_factory=...), consistently. Column.constraints and Column.sample_values both default to an empty list this way; FeatureFlags.enabled_features defaults to a set literal through a default_factory=lambda: {...}. None of these use a bare mutable literal as the default value itself, which is the classic version of this mistake (a default list built once at def-time and then shared, and silently mutated, across every call that does not pass its own).

A few of these dataclasses add an alternate constructor as a @classmethod instead of overriding __init__: BusinessRule.create builds a rule_id from a truncated uuid.uuid4() and fills in the rest, and FeatureFlags.safe_mode returns an instance with only the safest feature set enabled:

@classmethod
def create(cls, description: str, rule_type: str, columns: List[str],
           confidence: float = 0.8) -> 'BusinessRule':
    return cls(
        rule_id=str(uuid.uuid4())[:8],
        description=description,
        rule_type=rule_type,
        affected_columns=columns,
        confidence=confidence
    )

This keeps the generated __init__ as the one place field defaults live, while giving call sites a named, more specific way to build a common case.

Enums for closed vocabularies#

core/models.py defines two Enum classes for values that should never be arbitrary strings: EntityType (eleven members: USER, CUSTOMER, PRODUCT, and so on down to UNKNOWN) and ConstraintType (PRIMARY_KEY, FOREIGN_KEY, UNIQUE, NOT_NULL, CHECK, INDEX). They are used for real membership checks, not just labels; Column.is_primary_key is exactly ConstraintType.PRIMARY_KEY in self.constraints, which only works, and only means what it says, because constraints holds ConstraintType members rather than the strings "primary_key" or "pk" that a careless caller could otherwise substitute.

The same file does not apply this consistently. BusinessRule.rule_type is a bare str with a comment naming its actual closed set, # "format", "range", "dependency", "pattern"; Relationship.relationship_type is the same, # "one_to_one", "one_to_many", "many_to_many"; and PIIColumn.category is a str populated from the ten keys of PII_NAME_PATTERNS in adaptation/schema_analyzer.py ("email", "name", "phone", and so on). Every one of these is a closed, known, comment-documented vocabulary in the same file that defines EntityType for exactly that purpose, and none of them is an Enum. A value of rule_type="pattern-ish" or category="unknown" is accepted the same as a correct one; only the two constraint- and entity-facing fields get the type-safety Enum provides. The parallel case on the API side, plain str fields in api/models.py where a worked example implies a fixed set of choices, is covered in Pydantic.

Async: real concurrency in one place, none in another#

asyncio.gather appears in two places in this codebase, and only one of them overlaps real I/O.

SchemaAnalyzer.analyze_production_schema in adaptation/schema_analyzer.py is async def. It builds one _analyze_table coroutine per discovered table, capped at twenty (table_names[:20]  # Limit for rapid analysis), and runs them together:

tasks = [
    self._analyze_table(engine, inspector, table_name, metadata_only)
    for table_name in table_names[:20]  # Limit for rapid analysis
]
tables = await asyncio.gather(*tasks)

_analyze_table is also async def, and it awaits two helpers, _get_sample_values and _estimate_row_count, both declared async def in turn. Read their bodies and the async stops being real: both connect through the plain synchronous SQLAlchemy Engine created earlier with sa.create_engine(connection_string, pool_pre_ping=True), not sqlalchemy.ext.asyncio.create_async_engine, and both call conn.execute(query) directly.

async def _get_sample_values(self, engine: sa.Engine, table_name: str,
                            column_name: str, limit: int = 10) -> List:
    """Get sample column values for pattern detection."""
    try:
        with engine.connect() as conn:
            query = text(f'SELECT DISTINCT "{column_name}" FROM "{table_name}" LIMIT {limit}')
            result = conn.execute(query)
            return [row[0] for row in result if row[0] is not None]
    except Exception:
        return []

There is no await anywhere inside that function, or inside _estimate_row_count beside it. Both are async only by signature: marked async def so they can be called with await, but nothing in their own bodies ever reaches a genuine suspension point. The Python documentation is direct about the consequence of calling blocking code from inside a coroutine this way, in the “Running Blocking Code” section of Developing with asyncio: a blocking call delays every other concurrent task sharing that event loop, the same way a one-second CPU-bound calculation would. Fanning _analyze_table calls out through asyncio.gather cannot overlap their database queries the way it would if those queries were issued through an async driver and awaited properly; each table’s blocking work runs on the same single thread as everything else the event loop is doing, with no run_in_executor handoff to release it early.

Contrast that with RapidDeployment._write_file_async in adaptation/rapid_deployment.py, which does get real overlap:

try:
    import aiofiles
    async with aiofiles.open(file_path, 'w', encoding='utf-8') as f:
        await f.write(content)
except ImportError:
    with open(file_path, 'w', encoding='utf-8') as f:
        f.write(content)

aiofiles runs the actual file write on a background thread and awaits its completion, the same executor pattern the asyncio documentation recommends for blocking calls, so await f.write(content) is a real suspension point. _write_output_files fans several of these out through await asyncio.gather(*write_tasks, return_exceptions=True), and because each write genuinely yields control while its thread finishes, the writes do overlap. (The except ImportError fallback would itself be another async-only-by-signature call if it ever ran; it should not, since aiofiles>=23.2.1 is a pinned dependency.)

Either way, what asyncio buys here is concurrency for I/O-bound work sharing one event loop and one thread, never parallelism across CPU cores; Python’s own GIL rules that out for ordinary async code regardless of how many tasks gather schedules. The CLI is the synchronous/async boundary in both directions: cli/analyze.py, cli/deploy.py, and cli/demo.py each call into this async code with a single asyncio.run(...), so the rest of the CLI, argument parsing, Rich output, file saving, stays ordinary synchronous Python.

The standard library first#

The CLI entry point, cli/main.py, is built on argparse from the standard library. typer is not a dependency of this project; every subcommand is registered by hand through argparse.ArgumentParser.add_subparsers, one add_subparser(subparsers) function per module (analyze, deploy, validate, team, demo), each returning the argparse.ArgumentParser it configured:

def build_parser() -> argparse.ArgumentParser:
    # Imported lazily so that `tdw --version` stays fast and side-effect free.
    from . import analyze, deploy, demo, team, validate

    parser = argparse.ArgumentParser(
        prog=PROG_NAME,
        description="Test Data Workbench - schema-adaptive test data generator. "
        "Point it at a database, get realistic relational test data back.",
    )
    ...
    subparsers = parser.add_subparsers(dest="command", metavar="command")
    analyze.add_subparser(subparsers)
    ...

Each subcommand module wires its own arguments the same way; cli/analyze.py is representative, adding a positional connection, a --format choice among table/json/summary, an --output path, and a --metadata-only flag, then binding the subcommand to a handler with parser.set_defaults(func=_run). Subcommand modules are imported inside build_parser() rather than at module load time, so that tdw --version stays fast and does not pay the import cost, and pull in the dependencies, of every subcommand it is not running. Rich (see Rich) sits on top of this for output, tables, panels, and progress spinners, but it plays no part in defining or parsing arguments; that is argparse end to end. It mirrors the project’s broader dependency discipline described in Stack overview: reach for the standard library before adding a dependency, and a CLI parser is squarely inside what argparse already does well.

Version floor: 3.10, and what depends on it#

pyproject.toml declares requires-python = ">=3.10", with classifiers for 3.10, 3.11, and 3.12, and the CI matrix in Testing and Quality runs all three, on both Ubuntu and Windows.

Nothing in this codebase’s own source actually requires 3.10. Two syntax features shipped in that release: the match/case statement (PEP 634) and the X | Y union syntax in annotations (PEP 604). Neither appears anywhere under src/; there is no match statement in the project, and every union type is written the older way, Optional[int], Union[str, int]. Nothing else in 3.10 changes what syntax is legal either; the release’s other headline items, better SyntaxError and NameError messages (“did you forget parentheses around the comprehension target?”, “did you mean: …”), and parenthesized with statements spanning multiple lines, are interpreter and grammar conveniences a codebase benefits from just by running on 3.10, not features its own code has to opt into. So the floor is a deployment-target choice more than a language-feature requirement: 3.10 is old enough to be widely available and new enough to still receive attention, and the project simply has not needed anything past it.

Sharp edges and limits#

No static type checker runs anywhere in this project. There is no mypy, pyright, or ruff step in .github/workflows/ci.yml, and no [tool.mypy] or equivalent configuration in pyproject.toml. The only things that ever read these annotations are sphinx.ext.autodoc, for documentation, and Pydantic, at the one boundary described above. Everywhere else, “the code is type-hinted” describes intent and readability, not a verified property; register_generator’s dependencies: List[str] = None and create_scenario_template’s entities: List[str] = None are both live examples of an annotation and its own default disagreeing, in code that ships and runs correctly regardless.

Closed vocabularies are only sometimes closed. EntityType and ConstraintType are real Enum classes; BusinessRule.rule_type, Relationship.relationship_type, and PIIColumn.category are bare str fields with the same kind of fixed, comment-documented set of values, in the same core/models.py file, and accept anything.

Dataclasses carry no construction-time validation of their own. Column(name=123, data_type=None) builds without complaint; a dataclass’s generated __init__ assigns whatever it is given and never consults the annotations, the same erasure that motivates Pydantic at the API boundary in the first place (see Pydantic) applies with equal force here, just without Pydantic sitting in front of it to catch anything.

Several fields are typed Dict[str, Any] and stop there: SchemaInfo.analysis_metadata, GeneratorConfig.generation_rules, ScenarioConfig.constraints, DemoScenario.demo_data. The annotation confirms the outer container and says nothing about what is inside it; a reader has to go find where each one is populated to know its actual shape.

Broad except Exception handlers are the project’s error-handling strategy in both adaptation/schema_analyzer.py and core/defensive_manager.py, and in the latter it is the stated design, not an oversight; the module docstring calls it “defensive generator management with automatic fallbacks and error isolation.” The cost of that choice is real: a genuine bug and an expected failure both get caught by the same except Exception: and degrade the same way, whether that is _analyze_table returning None for a table, _get_sample_values returning an empty list, or analyze_production_schema’s own top-level handler replacing the whole result with an empty, database_name="error_fallback" SchemaInfo. The only trace left behind is a logged warning or an entry in an errors: List[str] field that a caller has to choose to read; the type of the failure is not distinguishable from the outside.

One specific method inside that same fallback machinery does not do what its name promises. DefensiveGeneratorManager._with_timeout takes a timeout_seconds: int = 30 parameter, and generate_safely calls it with timeout_seconds=30, but the body is:

def _with_timeout(self, func: Callable, timeout_seconds: int = 30) -> Any:
    """Execute function with timeout protection."""
    # Simple timeout implementation - in production would use more sophisticated approach
    try:
        return func()
    except Exception as e:
        # For now, just re-raise - timeout handling would be added here
        raise e

There is no timer, thread, or signal anywhere in it; timeout_seconds is accepted and never read. The comments are honest about this being a stub, but a caller reading only the signature would reasonably expect a hung generator to be cut off at thirty seconds. It will not be.

And as covered above, the asyncio.gather fan-out in analyze_production_schema does not overlap the per-table database work it schedules; the helpers it awaits never yield control themselves, so the concurrency is present in the code’s shape without being present in its behavior.

Learn more resources#

Official documentation

  • PEP 483: “The Theory of Type Hints,” the gradual-typing theory this page’s “idea underneath” section names.

  • PEP 484: “Type Hints,” the specification that made annotations erased-at-runtime by design, and the same PEP Pydantic cites for the boundary Pydantic enforces instead.

  • src layout vs flat layout: the Python Packaging User Guide’s own explanation of the failure mode the src layout closes off.

  • dataclasses: the standard-library module behind every internal model in core/models.py.

  • enum: the module behind EntityType and ConstraintType.

  • argparse: the module cli/main.py is built on.

  • asyncio - Coroutines and Tasks: documents asyncio.gather, used in both schema_analyzer.py and rapid_deployment.py.

  • Developing with asyncio: the “Running Blocking Code” section documents exactly the pitfall _get_sample_values and _estimate_row_count run into, and the run_in_executor escape hatch they do not use.

  • PEP 604: the X | Y union syntax that shipped in 3.10 and that this codebase does not use anywhere.

  • PEP 634: the match/case statement specification, the other 3.10 syntax feature absent from this codebase.

  • What’s New In Python 3.10: the full list of what the version floor actually adds, syntax and otherwise.

Tutorials and blogs

Videos