Jinja2#

Why Jinja2#

Test Data Workbench’s entire value proposition, laid out in Design Principles, is that it hands back readable, ownable Python source for each table, not rows and not a configuration DSL. Something has to turn a table’s shape (its columns, types, keys, and relationships) into working Python text, correctly, for however many tables a connected database has, in several structurally different shapes: a specialized template for recognized entity types, a generic template for unrecognized ones, a minimal always-works fallback, and three separate skill-level starting points for human contributors. Building that many code shapes out of manual string concatenation or nested f-strings would mean re-deriving loop and conditional logic in string form by hand, every time. A templating engine that already understands loops, conditionals, and variable substitution over text is the obvious fit, and Jinja2 is a mature, pure-Python, dependency-light choice for it (pyproject.toml pins jinja2>=3.1.2), consistent with the short dependency list described in Stack overview.

The idea underneath#

A template engine’s job is to separate the fixed structure of an output document from the variable data poured into it at render time. That separation is usually described in a web context (a page skeleton plus the values for one request), but the underlying idea does not care what kind of text comes out the other end. Jinja2’s own documentation is explicit that it is a general text-templating engine, not an HTML-only one.

What this project does with that general idea is code generation: a program that writes a program. That is a real, named tradeoff, not a free lunch. It buys exactly what Design Principles calls inspectability, correctability, and ownership: the output is an ordinary .py file a person can read, fix, and keep, with no hidden runtime model standing behind it. It costs synchronization: a rendered generator reflects the schema at the moment tdw deploy ran and nothing after. Add a column to the live database tomorrow and the generator file on disk does not know until someone reruns generation, and rerunning overwrites whatever hand edits were made in the meantime. The generated artifact is a snapshot, not a live view.

The alternative worth naming is building the output from an AST module instead of text: constructing ast.FunctionDef and ast.Assign nodes programmatically and unparsing or compiling them. That approach can never produce syntactically invalid Python, because the tree is built from a typed grammar rather than assembled character by character. But the “template” in that world is Python code that builds Python objects; there is no separate, readable shape of the target class to glance at, only code you would have to run to see what it produces. String templating is a reasonable choice here specifically because the project’s own stated bar is that the template, not only its output, stays close to ordinary readable Python. A Jinja2 template that is mostly literal Python text with a scattering of {{ }} and {% %} tags is something a contributor can read top to bottom and recognize as “the class, with some blanks.” An AST-builder version of the same generator would read as tree-construction code with no visual resemblance to the class it produces.

How it fits this project#

Two modules do all the rendering. src/test_data_workbench/adaptation/generator_factory.py’s GeneratorFactory renders the actual per-table generator source: for each Table discovered by SQLAlchemy reflection, it picks a specialized template (user, product, order, review) or falls back to a generic one, computes a small context dict in plain Python (class name, per-column generator expressions, a handful of booleans describing what helper methods the class needs), and calls .render() to get finished Python text. src/test_data_workbench/core/template_generator.py’s TeamTemplateGenerator uses the same mechanism for a different audience: copy-paste-and-edit starting points for team contributors at three skill levels, a scenario YAML template, and a generated project README. tdw deploy <connection> (cli/deploy.py, via RapidAdapter) is the command that ties this to a real database and writes the rendered files to disk.

Inline Template, not a configured environment#

Neither module builds a Jinja2 Environment with a FileSystemLoader pointed at template files on disk. Both import jinja2.Template directly and construct it from an inline triple-quoted string, right next to the Python that computes what goes into it:

from jinja2 import Template

template = Template('''
from faker import Faker
...
class {{ class_name }}Generator:
    ...
''')
return template.render(class_name=..., columns=..., ...)

Every generator-producing method in GeneratorFactory (_create_user_generator, _create_product_generator, _create_order_generator, _create_review_generator, _create_generic_generator, _create_fallback_generator) follows this shape. Keeping the template string and the Python that feeds it in the same method keeps each generation strategy self-contained and easy to find, at the cost of losing syntax highlighting on the embedded template text and of Jinja2 re-parsing the string on every call rather than compiling it once at import time. For templates this small, rendered once per table per tdw deploy run, that parsing cost is not something the project has needed to optimize away.

No autoescaping, and rightly so#

Nothing in this codebase sets autoescape. There is no Environment(autoescape=True), no select_autoescape(), and every Template(...) call in generator_factory.py and template_generator.py uses the plain constructor with no autoescape argument at all. Jinja2’s own default for that constructor is autoescape off; the project gets that behavior by never touching the setting, not by deliberately disabling something.

That default is correct here, and turning it on would be an active bug. Autoescaping exists to stop a value like <script> or a stray &, ", or ' from being interpreted as markup once it lands in a browser, by rewriting those characters into HTML entities (' becomes &#39;, < becomes &lt;, and so on). None of that transformation means anything to the Python interpreter that will eventually run this project’s output. An &#39; inside a Python string literal is not an apostrophe to the Python parser, it is four literal characters. If autoescape were on, a rendered '{{ column.name }}' whose column happened to contain an apostrophe, or an {{ entity_type }} value rendered anywhere in the output, would come back with HTML-entity noise baked into what needs to be valid Python. HTML escaping does not make code generation safer; it corrupts it, because the two output formats have entirely different ideas of what character needs protecting from what.

Whitespace control shapes the emitted Python#

In HTML, the whitespace {% if %}/{% endif %} tags leave behind is cosmetic: a browser collapses it. In generated Python, indentation is the block structure itself, so the same whitespace is load-bearing. These templates use Jinja2’s trim markers ({%- and -%}) throughout to control exactly what whitespace survives into the output.

The simplest pattern is an optional line inside a fixed block. In _create_user_generator, the sequential primary-key counter only appears when has_sequential_pk is true, and it is written so that when it is false, no blank line or stray indentation is left behind:

    self.fake = Faker()
    self._generated_emails = set()
{%- if has_sequential_pk %}
    self._pk_counter = 0
{%- endif %}
    if seed is not None:

A subtler pattern shows up in template_generator.py’s per-column generator bodies, where the same return statement has to land at a fixed eight-space indent regardless of which branch of an if/elif/else chain fires:

def _generate_{{ column.name }}(self, index: int):
    """Generate value for {{ column.name }} column."""
    {% if 'id' in column.name -%}
    return index + 1
    {%- elif 'name' in column.name -%}
    return self.fake.name()
    {%- elif 'email' in column.name -%}
    return self.fake.email()
    {%- else -%}
    return self.fake.text(max_nb_chars=50)  # Placeholder
    {%- endif %}

Here the opening {% if %} deliberately keeps its leading whitespace (no - on the left), so the indentation already sitting in the template source before the tag becomes the indentation of whichever return line actually fires. The trailing -%} and leading {%- on every other tag strip the newlines the tags would otherwise leave, so the if/elif/else lines themselves contribute nothing to the output; only the chosen return statement, at its already-correct indent, survives. Getting this wrong does not just look untidy, since Design Principles treats the generated file as something a person reads and edits, uneven or misindented output would directly undermine the property the whole generation strategy is built on.

Variables passed into the templates#

Every .render() call passes a small, purpose-built context computed in plain Python immediately beforehand, never the Table/Column schema objects wholesale. GeneratorFactory typically builds a columns list of plain dicts, one per column, each with a name and an already-decided generator expression string:

        for _ in range(count):
            record = {
{%- for column in columns %}
                '{{ column.name }}': {{ column.generator }},
{%- endfor %}
            }

column.generator is not template logic; it is a plain Python string such as "self.fake.email()" or "self._get_reference_id('categories')", decided entirely in Python (_get_column_generator, _infer_reference_table, _fk_relationships_for_table) before render time and interpolated into the output verbatim. The template performs no dispatch of its own here; it substitutes a value someone else already chose.

core/template_generator.py’s templates instead pass the real Column dataclass objects through and address their fields the same way, column.data_type, column.nullable, column.is_foreign_key. That works because Jinja2’s . operator tries attribute access first and falls back to item lookup: column.name resolves identically whether column is a dict with a "name" key or a dataclass instance with a name attribute. The two modules lean on opposite ends of that convenience without either one being wrong. Other context values vary by call site: class_name and entity_type (single strings), has_sequential_pk, uses_reference_helper, and uses_self_reference_helper (booleans computed from the schema), and, for the scenario and README templates, business_case, entities, default_counts, and tables.

Conditional blocks that change the generated class shape#

The generic-generator template (_create_generic_generator) is where conditionals stop being decoration and start doing real structural work. Whether the generated class’s __init__ accepts a reference_data parameter at all, and which helper methods it defines, is decided by booleans computed in Python and threaded into the template:

    def __init__(self, {% if has_relations %}reference_data: Dict[str, List[Any]] = None, {% endif %}seed: Optional[int] = None):
        self.fake = Faker()
{%- if has_relations %}
        self.reference_data = reference_data or {}
{%- endif %}
{%- if uses_self_reference_helper %}
        self._self_generated: Dict[str, list] = {}
{%- endif %}

A table with no foreign keys gets a generator with a plain, no-argument constructor. A table with an ordinary foreign key to another table gets a reference_data parameter and a _get_reference_id helper. A table with a self-referential foreign key (the method’s own docstring gives categories.parent_id -> categories.id as the example) additionally gets a _self_generated cache and a _self_reference_id helper that picks an already-generated value from earlier in the same batch, or None for the first rows, rather than pointing at a row that does not exist yet. The same {% if columns %} ... {% else %} ... {% endif %} shape appears in template_generator.py’s beginner template, which degrades to a plain id/name/created_at record shape when no Table was passed in at all. In both cases the template is choosing between materially different class shapes, not filling blanks in one fixed skeleton.

Where the templates actually live#

There are no .j2 or .jinja files anywhere in this repository, and no FileSystemLoader. Every Jinja2 template is an inline Python string literal, and there are exactly two places to find them: core/template_generator.py and adaptation/generator_factory.py.

The directory name generators/templates/ is a trap for anyone searching for “the templates”: it holds three files, basic_entity.py, relationship_aware.py, and adaptive_generator.py, and none of them contain any Jinja2 syntax at all. They are complete, directly runnable Python scripts, and their own module docstring says so plainly: “these files are reference examples for hackathon team members to copy and modify; they are standalone scripts… and are not imported by any code in this package.” They are a different artifact for a different audience, human contributors looking for a copy-paste starting point, not templates the tool renders.

Sharp edges and limits#

Templates here are project-authored, not user-supplied, and that is what makes the setup safe. Every Template(...) call renders a string written by this project’s own developers. The usual Jinja2 threat model, where a hostile or careless template author could use Jinja’s Python-like expression syntax to reach unsafe attributes or methods, does not apply, because nobody outside the project ever gets to write a template. That assumption would break the moment a template stopped being fixed project code, for instance if a future feature let a user supply their own scenario template or contribution scaffold. Jinja2 ships jinja2.sandbox.SandboxedEnvironment for exactly that situation, and nothing in this codebase uses it today, because nothing in this codebase needs it today.

The data flowing into templates is a different story: some of it comes from a real, external database. SchemaAnalyzer._analyze_table (see SQLAlchemy) builds every Column directly from inspector.get_columns(table_name) against whatever connection string tdw deploy was pointed at. Column.name is a bare str field with no validation, and nothing between reflection and .render() sanitizes it. In generator_factory.py that name only ever appears quoted, as a dict key or a string argument ('{{ column.name }}': ...), so the risk there is narrow: a column name containing a single quote would break the Python string literal it lands in, but nothing worse. In template_generator.py the same name is used unquoted, as part of an actual method identifier, def _generate_{{ column.name }}(self, ...):. A column named with a hyphen, a space, or a leading digit, all of which are legal in more than one real database’s column-naming rules, would render a syntactically invalid Python method definition. Autoescaping would not have prevented this either way; it targets HTML-unsafe characters, not Python-unsafe ones.

Nothing checks that the rendered text is valid Python before returning it. Jinja2 is a text engine; it has no concept of Python syntax and .render() will happily hand back a broken method definition. The project’s automated safety net is downstream and real, not aspirational: tests/test_generator_factory.py::test_generated_code_compiles runs compile(code, ..., "exec") on every rendered generator for the test fixtures, and tests/test_generated_execution.py goes further, executing the generated source and actually calling .generate() on it for every template kind (user, customer, product, order, review, generic, and fallback), asserting on the resulting records. That file’s own docstring describes a real bug this caught: a type-based fallback that emitted a bare fake.xxx() call which only worked in the one template that happened to define a module-level fake, and raised NameError in every other template the moment generate() actually ran, a defect compile()-only checking could never have caught. At deploy time, adaptation/constraint_check.py’s generate_sample_dataset execs rendered generator code again to build a sample dataset for its constraint report, but that path is explicitly best-effort: a table whose source fails to exec is silently skipped rather than surfaced as a failure, by design, so that one bad generator does not take down the rest of a deploy’s reporting. Separately, core/validators.py’s GeneratorValidator does run ast.parse and a set of quality-gate checks, but it is wired to the tdw validate CLI command for reviewing team-submitted contribution files, not into the schema-driven generation path itself.

Logic embedded in templates has a real maintainability cost. The generic-generator template folds three independent booleans (has_relations, uses_reference_helper, uses_self_reference_helper) into overlapping conditional blocks that each change part of the constructor signature or the method list. Reading what a specific table’s generator will actually look like means mentally executing the template against that table’s schema rather than reading a fixed class body; the method’s own docstring exists in prose largely because the branching is not obvious from the template text alone.

Learn more resources#

Official documentation

  • Template Designer Documentation: the full reference for whitespace control ({%-/-%}), variable and attribute access, and the conditional and loop syntax these templates use throughout.

  • API: documents the Template and Environment classes, including the autoescape default this project relies on simply by never overriding it.

  • Sandbox: describes SandboxedEnvironment, the mechanism this project would need if it ever rendered a template written by someone other than the project itself.

Tutorials and blogs

Videos

  • Jinja2 Templating Engine Tutorial (Mr. Rigden): covers loading templates, variable substitution, control structures, and template inheritance, the general mechanics this project’s inline templates build on.