Faker#

Faker is the value engine: whatever ends up in a generated column, in the overwhelming majority of cases, came out of a Faker provider method. It shows up in this project in two genuinely different roles, and telling them apart is the point of this chapter.

Why Faker#

A schema-adaptive generator meets tables it has never seen and has to produce column values for them without a human writing per-field fixtures. Hand-rolled randomness (random.randint, random.choice over a short list) can fill a column with something type-valid, but it cannot produce a string that reads as a name, an email address, or a street. Faker exists to close exactly that gap: a library of typed data providers whose method names already line up with the kinds of fields a real schema tends to have (email, name, address, phone_number, and so on).

pyproject.toml pins faker>=20.1.0 as a plain runtime dependency, listed alongside sqlalchemy and code-generation in the project’s own keywords. That placement is a signal worth noting: Faker is not a development-only convenience the tool uses internally and then discards. As the next section explains, most of the project’s use of Faker ships in the code it hands back to you.

The idea underneath#

Faker does not invent anything from entropy. Every Faker() instance holds a random.Random object internally and calls it to pick which template, which name, which number to return, per Faker’s own description of the Faker class. Python’s random module, in turn, is explicit that it is a pseudo-random number generator: deterministic, not truly random. Its documentation names the algorithm directly: “Python uses the Mersenne Twister as the core generator,” and warns plainly that this makes it “completely unsuitable for cryptographic purposes.” A dedicated warning box in the same page tells readers who need real security to reach for the secrets module instead.

That distinction matters more for test fixtures than it first looks. Because the sequence is deterministic given a starting state, seeding the generator (Faker.seed(n)) fixes that starting state, and every subsequent call becomes reproducible: two runs with the same seed produce byte-identical output. That is precisely what a test suite needs. A fixture that changes between runs turns a reproducible failure into a flaky one; a fixture that reproduces means a CI failure can be recreated exactly on a developer’s own machine by passing the same seed. “Random” here means “replayable,” not “unpredictable,” and that is a feature, not a limitation.

The same determinism is also the reason Faker (and the random module it sits on) must never generate anything security-sensitive: a token, a password, an API key. If the seed or internal state is ever knowable or guessable, every value the generator has produced or will produce is knowable too. Faker is a plausibility engine, not an entropy source.

Underneath the Faker object, data is organized into providers: classes that each own one topical slice of fake data (people, addresses, the internet, lorem text) and the methods that generate it. Locale-specific data lives underneath the same provider interface, so a locale-aware provider can back the same method name, address() for instance, with a different underlying data table per locale, without the calling code changing at all. This project’s use of that architecture, or rather its non-use of the locale half of it, is covered in Locale: none configured below.

How it fits this project#

Faker plays two distinct roles here, and only one of them is Faker being called by this project’s own code. src/test_data_workbench/cli/demo.py instantiates a Faker() and calls it directly to seed the bundled SQLite demo database. Everywhere else that matters more, in src/test_data_workbench/adaptation/generator_factory.py and src/test_data_workbench/core/template_generator.py, this project does not call Faker at all. Those modules write Python source text that calls Faker, through Jinja2 templates (see Jinja2). The generator classes that result, and that tdw deploy writes to disk, contain their own from faker import Faker and their own calls into it; Faker is a runtime dependency of the code you receive, exactly as the design decision described in Design Principles (“generated code is the product”) would predict.

Two call sites#

Faker is called from two genuinely different places in this codebase: the tool’s own code, directly, and the generator code the tool writes and hands back to you.

Called directly, at generation time#

build_demo_database in cli/demo.py is the one place in this codebase where the tool itself runs Faker to produce values, rather than writing code that will. It seeds once via the shared seed_generation helper (see Seeding and reproducibility: two seed calls, not one), creates a single Faker() instance, and passes it into a handful of small insert helpers:

conn.execute(
    "INSERT INTO users (id, username, email, first_name, last_name, created_at) "
    "VALUES (?, ?, ?, ?, ?, ?)",
    (
        i,
        fake.user_name(),
        fake.unique.email(),
        fake.first_name(),
        fake.last_name(),
        fake.date_time_this_year().isoformat(),
    ),
)

This is the tdw demo zero-setup path: it exists purely to build a small, plausible-looking SQLite database for someone to point the rest of the tool at without a database of their own. It is a legitimate but minor role: five insert helpers, one shared Faker instance, done.

Called from the code it writes#

The role that matters is the opposite direction. generator_factory.py’s _create_generic_generator (and its four sibling entity-specific methods) builds a Jinja2 Template whose rendered output is a complete Python module:

from faker import Faker
from typing import Dict, Any, Optional, List
import random

class {{ class_name }}Generator:
    """Generated generic data generator."""

    def __init__(self, {% if has_relations %}reference_data: Dict[str, List[Any]] = None, {% endif %}seed: Optional[int] = None):
        self.fake = Faker()
{%- if has_relations %}
        self.reference_data = reference_data or {}
{%- endif %}
{%- if has_sequential_pk %}
        self._pk_counter = 0
{%- endif %}
{%- if uses_self_reference_helper %}
        self._self_generated: Dict[str, list] = {}
{%- endif %}
        if seed is not None:
            Faker.seed(seed)
            random.seed(seed)

The {{ }} and {% %} markers are Jinja2 syntax, not Python; once rendered for a real table, this becomes an ordinary .py file that imports faker on its own account. GeneratorFactory (the tool) never imports or calls Faker to produce a value; it only ever emits text that, once written to disk and run, does. The same is true of core/template_generator.py’s TeamTemplateGenerator, which renders the three skill-level contribution templates described in Jinja2; every one of them also opens with from faker import Faker. Treat every self.fake.something() you read in this chapter as a description of code the tool writes, unless the surrounding text says otherwise.

Provider selection per column: names before types#

generator_factory.py decides which Faker method a column gets through two layers, checked in order, in _get_column_generator and duplicated inline in the entity-specific generator methods. Column-name patterns are checked first, regardless of the column’s declared SQL type:

if 'email' in col_name_lower:
    return 'self.fake.email()'
elif 'phone' in col_name_lower:
    return 'self.fake.phone_number()'
elif 'name' in col_name_lower:
    return 'self.fake.name()'
elif 'address' in col_name_lower:
    return 'self.fake.address()'
elif 'city' in col_name_lower:
    return 'self.fake.city()'
elif 'country' in col_name_lower:
    return 'self.fake.country()'
elif 'url' in col_name_lower:
    return 'self.fake.url()'

The declared type is the fallback, via a fixed dictionary:

self.faker_mapping = {
    'varchar': 'self.fake.text',
    'text': 'self.fake.text',
    'integer': 'self.fake.random_int',
    'bigint': 'self.fake.random_int',
    'decimal': 'self.fake.pydecimal',
    'numeric': 'self.fake.pydecimal',
    'boolean': 'self.fake.boolean',
    'date': 'self.fake.date',
    'timestamp': 'self.fake.date_time',
    'timestamptz': 'self.fake.date_time_this_year',
    'uuid': 'self.fake.uuid4'
}

A generic varchar/text type is deferred to name-based temporal detection (_temporal_kind) before this mapping is trusted, because SQLite in particular stores timestamps with TEXT affinity, and a declared type of “text” alone cannot tell the factory whether a column is really a date. If neither layer matches anything, the column falls all the way through to self.fake.text(max_nb_chars=50), the same “always produces something” floor described in The Fallback Ladder, applied one column at a time rather than one table at a time.

Primary keys: a counter or a uuid, and neither one is random#

Primary keys are the one case where the type-based mapping above is deliberately overridden before it ever runs. An integer-typed primary key does not get self.fake.random_int() from the table above; it gets a contiguous, instance-level counter instead:

def _needs_sequential_pk(self, column: Column) -> bool:
    """True if this column is an integer-typed primary key and should
    therefore get a contiguous sequence (1, 2, 3, ...) instead of a
    random or uuid value. Non-integer primary keys (uuid, string,
    etc.) are left to their existing type/name-based strategy."""
    return column.is_primary_key and self._is_integer_column(column)

A real auto-increment primary key is contiguous across the whole table, and a randomly-drawn integer is not, so this is a case where the correct choice is to bypass Faker’s randomness entirely for the sake of realism. A non-integer primary key (typically uuid) keeps a different, still-not-random-looking strategy: self.fake.uuid4(), unconditionally, whenever the column’s own name contains id. Both are exceptions to “let the type or name pick a provider” carved out specifically because a primary key’s shape matters more than its plausibility as fake data.

Entity-specialized calls#

The four entity-specific generator methods (user, product, order, review) layer domain-appropriate Faker calls on top of the generic mapping, choices about which provider reads right for a field rather than which one merely type-checks: self.fake.catch_phrase() for a product name or title (a punchy business-speak phrase reads more like a product name than lorem ipsum does), self.fake.text(max_nb_chars=200) for a product description, and self.fake.text(max_nb_chars=300) for a review comment. A review’s numeric rating is not Faker-derived at all; it is generated once per record with plain random.randint(1, 5) and reused for both the rating column and any validation that checks it stayed in range.

Seeding and reproducibility: two seed calls, not one#

Both tdw deploy and tdw demo accept an optional --seed. It flows through one shared helper, called once, before any schema analysis or generation begins:

def seed_generation(seed: Optional[int]) -> None:
    """Seed Python's random module and Faker for reproducible output.

    A no-op when ``seed`` is ``None`` (the default, unseeded behavior).
    """
    if seed is None:
        return
    random.seed(seed)
    Faker.seed(seed)

Every generated generator’s __init__ repeats the same pair of calls:

if seed is not None:
    Faker.seed(seed)
    random.seed(seed)

Both calls are necessary, and it is not redundancy. Faker.seed() is a classmethod: it reseeds Faker’s shared random source process-wide, which is what makes every Faker() instantiated afterward, in every generator, draw from the same seeded state rather than an independently-seeded one. But not every random choice in the generated code goes through Faker at all. The product generator’s price, for instance, is plain standard-library random, not a Faker provider:

gen = f"round(random.uniform(10, 1000), 2)"

and category selection is random.choice(self.categories) over a hardcoded list, again bypassing Faker entirely. Seeding only Faker would leave every one of these random.* calls, and the self-referential foreign-key selection logic in the generic template (which calls random.random() and random.choice() directly), unseeded and therefore non-reproducible. Both seed calls exist because both random sources are genuinely in use, and _seeding.py’s module docstring says as much: seeding happens once, at CLI entry, “before any schema analysis, database population, or code generation happens,” specifically so that no call has already consumed random state before the seed lands.

Uniqueness: two different strategies, and one real gap#

Faker does not guarantee unique output by default; two calls to fake.email() can return the same address, especially as a batch grows. This project handles that in two different ways depending on which of the two call sites (above) is doing the generating.

In the code the tool writes, the user and customer templates track what they have already produced themselves, in a plain instance-level set, rather than relying on Faker’s own uniqueness machinery:

def _unique_email(self) -> str:
    """Generate unique email address."""
    while True:
        email = self.fake.email()
        if email not in self._generated_emails:
            self._generated_emails.add(email)
            return email

This is a deliberate readability choice, not an oversight: Design Principles treats generated code as something a reader unfamiliar with Faker has to be able to follow and edit, and a loop-and-retry against a local set reads more plainly than the semantics of a stateful proxy object.

In the code the tool runs itself (build_demo_database, see Two call sites above), there is no such concern about an unfamiliar reader, and the demo seeder reaches for Faker’s built-in uniqueness proxy directly instead: fake.unique.email(). Per Faker’s own documentation, .unique guarantees a returned value has not been produced before for the lifetime of that Faker instance, and raises UniquenessException if it runs out of room to keep that promise.

The real gap is that neither mechanism is driven by the schema’s actual unique constraints. core/models.py defines ConstraintType.UNIQUE as a value schema metadata can carry, but the schema analyzer never populates it (there is no call to SQLAlchemy’s get_unique_constraints anywhere in this codebase), and generator_factory.py never checks it either. The dedup handling above is triggered purely by a column being named something containing email; a schema with a real unique constraint on, say, a username or a sku column gets no deduplication at all, and the two values can collide the same way any unseeded Faker output can. Worth noting: the team-facing config.yaml that ConfigurationBuilder writes reports 'unique_constraints': True unconditionally, for every schema, regardless of whether any dedup logic actually applies to that schema’s unique columns.

Locale: none configured#

Nowhere in this codebase does a Faker() call pass a locale. Every instantiation, in the generated-user, generated-product, generated-order, generated-review, and generic templates in generator_factory.py, in the three skill-level templates in template_generator.py, in the demo-database seeder in cli/demo.py, and in the three static example templates under generators/templates/, is a bare Faker(). That means every value this project produces, across every table and every generator, comes from Faker’s default locale (en_US). If a project needed locale-appropriate data, for instance Dutch-shaped addresses and phone numbers for a Netherlands-facing schema, nothing here currently selects one; per the “correctability” principle in Design Principles, that would be a one-line edit to the emitted Faker() call in the generated file, not a configuration option this tool currently exposes.

A custom provider that is defined but never registered#

ConfigurationBuilder.generate_faker_providers builds a custom Faker provider, structurally correct Faker usage that this project’s generated code never actually gets to use:

def generate_faker_providers(self, entity_types: Dict[str, EntityType]) -> List[BaseProvider]:
    """Create custom Faker providers matching schema patterns."""
    providers = []

    # E-commerce specific provider
    class EcommerceProvider(BaseProvider):
        """Custom provider for e-commerce data."""

        def product_category(self):
            ...

        def order_status(self):
            ...

        def product_sku(self):
            ...

        def price(self, min_price=5.0, max_price=500.0):
            ...

        def customer_tier(self):
            ...

    providers.append(EcommerceProvider)

    return providers

EcommerceProvider subclasses Faker’s own faker.providers.BaseProvider correctly, and its five methods (bodies omitted above for length) are reasonable domain-specific generators. But nothing in this codebase ever calls some_faker_instance.add_provider(EcommerceProvider), the step that would actually attach it to a Faker object and make fake.product_sku() callable. No generated generator template references it, and rapid_deployment.py never wires its return value into anything downstream. The one test that touches it, test_faker_providers_returned, only asserts len(providers) >= 1, which passes whether or not the provider is ever registered anywhere. Product categories and prices in the actual generated product template come from a hardcoded self.categories list and random.uniform, not from this provider. As written, EcommerceProvider is real, correct, tested-to-exist Faker provider code that nothing in the system currently uses.

Sharp edges and limits#

Plausible is not real. Every value Faker produces is shaped like real data and is not real data. A self.fake.email() result is not deliverable; self.fake.phone_number() does not reach anyone; self.fake.address() will not geocode. Faker’s provider set includes fields with even sharper failure modes if this project ever reached for them (an IBAN provider exists in Faker itself, unused here, and its output looks exactly like a real IBAN without being tied to any real bank). Nothing generated by this project, or by Faker generally, should be fed into anything that validates real-world identity, address, or payment data.

Uniqueness is name-driven, not constraint-driven. As covered above, only columns whose name contains email get any deduplication in generated code. A genuinely unique-constrained column with a different name gets none, and Faker does not enforce uniqueness on its own; collisions in a large enough batch are a real, currently unguarded possibility.

The username trap. A column literally named username does not get a username. The name-pattern check in _get_column_generator and in the inline checks in _create_user_generator is a plain substring test: 'name' in col_name_lower. The word “username” contains “name” as a substring, and the email/phone checks are evaluated first but do not match it, so a column named username is routed to self.fake.name(), a full display name such as “Jordan Alvarez,” rather than anything shaped like a login handle. This is present in the current code, not a historical bug; a schema with a real username column will get full names in that column.

Locale coverage is unexercised, not unsupported. Faker itself supports a large number of locales, but as covered above, this project always uses the default (en_US). Anyone relying on generated data to look correct for a non-US locale needs to know that today it will not, without a manual edit to the emitted code.

Never for secrets. Faker sits on a Mersenne Twister PRNG (see The idea underneath), the same generator Python’s own documentation calls “completely unsuitable for cryptographic purposes.” Nothing in this codebase currently misuses Faker this way, but a self.fake.uuid4() or a Faker-derived string can look enough like a token or a credential that the temptation exists. It should never be used for one.

Learn more resources#

Official documentation#

  • Welcome to Faker’s documentation!: the project’s own overview, including how the shared random.Random instance and the seed() method work.

  • Using the Faker Class: covers Faker.seed() as a classmethod, seed_instance() for per-instance seeding, and the .unique proxy plus UniquenessException.

  • Standard Providers: the full list of built-in providers, useful when picking a method by hand instead of relying on this project’s name/type mapping.

  • The Python random module: Python’s own documentation for the module Faker is built on, including the Mersenne Twister attribution and the explicit warning that it is not suitable for cryptographic or security use.

Tutorials and blogs#

  • Python Faker Library (GeeksforGeeks): a practical walkthrough of installation, common provider methods, locale selection, and seeding, useful as a broader tour than this chapter’s project-specific focus.

Videos#