Faker#
Faker is the value engine: whatever ends up in a generated column, in the overwhelming majority of cases, came out of a Faker provider method. It shows up in this project in two genuinely different roles, and telling them apart is the point of this chapter.
Why Faker#
A schema-adaptive generator meets tables it has never seen and has to produce
column values for them without a human writing per-field fixtures. Hand-rolled
randomness (random.randint, random.choice over a short list) can fill a
column with something type-valid, but it cannot produce a string that reads as
a name, an email address, or a street. Faker exists to close exactly that gap:
a library of typed data providers whose method names already line up with the
kinds of fields a real schema tends to have (email, name, address,
phone_number, and so on).
pyproject.toml pins faker>=20.1.0 as a plain runtime dependency, listed
alongside sqlalchemy and code-generation in the project’s own
keywords. That placement is a signal worth noting: Faker is not a
development-only convenience the tool uses internally and then discards. As
the next section explains, most of the project’s use of Faker ships in the
code it hands back to you.
The idea underneath#
Faker does not invent anything from entropy. Every Faker() instance holds
a random.Random object internally and calls it to pick which template,
which name, which number to return, per Faker’s own description of the
Faker class. Python’s random module, in turn, is explicit that it is a
pseudo-random number generator: deterministic, not truly random. Its
documentation names the algorithm directly: “Python uses the Mersenne Twister
as the core generator,” and warns plainly that this makes it “completely
unsuitable for cryptographic purposes.” A dedicated warning box in the same
page tells readers who need real security to reach for the secrets module
instead.
That distinction matters more for test fixtures than it first looks. Because
the sequence is deterministic given a starting state, seeding the generator
(Faker.seed(n)) fixes that starting state, and every subsequent call
becomes reproducible: two runs with the same seed produce byte-identical
output. That is precisely what a test suite needs. A fixture that changes
between runs turns a reproducible failure into a flaky one; a fixture that
reproduces means a CI failure can be recreated exactly on a developer’s own
machine by passing the same seed. “Random” here means “replayable,” not
“unpredictable,” and that is a feature, not a limitation.
The same determinism is also the reason Faker (and the random module it
sits on) must never generate anything security-sensitive: a token, a password,
an API key. If the seed or internal state is ever knowable or guessable, every
value the generator has produced or will produce is knowable too. Faker is a
plausibility engine, not an entropy source.
Underneath the Faker object, data is organized into providers: classes
that each own one topical slice of fake data (people, addresses, the internet,
lorem text) and the methods that generate it. Locale-specific data lives
underneath the same provider interface, so a locale-aware provider can back
the same method name, address() for instance, with a different underlying
data table per locale, without the calling code changing at all. This
project’s use of that architecture, or rather its non-use of the locale half of
it, is covered in Locale: none configured below.
How it fits this project#
Faker plays two distinct roles here, and only one of them is Faker being
called by this project’s own code. src/test_data_workbench/cli/demo.py
instantiates a Faker() and calls it directly to seed the bundled SQLite
demo database. Everywhere else that matters more, in
src/test_data_workbench/adaptation/generator_factory.py and
src/test_data_workbench/core/template_generator.py, this project does not
call Faker at all. Those modules write Python source text that calls Faker,
through Jinja2 templates (see Jinja2). The generator classes that
result, and that tdw deploy writes to disk, contain their own
from faker import Faker and their own calls into it; Faker is a runtime
dependency of the code you receive, exactly as the design decision described
in Design Principles (“generated code is the product”)
would predict.
Two call sites#
Faker is called from two genuinely different places in this codebase: the tool’s own code, directly, and the generator code the tool writes and hands back to you.
Called directly, at generation time#
build_demo_database in cli/demo.py is the one place in this codebase
where the tool itself runs Faker to produce values, rather than writing code
that will. It seeds once via the shared seed_generation helper (see
Seeding and reproducibility: two seed calls, not one), creates a single
Faker() instance, and passes it into a handful of small insert helpers:
conn.execute(
"INSERT INTO users (id, username, email, first_name, last_name, created_at) "
"VALUES (?, ?, ?, ?, ?, ?)",
(
i,
fake.user_name(),
fake.unique.email(),
fake.first_name(),
fake.last_name(),
fake.date_time_this_year().isoformat(),
),
)
This is the tdw demo zero-setup path: it exists purely to build a small,
plausible-looking SQLite database for someone to point the rest of the tool at
without a database of their own. It is a legitimate but minor role: five
insert helpers, one shared Faker instance, done.
Called from the code it writes#
The role that matters is the opposite direction. generator_factory.py’s
_create_generic_generator (and its four sibling entity-specific methods)
builds a Jinja2 Template whose rendered output is a complete Python module:
from faker import Faker
from typing import Dict, Any, Optional, List
import random
class {{ class_name }}Generator:
"""Generated generic data generator."""
def __init__(self, {% if has_relations %}reference_data: Dict[str, List[Any]] = None, {% endif %}seed: Optional[int] = None):
self.fake = Faker()
{%- if has_relations %}
self.reference_data = reference_data or {}
{%- endif %}
{%- if has_sequential_pk %}
self._pk_counter = 0
{%- endif %}
{%- if uses_self_reference_helper %}
self._self_generated: Dict[str, list] = {}
{%- endif %}
if seed is not None:
Faker.seed(seed)
random.seed(seed)
The {{ }} and {% %} markers are Jinja2 syntax, not Python; once
rendered for a real table, this becomes an ordinary .py file that imports
faker on its own account. GeneratorFactory (the tool) never imports or
calls Faker to produce a value; it only ever emits text that, once written to
disk and run, does. The same is true of core/template_generator.py’s
TeamTemplateGenerator, which renders the three skill-level contribution
templates described in Jinja2; every one of them also opens with
from faker import Faker. Treat every self.fake.something() you read in
this chapter as a description of code the tool writes, unless the surrounding
text says otherwise.
Provider selection per column: names before types#
generator_factory.py decides which Faker method a column gets through two
layers, checked in order, in _get_column_generator and duplicated inline in
the entity-specific generator methods. Column-name patterns are checked
first, regardless of the column’s declared SQL type:
if 'email' in col_name_lower:
return 'self.fake.email()'
elif 'phone' in col_name_lower:
return 'self.fake.phone_number()'
elif 'name' in col_name_lower:
return 'self.fake.name()'
elif 'address' in col_name_lower:
return 'self.fake.address()'
elif 'city' in col_name_lower:
return 'self.fake.city()'
elif 'country' in col_name_lower:
return 'self.fake.country()'
elif 'url' in col_name_lower:
return 'self.fake.url()'
The declared type is the fallback, via a fixed dictionary:
self.faker_mapping = {
'varchar': 'self.fake.text',
'text': 'self.fake.text',
'integer': 'self.fake.random_int',
'bigint': 'self.fake.random_int',
'decimal': 'self.fake.pydecimal',
'numeric': 'self.fake.pydecimal',
'boolean': 'self.fake.boolean',
'date': 'self.fake.date',
'timestamp': 'self.fake.date_time',
'timestamptz': 'self.fake.date_time_this_year',
'uuid': 'self.fake.uuid4'
}
A generic varchar/text type is deferred to name-based temporal
detection (_temporal_kind) before this mapping is trusted, because SQLite
in particular stores timestamps with TEXT affinity, and a declared type of
“text” alone cannot tell the factory whether a column is really a date. If
neither layer matches anything, the column falls all the way through to
self.fake.text(max_nb_chars=50), the same “always produces something”
floor described in The Fallback Ladder, applied one column at a
time rather than one table at a time.
Primary keys: a counter or a uuid, and neither one is random#
Primary keys are the one case where the type-based mapping above is
deliberately overridden before it ever runs. An integer-typed primary key does
not get self.fake.random_int() from the table above; it gets a contiguous,
instance-level counter instead:
def _needs_sequential_pk(self, column: Column) -> bool:
"""True if this column is an integer-typed primary key and should
therefore get a contiguous sequence (1, 2, 3, ...) instead of a
random or uuid value. Non-integer primary keys (uuid, string,
etc.) are left to their existing type/name-based strategy."""
return column.is_primary_key and self._is_integer_column(column)
A real auto-increment primary key is contiguous across the whole table, and a
randomly-drawn integer is not, so this is a case where the correct choice is
to bypass Faker’s randomness entirely for the sake of realism. A non-integer
primary key (typically uuid) keeps a different, still-not-random-looking
strategy: self.fake.uuid4(), unconditionally, whenever the column’s own
name contains id. Both are exceptions to “let the type or name pick a
provider” carved out specifically because a primary key’s shape matters more
than its plausibility as fake data.
Entity-specialized calls#
The four entity-specific generator methods (user, product, order, review)
layer domain-appropriate Faker calls on top of the generic mapping, choices
about which provider reads right for a field rather than which one merely
type-checks: self.fake.catch_phrase() for a product name or title (a
punchy business-speak phrase reads more like a product name than lorem ipsum
does), self.fake.text(max_nb_chars=200) for a product description, and
self.fake.text(max_nb_chars=300) for a review comment. A review’s numeric
rating is not Faker-derived at all; it is generated once per record with
plain random.randint(1, 5) and reused for both the rating column and
any validation that checks it stayed in range.
Seeding and reproducibility: two seed calls, not one#
Both tdw deploy and tdw demo accept an optional --seed. It flows
through one shared helper, called once, before any schema analysis or
generation begins:
def seed_generation(seed: Optional[int]) -> None:
"""Seed Python's random module and Faker for reproducible output.
A no-op when ``seed`` is ``None`` (the default, unseeded behavior).
"""
if seed is None:
return
random.seed(seed)
Faker.seed(seed)
Every generated generator’s __init__ repeats the same pair of calls:
if seed is not None:
Faker.seed(seed)
random.seed(seed)
Both calls are necessary, and it is not redundancy. Faker.seed() is a
classmethod: it reseeds Faker’s shared random source process-wide, which is
what makes every Faker() instantiated afterward, in every generator, draw
from the same seeded state rather than an independently-seeded one. But not
every random choice in the generated code goes through Faker at all. The
product generator’s price, for instance, is plain standard-library
random, not a Faker provider:
gen = f"round(random.uniform(10, 1000), 2)"
and category selection is random.choice(self.categories) over a hardcoded
list, again bypassing Faker entirely. Seeding only Faker would leave every
one of these random.* calls, and the self-referential foreign-key
selection logic in the generic template (which calls random.random() and
random.choice() directly), unseeded and therefore non-reproducible. Both
seed calls exist because both random sources are genuinely in use, and
_seeding.py’s module docstring says as much: seeding happens once, at CLI
entry, “before any schema analysis, database population, or code generation
happens,” specifically so that no call has already consumed random state
before the seed lands.
Uniqueness: two different strategies, and one real gap#
Faker does not guarantee unique output by default; two calls to
fake.email() can return the same address, especially as a batch grows.
This project handles that in two different ways depending on which of the two
call sites (above) is doing the generating.
In the code the tool writes, the user and customer templates track what they have already produced themselves, in a plain instance-level set, rather than relying on Faker’s own uniqueness machinery:
def _unique_email(self) -> str:
"""Generate unique email address."""
while True:
email = self.fake.email()
if email not in self._generated_emails:
self._generated_emails.add(email)
return email
This is a deliberate readability choice, not an oversight: Design Principles treats generated code as something a reader unfamiliar with Faker has to be able to follow and edit, and a loop-and-retry against a local set reads more plainly than the semantics of a stateful proxy object.
In the code the tool runs itself (build_demo_database, see Two call
sites above), there is no such concern about an unfamiliar reader, and the
demo seeder reaches for Faker’s built-in uniqueness proxy directly instead:
fake.unique.email(). Per Faker’s own documentation, .unique guarantees
a returned value has not been produced before for the lifetime of that
Faker instance, and raises UniquenessException if it runs out of room
to keep that promise.
The real gap is that neither mechanism is driven by the schema’s actual unique
constraints. core/models.py defines ConstraintType.UNIQUE as a value
schema metadata can carry, but the schema analyzer never populates it (there
is no call to SQLAlchemy’s get_unique_constraints anywhere in this
codebase), and generator_factory.py never checks it either. The dedup
handling above is triggered purely by a column being named something
containing email; a schema with a real unique constraint on, say, a
username or a sku column gets no deduplication at all, and the two
values can collide the same way any unseeded Faker output can. Worth noting:
the team-facing config.yaml that ConfigurationBuilder writes reports
'unique_constraints': True unconditionally, for every schema, regardless
of whether any dedup logic actually applies to that schema’s unique columns.
Locale: none configured#
Nowhere in this codebase does a Faker() call pass a locale. Every
instantiation, in the generated-user, generated-product, generated-order,
generated-review, and generic templates in generator_factory.py, in the
three skill-level templates in template_generator.py, in the demo-database
seeder in cli/demo.py, and in the three static example templates under
generators/templates/, is a bare Faker(). That means every value this
project produces, across every table and every generator, comes from Faker’s
default locale (en_US). If a project needed locale-appropriate data, for
instance Dutch-shaped addresses and phone numbers for a Netherlands-facing
schema, nothing here currently selects one; per the “correctability” principle
in Design Principles, that would be a one-line edit to the
emitted Faker() call in the generated file, not a configuration option
this tool currently exposes.
A custom provider that is defined but never registered#
ConfigurationBuilder.generate_faker_providers builds a custom Faker
provider, structurally correct Faker usage that this project’s generated code
never actually gets to use:
def generate_faker_providers(self, entity_types: Dict[str, EntityType]) -> List[BaseProvider]:
"""Create custom Faker providers matching schema patterns."""
providers = []
# E-commerce specific provider
class EcommerceProvider(BaseProvider):
"""Custom provider for e-commerce data."""
def product_category(self):
...
def order_status(self):
...
def product_sku(self):
...
def price(self, min_price=5.0, max_price=500.0):
...
def customer_tier(self):
...
providers.append(EcommerceProvider)
return providers
EcommerceProvider subclasses Faker’s own faker.providers.BaseProvider
correctly, and its five methods (bodies omitted above for length) are
reasonable domain-specific generators. But nothing in this codebase ever calls
some_faker_instance.add_provider(EcommerceProvider), the step that would
actually attach it to a Faker object and make fake.product_sku()
callable. No generated generator template references it, and
rapid_deployment.py never wires its return value into anything downstream.
The one test that touches it, test_faker_providers_returned, only asserts
len(providers) >= 1, which passes whether or not the provider is ever
registered anywhere. Product categories and prices in the actual generated
product template come from a hardcoded self.categories list and
random.uniform, not from this provider. As written, EcommerceProvider
is real, correct, tested-to-exist Faker provider code that nothing in the
system currently uses.
Learn more resources#
Official documentation#
Welcome to Faker’s documentation!: the project’s own overview, including how the shared
random.Randominstance and theseed()method work.Using the Faker Class: covers
Faker.seed()as a classmethod,seed_instance()for per-instance seeding, and the.uniqueproxy plusUniquenessException.Standard Providers: the full list of built-in providers, useful when picking a method by hand instead of relying on this project’s name/type mapping.
The Python random module: Python’s own documentation for the module Faker is built on, including the Mersenne Twister attribution and the explicit warning that it is not suitable for cryptographic or security use.
Tutorials and blogs#
Python Faker Library (GeeksforGeeks): a practical walkthrough of installation, common provider methods, locale selection, and seeding, useful as a broader tour than this chapter’s project-specific focus.
Videos#
How to create fake data with Python and Faker | Aiven Developer Tips (Aiven): a short, focused walkthrough of installing Faker and generating fake records, verified via YouTube’s oEmbed API.