Skip to content

Bloviate Architecture

This document is a technical deep-dive into how Bloviate works — the design decisions and the genuinely interesting machinery behind the scenes. If you just want to use Bloviate, start with the README. If you want to understand it, extend it, or contribute, you’re in the right place.

All class references below link to real code. Open paths are relative to the repository root.

At its core, Bloviate does something deceptively simple to describe: point it at a JDBC database, and it fills every table with type-appropriate, constraint-respecting, reproducible data. The interesting part is everything required to make that “just work” without you writing a single line of generation code.

The end-to-end pipeline:

flowchart LR
    A[JDBC Connection] --> B[DatabaseUtils.getMetadata]
    B --> C["Database / Table / Column<br/>(immutable records)"]
    C --> D[buildReversedDependencyGraph]
    D --> E[TopologicalOrderIterator]
    E --> F["TableFiller<br/>(per table, in order)"]
    F --> G[("Populated<br/>database")]

    C -.same generators.-> H[FlatFileGenerator]
    H -.-> I[("CSV / TSV / pipe<br/>files")]

Two entry points share the same generator engine: DatabaseFiller (fill a live database) and FlatFileGenerator (write flat files, no database required).

Schema introspection — reading the database

Section titled “Schema introspection — reading the database”

Bloviate never asks you to describe your schema — it reads it. DatabaseUtils.getMetadata(Connection) walks the standard JDBC DatabaseMetaData API to discover tables, columns (with type, size, precision, nullability), primary keys, and foreign keys, then assembles them into an immutable model:

RecordRepresents
DatabaseCatalog + the set of tables
TableColumns, primary key, foreign keys, generated INSERT SQL
ColumnJDBCType, vendor typeName, size/precision, ordinal position
PrimaryKey / ForeignKeyKey relationships used to build the dependency graph

These are all Java records — immutable, boilerplate-free value types. The metadata model is the single source of truth that every later stage reads from.

Discovery goes through the DatabaseSupport (discoveredTableTypes(), readPartitions(...)): on PostgreSQL a declaratively partitioned table is discovered as one table (its parent) and its partitions, at any depth, are left out of the model, so rows are inserted through the parent and routed by the database. A foreign key to a partitioned table is read once against it; the per-partition copies PostgreSQL clones onto the referencing table are dropped. Other databases keep discovering plain TABLEs only. See Partitioned tables.

The dependency DAG — fill order via topological sort

Section titled “The dependency DAG — fill order via topological sort”

This is the headline feature. You can’t insert an order row before the customer it references exists, so Bloviate has to fill parent tables before child tables. It figures the order out automatically by modeling the schema as a directed graph and topologically sorting it, using the JGraphT library.

DatabaseFiller.buildReversedDependencyGraph builds a DefaultDirectedGraph<Table, DefaultEdge> where each foreign key adds an edge from the child table (the one holding the FK) to the parent table it references. It then wraps the result in an EdgeReversedGraph so that a TopologicalOrderIterator yields parents before the children that depend on them:

flowchart TD
    subgraph build["1 - edges follow foreign keys (child to parent)"]
        direction LR
        stock1[stock] --> warehouse1[warehouse]
        stock1 --> item1[item]
        district1[district] --> warehouse1
        customer1[customer] --> district1
        history1[history] --> customer1
        history1 --> district1
        open_order1[open_order] --> customer1
        new_order1[new_order] --> open_order1
        order_line1[order_line] --> open_order1
        order_line1 --> stock1
    end

    subgraph sort["2 - reverse + topological sort = safe fill order"]
        direction LR
        warehouse2[warehouse] --> stock2[stock]
        item2[item] --> stock2
        warehouse2 --> district2[district]
        district2 --> customer2[customer]
        district2 --> history2[history]
        customer2 --> history2
        customer2 --> open_order2[open_order]
        open_order2 --> new_order2[new_order]
        open_order2 --> order_line2[order_line]
        stock2 --> order_line2
    end

    build --> sort

The fill loop is then just:

TopologicalOrderIterator<Table, DefaultEdge> iterator = new TopologicalOrderIterator<>(reversedGraph);
while (iterator.hasNext()) {
new TableFiller.Builder(connection, database, configuration)
.table(iterator.next())
.build().fill();
}

Table selection. The graph is built from the tables the fill selected, not necessarily every table in the schema: DatabaseFiller.Builder#includeTables/#excludeTables narrow what DatabaseUtils.getMetadata returns (before any column metadata is read), and #schema/#catalog point every connection the fill uses at another schema. A foreign key from a selected table to one that was left out has no parent to order after, so buildReversedDependencyGraph first checks every foreign key’s target is present and fails with a message naming the child table, its column(s) and the missing parent, before a single row is written. See Selecting tables and schema.

Visualize it for free. As a nice touch, DatabaseFiller exports the graph to Graphviz DOT notation with JGraphT’s DOTExporter, URL-encodes it, and logs a clickable GraphvizOnline link so you can see your schema’s dependency graph rendered in the browser — no tooling required.

See it on a real schema. Here’s the graph Bloviate emits for the TPC-C schema — the exact DOT its DOTExporter produces, rendered live below. Each parent points to the children that depend on it, which is the order tables get filled:

strict digraph tpcc {
  warehouse [ label="warehouse" ];
  item [ label="item" ];
  stock [ label="stock" ];
  district [ label="district" ];
  customer [ label="customer" ];
  history [ label="history" ];
  open_order [ label="open_order" ];
  new_order [ label="new_order" ];
  order_line [ label="order_line" ];
  warehouse -> stock;
  item -> stock;
  warehouse -> district;
  district -> customer;
  district -> history;
  customer -> history;
  customer -> open_order;
  open_order -> new_order;
  open_order -> order_line;
  stock -> order_line;
}

That diagram is rendered straight from the DOT above. Open it in GraphvizOnline → — the very link DatabaseFiller logs at fill time, where you can pan, zoom, and edit the DOT yourself.

Cycle handling. A self-referencing foreign key (a table pointing at itself) constrains the order of rows within one table, not the order of tables, so the self-edge is left out of the graph and logged with a warning — the fill proceeds, and parent/child ordering inside that table is yours to arrange.

A mutual cycle (two or more tables referencing each other, directly or through a chain) is different: no order satisfies it, because whichever table is filled first has a foreign key with no parent row to point at. Both ordering paths reject it before a row is written, with an error naming every cycle. The one way to fill such a schema is BulkLoadStrategy.unorderedBulk(), which needs no order at all because it does not enforce constraints. It is only selected for a DataSource running more than one worker thread, and where the support implements bulk loading: PostgreSQL, MySQL and MariaDB (by suspending enforcement) and BigQuery (whose key constraints are always NOT ENFORCED, so there is nothing to suspend).

The loop above is sequential and runs on a single Connection — the default, and the only option when you hand DatabaseFiller a Connection. Construct it from a pooled DataSource instead and call threads(n), and the fill runs in parallel.

The key insight: tables in the same topological “level” have no foreign key between them, so they can fill concurrently. DatabaseFiller.fillLevels groups the graph into levels with Kahn’s algorithm — level 0 is every table that references nothing, level 1 the tables whose parents are all in level 0, and so on:

flowchart TD
    subgraph L0["level 0 — fill concurrently"]
        warehouse & item
    end
    subgraph L1["level 1 — fill concurrently"]
        district & stock
    end
    subgraph L2["level 2"]
        customer
    end
    subgraph L3["level 3 — fill concurrently"]
        history & open_order
    end
    subgraph L4["level 4 — fill concurrently"]
        new_order & order_line
    end
    L0 -->|barrier| L1
    L1 -->|barrier| L2
    L2 -->|barrier| L3
    L3 -->|barrier| L4

Each table is filled by a worker that borrows its own Connection from the pool (JDBC connections are not thread-safe) inside its own transaction, and a barrier between levels guarantees a child table never starts before its parents are committed. The fill stays fully reproducible: a table’s data depends only on its own per-column seeds and row order, never on which tables fill alongside it, so for the same seed a parallel fill yields the same row content as a sequential one across every deterministic column — physical row order and wall-clock columns aside (reproducible seeds). The win is largest for wide schemas of independent tables and small for deep, narrow FK chains (little fans out within a level).

Filling a table — generators, batching, and FK fidelity

Section titled “Filling a table — generators, batching, and FK fidelity”

TableFiller handles one table. For each column it resolves a DataGenerator, then loops rowCount times generating values and binding them to a PreparedStatement.

Batch inserts. Rows are accumulated with addBatch() and flushed with executeBatch() every batchSize rows (default 1000), with a final flush for the remainder — keeping inserts efficient on large datasets.

Foreign-key fidelity. This is subtle and clever. Bloviate never reads a parent’s rows back: a foreign-key column is made to replay the key it points at, generating from the same seed at the same row index (see reproducibility), so the values line up by construction. ForeignKeyPlan works that out once for the whole database, from three rules:

  • Which key. Edges come from the columns the constraint actually names (PKCOLUMN_NAME), so a key to a UNIQUE target — the tenant pattern, PRIMARY KEY (id) alongside UNIQUE (tenant_id, id) — resolves to the columns it really references rather than to the parent’s primary key by position.
  • Which seed. Columns linked by a foreign key, in either direction, form one class, and the whole class generates from a single seed: the class’s seed source. For an ordinary parent/child chain that is the column at the far end, exactly as before. It matters when one column belongs to several keys — its value must exist in every key it references, which only holds if those keys carry the same values, so the rule ties them into one class. A schema that points one column at two unrelated keys therefore makes those two keys equal, column for column; that is inherent, since a value cannot be in both key spaces unless the key spaces overlap.
  • How far. Row i of a child reads row i % parentRows of its parent, which itself reads row (i % parentRows) % grandparentRows of the grandparent, and so on up the chain — so a child sized larger than any of its ancestors cycles through their keys rather than running past them. The counts come from each table’s TableConfiguration when it has one and the default row count otherwise. Folding in order is what keeps a composite key’s columns on one parent row: reduce each column by the smallest count in its chain instead and, where the counts are not multiples of each other, the columns land on different parent rows and the tuple matches nothing. A level is bounded by several tables only when a column references keys that do not imply one another, and then the smallest wins.

A key the database assigns (serial, IDENTITY) is left out of its own insert, so there is no seed to share with it. A table filled from empty is assigned 1..N in insertion order, which the child reproduces by counting through the same range — and, because a class carries one set of values, every column of that class counts rather than only the ones referencing the generated key.

flowchart TD
    A[For each column in table] --> B{"In a foreign key?"}
    B -->|no| D["seed = columnSeed(this column)"]
    B -->|yes| C{"Is the referenced key database-generated?"}
    C -->|yes| I["count 1..N of the referenced table"]
    C -->|no| J["seed = columnSeed(seed source of its class)<br/>wrap at the smallest referenced row count"]
    D --> E[resolve generator by precedence]
    I --> E
    J --> E
    E --> F["generate value -> bind to PreparedStatement"]
    F --> G{rowCounter % batchSize == 0?}
    G -->|yes| H[executeBatch]
    G -->|no| A

The inner loop is the hot path (it runs rowCount × columnCount times), so TableFiller resolves every column’s generator, seed, and FK wrap limit once into positional arrays indexed by column ordinal — the loop then does array reads instead of hashing the Column on every cell.

By default the engine leaves the connection’s autocommit state untouched — an autocommit connection commits per executeBatch(). A CommitStrategy on DatabaseConfiguration lets you cut that overhead: perTable() disables autocommit and commits once when the table is filled, and everyNBatches(n) commits every n batches to bound the open transaction. When a strategy is set, TableFiller owns the transaction (autocommit off, commit at the configured cadence, rollback on error, prior autocommit restored afterward); the default connectionDefault() preserves the original behavior. The parallel path’s per-table commit is the same mechanism — its workers run with an effective perTable() strategy.

Intra-table partitioning — seeking to a row range

Section titled “Intra-table partitioning — seeking to a row range”

Parallel fill (topological levels) parallelizes across tables, which doesn’t help when a single huge table dominates — it fills alone in its level. Set partitions on that table’s TableConfiguration and, on the parallel path, DatabaseFiller splits its [0, rowCount) rows into that many contiguous ranges filled concurrently, one Connection per range.

The challenge is reproducibility: a worker starting at absolute row N must produce the value the sequential fill produces at row N, without replaying rows 0..N. The engine solves this with per-index derivation: each column’s random source is an IndexedRandom that TableFiller repositions to the absolute row index before every cell, so every generated value is a pure function of (columnSeed, rowIndex). Seeking to any row is O(1) — there is nothing to replay. A foreign-key column positions at the parent’s index (rowIndex % parentRows) instead, which replays the parent’s exact value and makes key-space wraparound a formula. Counter/cursor generators additionally implement IndexedDataGenerator, whose seek(rowIndex) positions their counters directly (the composite/sequence/permutation generators are closed-form O(1); the variable-cardinality child-key generator locates the owning parent via the ChildCardinality cumulative).

Column kindPositioningResult vs. sequential
Positionable (all built-in generators)IndexedRandom.position(rowIndex) before every cellbyte-identical, any partition count
Foreign-key on a positionable generatorposition(rowIndex % parentRows)byte-identical (FK-valid), any partition count
Positional counters (IndexedDataGenerator — keys, sequences, permutations, prefixes)seek(start) at the partition boundarybyte-identical
Non-positionable (opt-out custom generators, datafaker)legacy: per-partition reseed; foreign keys replay start % parentRows drawsdeterministic per partition count

Because every built-in column — not just keys — is a pure function of the row index, foreign-key validity always holds and a partitioned fill is byte-identical to the sequential fill of the same seed, for any partition count. Only non-positionable columns (a custom DataGenerator that opts out of positionable(), such as the datafaker integration whose values come from an internal Faker RNG) fall back to the legacy per-partition behavior. One edge remains unsupported: partitioning a parent whose primary key comes from a non-positionable custom generator referenced by a foreign key (partition the child instead, or use the positional key generators). A custom generator with internal positional state must implement IndexedDataGenerator to stay aligned under partitioning.

Bulk load — unordered fill with constraints disabled

Section titled “Bulk load — unordered fill with constraints disabled”

Parallel fill barriers between topological levels, which costs the most on a deep, narrow foreign-key chain: each level holds few tables, so little fans out and the fill effectively serializes down the chain. BulkLoadStrategy.unorderedBulk() collapses every level into a single wave — one task per table (or partition), submitted at once with no barrier — and disables foreign-key enforcement for the duration.

This is only sound because of foreign-key fidelity: an FK column is seeded from its referenced PK column, so the data is referentially consistent regardless of insert order. Disabling enforcement removes the per-row check and the ordering requirement without changing a single generated value — so for the same seed the result has the same row content as an ordered fill across every deterministic column (physical row order aside).

flowchart TD
    P["probe once: disable + re-enable on a borrowed connection"] -->|privilege ok| W
    P -->|BulkLoadUnsupportedException| FB["fall back to ordered level-parallel fill"]
    subgraph W["single wave — every table at once, no barrier"]
        direction LR
        T1["worker: disable → fill table → re-enable (finally)"]
        T2["worker: disable → fill table → re-enable (finally)"]
        T3["worker: disable → fill partition → re-enable (finally)"]
    end

The mechanism is database-specific and lives behind two DatabaseSupport SPI methods, disableConstraints / enableConstraints, guarded by supportsBulkLoad():

SupportMechanism (per session)Notes
PostgreSQLSET session_replication_role = replica → originneeds a superuser/rds_superuser role; privilege failure raises BulkLoadUnsupportedException
MySQLSET FOREIGN_KEY_CHECKS=0/UNIQUE_CHECKS=0 → 1no special privilege
CockroachDBunsupported (supportsBulkLoad() is false)no session_replication_role; falls back to ordered

Because these settings are per connection and each worker borrows its own from the pool, the disable/enable runs inside every worker task, wrapped in a try/finally that restores the session before the connection returns to the pool — so a constraint-disabled connection never leaks to other pool users, even if a fill throws. Privilege is probed once up front on a throwaway connection; if it fails, the engine logs a warning and runs the ordered level-parallel path instead of fanning out half-disabled. Bulk mode only applies to the DataSource + threads > 1 path; it is ignored with a warning elsewhere. The default BulkLoadStrategy.ordered() keeps the dependency-ordered behavior with constraints always enforced.

Reproducibility — deterministic seeds from schema identity

Section titled “Reproducibility — deterministic seeds from schema identity”

Bloviate datasets are reproducible across JVM runs, machines, and time — run it twice against the same schema with the same base seed and you get byte-for-byte identical data. This is not done by seeding one global Random; it’s done per column.

DatabaseUtils.columnSeed(Column, baseSeed) derives a stable seed by hashing the column’s identity — its name, table, schema, catalog, JDBC type name, and ordinal position — and mixing it with the base seed:

int identity = Objects.hash(
column.name(), column.tableName(), column.schema(), column.catalog(),
column.jdbcType() == null ? null : column.jdbcType().getName(),
column.ordinalPosition());
return baseSeed * 1_000_003L + identity;

Two important properties fall out of this design:

  • Run-independence. The seed depends only on schema identity, never on iteration order, hash-map ordering, or wall-clock time. The JDBC type’s name is hashed (not the enum’s ordinal()), so the seed is stable even if the enum changes.
  • FK alignment for free. Because the seed is a pure function of the column, a foreign key column and the primary key it references resolve to the same seed — which is exactly what makes the FK-fidelity trick in the table-fill section work.

Every generator — built-in, registry-supplied, or per-column override — is constructed with this engine-managed seed, so reproducibility holds no matter how a column’s generator was chosen.

The one other input a factory can receive is the fill’s GenerationContext, created once when DatabaseFiller.fill() starts and handed to every TableFiller (so to every table, partition and worker): it carries the asOf anchor that relative date windows resolve against. ColumnGeneratorFactory and GeneratorFactory gained a default create(..., GenerationContext) overload that ignores it, so existing factories are untouched. A pinned asOf keeps the data a pure function of seed and anchor; wall-clock time is only read, once, when none is pinned (see Relative date ranges and asOf).

The per-column seed feeds an IndexedRandom — a repositionable SplitMix64 stream (the construction behind java.util.SplittableRandom) that the engine positions to the absolute row index before every cell, so each value is a pure function of (columnSeed, rowIndex). The seeding architecture is one isolated, deterministically-seeded random source per column, so output is reproducible and order-independent. (Generators used outside the fill engine — flat files, caller-constructed — draw from RandomGenerators.create(seed), the JDK general-purpose default L64X128MixRandom.)

Because a value depends only on its column’s seed and row index — never on timing, prior rows, or which tables fill alongside it — reproducibility survives concurrency. A parallel table fill (parallel fill) yields the same row content as a sequential one across every deterministic column. An intra-table partitioned fill (intra-table partitioning) is byte-identical to the sequential fill for any partition count across every built-in generator; only non-positionable custom generators vary with the partition count.

Different databases expose different types. DatabaseSupport is the strategy interface that maps a Column to a generator; AbstractDatabaseSupport holds an EnumMap<JDBCType, GeneratorFactory> of cross-database defaults and exposes a configure() hook that subclasses override to add or replace entries for vendor-specific types.

classDiagram
    class DatabaseSupport {
        <<interface>>
        +getDataGenerator(Column, RandomGenerator) DataGenerator
        +forConnection(Connection)$ DatabaseSupport
        +forProduct(String)$ DatabaseSupport
    }
    class AbstractDatabaseSupport {
        <<abstract>>
        -EnumMap~JDBCType,GeneratorFactory~ registry
        #configure(Map) void
    }
    DatabaseSupport <|.. AbstractDatabaseSupport
    AbstractDatabaseSupport <|-- DefaultSupport
    AbstractDatabaseSupport <|-- PostgresSupport
    AbstractDatabaseSupport <|-- MySQLSupport
    AbstractDatabaseSupport <|-- H2Support
    AbstractDatabaseSupport <|-- SQLiteSupport
    PostgresSupport <|-- CockroachDBSupport
    MySQLSupport <|-- MariaDBSupport
ImplementationAdds on top of the JDBC defaults
DefaultSupportNothing — cross-database JDBC types only
PostgresSupportuuid, json/jsonb, inet, cidr, macaddr/macaddr8, interval, bit/varbit, xml, and text/int arrays
MySQLSupportJSON columns generate valid JSON instead of arbitrary text; BIT(n) generates an n-bit number, not a bit string
MariaDBSupportExtends MySQLSupport — MariaDB columns surface through JDBC essentially as MySQL’s
CockroachDBSupportExtends PostgresSupport (CockroachDB is PG wire-compatible), including its CHECK/ENUM constraint reading
H2SupportUUID (reported as BINARY) and valid JSON
SQLiteSupportNothing — SQLite’s affinity types collapse onto INTEGER/FLOAT/VARCHAR, already covered by the defaults

You don’t have to pick manually. DatabaseSupport.forConnection(connection) reads DatabaseMetaData.getDatabaseProductName() and selects the right strategy by substring match, falling back to DefaultSupport. (Note: CockroachDB reached via the PG driver reports as PostgreSQL and resolves to PostgresSupport — equivalent, since CockroachDBSupport adds no extra behavior.)

DatabaseSupport also reads value constraints for a table — DatabaseSupport.readConstraints(...). The default is none; PostgresSupport queries pg_constraint/pg_enum, parses the common CHECK forms (IN, BETWEEN, comparisons, first-of-month date_trunc/EXTRACT on dates) and enum labels into a ColumnConstraint, and TableFiller then prefers a constraint-satisfying generator (categorical, bounded numeric, or TruncatedDateGenerator) over the type default — so generated values conform instead of being rejected. The parser recognises only exact forms (a bare, optionally cast, column against literals; never an expression containing a function call other than the first-of-month spellings); anything else is logged and skipped.

Pluggable generators — Registry + ServiceLoader

Section titled “Pluggable generators — Registry + ServiceLoader”

Sometimes the type-based default isn’t what you want — you want every column named email to look like an email, regardless of its SQL type. The GeneratorRegistry lets you override generation without subclassing DatabaseSupport, and external jars can contribute rules automatically via Java’s ServiceLoader.

A registry supports three matcher kinds, and the fill engine resolves each column through a fixed precedence chain:

flowchart TD
    Col[Column to fill] --> CC{per-column<br/>ColumnConfiguration?}
    CC -->|yes| Use1[use it]
    CC -->|no| NP{registry name<br/>pattern match?}
    NP -->|yes| Use2[use it]
    NP -->|no| TN{registry vendor<br/>typeName match?}
    TN -->|yes| Use3[use it]
    TN -->|no| JT{registry<br/>JDBCType match?}
    JT -->|yes| Use4[use it]
    JT -->|no| Def["DatabaseSupport default"]

Plugins implement the single-method GeneratorPlugin SPI and declare themselves in META-INF/services/io.bloviate.ext.GeneratorPlugin. Calling GeneratorRegistry.Builder.discover() loads every plugin on the classpath:

GeneratorRegistry registry = new GeneratorRegistry.Builder()
.registerColumnNamePattern(".*email", (column, random) -> new EmailGenerator.Builder(random).build())
.registerTypeName("uuid", (column, random) -> new UUIDGenerator.Builder(random).build())
.discover() // pick up GeneratorPlugin services from the classpath
.build();

Crucially, registry- and plugin-supplied generators are still constructed with the engine’s seeded RandomGenerator, so they remain just as reproducible as the built-ins. The optional bloviate-datafaker module is exactly this pattern in practice: one GeneratorPlugin that maps column names (email, first_name, phone, …) to realistic Datafaker values, seeded from the engine’s column seed for reproducibility — keeping Datafaker out of the core. It also offers referential realism via a RowContext: correlated columns project fields of one per-row entity — a Person whose email/username derive from its name, or a Geo tuple (city/state/zip/area-code that agree) drawn from a bundled reference dataset — computed as a pure function of (seed, rowIndex) so consistency survives parallel and partitioned fills.

Every generator implements DataGenerator<T>, which can generate() a typed value, generateAsString() for flat files, bind itself to a PreparedStatement (generateAndSet), and read a value back from a ResultSet (get). Each is constructed through a static inner Builder seeded with a RandomGenerator:

new SimpleStringGenerator.Builder(random).size(100).build();
new BigDecimalGenerator.Builder(random).precision(10).digits(2).build();

The library ships ~50 generators in io.bloviate.gen, covering everything from primitives and dates to uuid, jsonb, inet/cidr, MAC addresses, intervals, arrays, and XML. A few are worth calling out:

  • Referential-fidelity generators — CompositeKeyComponentGenerator, ChildKeyComponentGenerator, and ChildCountGenerator produce collision-free composite keys and variable parent/child cardinalities that keep foreign keys consistent.
  • GroupedPermutationGenerator — emits a deterministic permutation per group (e.g. TPC-C’s shuffled o_c_id) using a Feistel network with cycle-walking, achieving a unique pseudo-random permutation in O(1) memory without materializing or shuffling an array.
  • TPC-C generators — io.bloviate.gen.tpcc provides benchmark-faithful fields (customer last names, zip codes, credit, delivery dates).
  • Distribution generators — WeightedCategoricalGenerator, NormalDoubleGenerator/NormalIntegerGenerator, ZipfianIntegerGenerator, and SkewedTimestampGenerator emit non-uniform values (categorical, bounded-Gaussian, power-law, recency-skewed). A column opts in through the Distributions convenience without writing a factory. These are specified distributions, not learned from data — so they stay deterministic by seed and compose with FK reseeding and parallel fills like any other generator.

The same generators power FlatFileGenerator, which needs no database at all. You describe columns with ColumnDefinition (name + generator), pick a FileType (CSV, TDV, or PIPE), and generate(). Output is written via Apache Commons CSV, with headers derived from column names:

new FlatFileGenerator.Builder("output/users")
.add(new ColumnDefinition("id", new IntegerGenerator.Builder(random).build()))
.add(new ColumnDefinition("email", new SimpleStringGenerator.Builder(random).build()))
.rows(1000)
.build()
.generate();

Bloviate is a Maven reactor. The engine is self-contained; the integration modules pull in their framework as a provided dependency so you bring your own version.

graph TD
    core["bloviate-core<br/><i>self-contained engine</i><br/>db · ext · gen · file · util"]
    junit["bloviate-junit<br/><i>@FillDatabase / @FillSource</i><br/>BloviateExtension"]
    tc["bloviate-testcontainers<br/><i>BloviateContainers</i>"]
    df["bloviate-datafaker<br/><i>DatafakerGeneratorPlugin</i>"]
    junit -->|depends on| core
    tc -->|depends on| core
    df -->|depends on| core
    junit -.->|provided| j[JUnit Jupiter]
    tc -.->|provided| t[Testcontainers]
    df -->|depends on| d[Datafaker]

Bloviate leans deliberately on a small, well-understood set of design patterns. They aren’t applied for their own sake — each one buys a concrete property (extensibility, immutability, testability) and they compose cleanly. If you’ve read the Gang of Four, this table is a fast map of where each pattern lives and what it’s doing for us.

PatternWhere it livesWhat it buys
BuilderNearly everything: DatabaseFiller.Builder, TableFiller.Builder, every *Generator.Builder, GeneratorRegistry.Builder, FlatFileGenerator.Builder, BloviateContainers.BuilderReadable construction of objects with many optional, defaulted parameters; immutable results with no telescoping constructors
StrategyDatabaseSupport + per-database implementationsDatabase-specific behavior is swappable at runtime; adding a database means adding a class, not editing the engine
Template MethodAbstractDatabaseSupport seeds defaults, then calls the configure() hookSubclasses customize only the vendor-specific slice; the invariant default registry is defined once
Factory (functional)GeneratorFactory, ColumnGeneratorFactoryDefers generator creation until the engine can supply a column-seeded RandomGenerator, preserving reproducibility
RegistryGeneratorRegistry with documented precedenceOverride generation by name/type without subclassing; rules resolved in a fixed, predictable order
Service Provider (SPI)GeneratorPlugin via ServiceLoaderThird-party jars contribute generators by dropping a file on the classpath — zero engine coupling
Strategy / polymorphismThe DataGenerator<T> hierarchy; CommitStrategy for transaction cadenceOne uniform interface (generate / bind / read-back) over ~50 type-specific implementations; swappable commit behavior
Seekable cursorIndexedDataGenerator.seek(long) on the positional generatorsPosition a generator at an absolute row index without replaying — the basis for intra-table partitioning
IteratorJGraphT’s TopologicalOrderIterator in DatabaseFillerFill order is expressed as a traversal, decoupled from graph construction
Adapterbloviate-junit and bloviate-testcontainersWrap the same core engine behind framework-native front-ends (@FillDatabase, BloviateContainers)

The payoff is that the two extension axes you actually care about — “support a new database” and “generate a new kind of value” — are both open for extension without modifying a line of the core engine (Open/Closed Principle). Strategy handles the first; Registry + SPI handle the second.

A few themes recur throughout the codebase and explain most of the “why” — they’re the result of deliberate design effort, not incidental:

  1. Read, don’t configure. Schema is introspected from JDBC metadata, not hand-described.
  2. Immutability. The metadata and configuration models are Java records; registries are built once and copied defensively.
  3. Determinism by construction. Seeds derive from schema identity, never runtime state — so data is reproducible and foreign keys align without bookkeeping.
  4. Open for extension, closed for modification. The Strategy pattern (DatabaseSupport) handles new databases; the Registry + ServiceLoader (GeneratorPlugin) handles new generation rules — neither requires touching the engine.
  5. One engine, many front-ends. Database filling, flat files, JUnit, and Testcontainers all sit on the same bloviate-core generators.

These principles aren’t abstract — they show up as concrete quality properties you can rely on:

  • A lean core. bloviate-core keeps its dependency surface small and pushes JUnit and Testcontainers to provided scope, so integrating Bloviate doesn’t drag a testing framework into your runtime classpath, and you bring your own versions.
  • Tested against real databases. Integration tests run against actual PostgreSQL, MySQL, MariaDB, and CockroachDB instances via Testcontainers — plus embedded H2 and SQLite — not mocks — over real benchmark schemas (TPC-C, AuctionMark, Wikipedia). Behavior is verified end-to-end, including the FK ordering and round-trip read-back through each generator’s get(ResultSet, ...).
  • Reproducibility as a guarantee, not a hope. Because seeds are pure functions of schema identity, “it worked on my machine” datasets are byte-for-byte portable — which is exactly what you want from test fixtures and benchmark data.
  • Modern Java, used deliberately. Records for the immutable model, sealed extension points, functional SPIs (@FunctionalInterface), and EnumMap-backed registries reflect a codebase built on Java 25 idioms rather than retrofitted onto them.
  • Thorough documentation. The public types carry real Javadoc — including the precedence rules for generator resolution and the CockroachDB/PostgreSQL driver caveat — so the contracts are written down, not folklore.

The throughline: Bloviate is a small library that takes its own design seriously. The patterns and principles above are what let it stay simple to use while remaining genuinely extensible underneath.


See also: Quick Start to start filling databases.