Bloviate Architecture
This document is a technical deep-dive into how Bloviate works — the design decisions and the genuinely interesting machinery behind the scenes. If you just want to use Bloviate, start with the README. If you want to understand it, extend it, or contribute, you’re in the right place.
All class references below link to real code. Open paths are relative to the repository root.
The big picture
Section titled “The big picture”At its core, Bloviate does something deceptively simple to describe: point it at a JDBC database, and it fills every table with type-appropriate, constraint-respecting, reproducible data. The interesting part is everything required to make that “just work” without you writing a single line of generation code.
The end-to-end pipeline:
flowchart LR
A[JDBC Connection] --> B[DatabaseUtils.getMetadata]
B --> C["Database / Table / Column<br/>(immutable records)"]
C --> D[buildReversedDependencyGraph]
D --> E[TopologicalOrderIterator]
E --> F["TableFiller<br/>(per table, in order)"]
F --> G[("Populated<br/>database")]
C -.same generators.-> H[FlatFileGenerator]
H -.-> I[("CSV / TSV / pipe<br/>files")]
Two entry points share the same generator engine: DatabaseFiller
(fill a live database) and FlatFileGenerator
(write flat files, no database required).
Schema introspection — reading the database
Section titled “Schema introspection — reading the database”Bloviate never asks you to describe your schema — it reads it. DatabaseUtils.getMetadata(Connection)
walks the standard JDBC DatabaseMetaData
API to discover tables, columns (with type, size, precision, nullability), primary keys, and
foreign keys, then assembles them into an immutable model:
| Record | Represents |
|---|---|
Database | Catalog + the set of tables |
Table | Columns, primary key, foreign keys, generated INSERT SQL |
Column | JDBCType, vendor typeName, size/precision, ordinal position |
PrimaryKey / ForeignKey | Key relationships used to build the dependency graph |
These are all Java records — immutable, boilerplate-free value types. The metadata model is the single source of truth that every later stage reads from.
Discovery goes through the DatabaseSupport (discoveredTableTypes(), readPartitions(...)): on
PostgreSQL a declaratively partitioned table is discovered as one table (its parent) and its partitions,
at any depth, are left out of the model, so rows are inserted through the parent and routed by the
database. A foreign key to a partitioned table is read once against it; the per-partition copies PostgreSQL
clones onto the referencing table are dropped. Other databases keep discovering plain TABLEs only. See
Partitioned tables.
The dependency DAG — fill order via topological sort
Section titled “The dependency DAG — fill order via topological sort”This is the headline feature. You can’t insert an order row before the customer it references exists, so Bloviate has to fill parent tables before child tables. It figures the order out automatically by modeling the schema as a directed graph and topologically sorting it, using the JGraphT library.
DatabaseFiller.buildReversedDependencyGraph
builds a DefaultDirectedGraph<Table, DefaultEdge> where each foreign key adds an edge from the
child table (the one holding the FK) to the parent table it references. It then wraps the
result in an EdgeReversedGraph so that a TopologicalOrderIterator yields parents before the
children that depend on them:
flowchart TD
subgraph build["1 - edges follow foreign keys (child to parent)"]
direction LR
stock1[stock] --> warehouse1[warehouse]
stock1 --> item1[item]
district1[district] --> warehouse1
customer1[customer] --> district1
history1[history] --> customer1
history1 --> district1
open_order1[open_order] --> customer1
new_order1[new_order] --> open_order1
order_line1[order_line] --> open_order1
order_line1 --> stock1
end
subgraph sort["2 - reverse + topological sort = safe fill order"]
direction LR
warehouse2[warehouse] --> stock2[stock]
item2[item] --> stock2
warehouse2 --> district2[district]
district2 --> customer2[customer]
district2 --> history2[history]
customer2 --> history2
customer2 --> open_order2[open_order]
open_order2 --> new_order2[new_order]
open_order2 --> order_line2[order_line]
stock2 --> order_line2
end
build --> sort
The fill loop is then just:
TopologicalOrderIterator<Table, DefaultEdge> iterator = new TopologicalOrderIterator<>(reversedGraph);while (iterator.hasNext()) { new TableFiller.Builder(connection, database, configuration) .table(iterator.next()) .build().fill();}Table selection. The graph is built from the tables the fill selected, not necessarily every table
in the schema: DatabaseFiller.Builder#includeTables/#excludeTables narrow what
DatabaseUtils.getMetadata returns (before any column metadata is read), and #schema/#catalog
point every connection the fill uses at another schema. A foreign key from a selected table to one
that was left out has no parent to order after, so buildReversedDependencyGraph first checks every
foreign key’s target is present and fails with a message naming the child table, its column(s) and
the missing parent, before a single row is written. See
Selecting tables and schema.
Visualize it for free. As a nice touch, DatabaseFiller exports the graph to
Graphviz DOT notation with JGraphT’s DOTExporter,
URL-encodes it, and logs a clickable GraphvizOnline
link so you can see your schema’s dependency graph rendered in the browser — no tooling required.
See it on a real schema. Here’s the graph Bloviate emits for the
TPC-C schema — the exact DOT its DOTExporter produces, rendered live
below. Each parent points to the children that depend on it, which is the order tables get filled:
strict digraph tpcc {
warehouse [ label="warehouse" ];
item [ label="item" ];
stock [ label="stock" ];
district [ label="district" ];
customer [ label="customer" ];
history [ label="history" ];
open_order [ label="open_order" ];
new_order [ label="new_order" ];
order_line [ label="order_line" ];
warehouse -> stock;
item -> stock;
warehouse -> district;
district -> customer;
district -> history;
customer -> history;
customer -> open_order;
open_order -> new_order;
open_order -> order_line;
stock -> order_line;
}
That diagram is rendered straight from the DOT above.
Open it in GraphvizOnline →
— the very link DatabaseFiller logs at fill time, where you can pan, zoom, and edit the DOT yourself.
Cycle handling. A self-referencing foreign key (a table pointing at itself) constrains the order of rows within one table, not the order of tables, so the self-edge is left out of the graph and logged with a warning — the fill proceeds, and parent/child ordering inside that table is yours to arrange.
A mutual cycle (two or more tables referencing each other, directly or through a chain) is different:
no order satisfies it, because whichever table is filled first has a foreign key with no parent row to
point at. Both ordering paths reject it before a row is written, with an error naming every cycle. The
one way to fill such a schema is BulkLoadStrategy.unorderedBulk(), which needs no order at all
because it does not enforce constraints. It is only selected for a DataSource running more than one
worker thread, and where the support implements bulk loading: PostgreSQL, MySQL and MariaDB (by
suspending enforcement) and BigQuery (whose key constraints are always NOT ENFORCED, so there is
nothing to suspend).
Parallel fill — topological levels
Section titled “Parallel fill — topological levels”The loop above is sequential and runs on a single Connection — the default, and the only option
when you hand DatabaseFiller a Connection. Construct it from a pooled DataSource instead and
call threads(n), and the fill runs in parallel.
The key insight: tables in the same topological “level” have no foreign key between them, so they
can fill concurrently. DatabaseFiller.fillLevels
groups the graph into levels with Kahn’s algorithm — level 0 is every table that references nothing,
level 1 the tables whose parents are all in level 0, and so on:
flowchart TD
subgraph L0["level 0 — fill concurrently"]
warehouse & item
end
subgraph L1["level 1 — fill concurrently"]
district & stock
end
subgraph L2["level 2"]
customer
end
subgraph L3["level 3 — fill concurrently"]
history & open_order
end
subgraph L4["level 4 — fill concurrently"]
new_order & order_line
end
L0 -->|barrier| L1
L1 -->|barrier| L2
L2 -->|barrier| L3
L3 -->|barrier| L4
Each table is filled by a worker that borrows its own Connection from the pool (JDBC connections
are not thread-safe) inside its own transaction, and a barrier between levels guarantees a child
table never starts before its parents are committed. The fill stays fully reproducible: a table’s
data depends only on its own per-column seeds and row order, never on which tables fill alongside it,
so for the same seed a parallel fill yields the same row content as a sequential one across every
deterministic column — physical row order and wall-clock columns aside
(reproducible seeds). The win is largest for wide
schemas of independent tables and small for deep, narrow FK chains (little fans out within a level).
Filling a table — generators, batching, and FK fidelity
Section titled “Filling a table — generators, batching, and FK fidelity”TableFiller handles one table. For
each column it resolves a DataGenerator,
then loops rowCount times generating values and binding them to a PreparedStatement.
Batch inserts. Rows are accumulated with addBatch() and flushed with executeBatch() every
batchSize rows (default 1000), with a final flush for the remainder — keeping inserts efficient on
large datasets.
Foreign-key fidelity. This is subtle and clever. Bloviate never reads a parent’s rows back: a
foreign-key column is made to replay the key it points at, generating from the same seed at the same
row index (see reproducibility), so the
values line up by construction.
ForeignKeyPlan
works that out once for the whole database, from three rules:
- Which key. Edges come from the columns the constraint actually names (
PKCOLUMN_NAME), so a key to aUNIQUEtarget — the tenant pattern,PRIMARY KEY (id)alongsideUNIQUE (tenant_id, id)— resolves to the columns it really references rather than to the parent’s primary key by position. - Which seed. Columns linked by a foreign key, in either direction, form one class, and the whole class generates from a single seed: the class’s seed source. For an ordinary parent/child chain that is the column at the far end, exactly as before. It matters when one column belongs to several keys — its value must exist in every key it references, which only holds if those keys carry the same values, so the rule ties them into one class. A schema that points one column at two unrelated keys therefore makes those two keys equal, column for column; that is inherent, since a value cannot be in both key spaces unless the key spaces overlap.
- How far. Row i of a child reads row
i % parentRowsof its parent, which itself reads row(i % parentRows) % grandparentRowsof the grandparent, and so on up the chain — so a child sized larger than any of its ancestors cycles through their keys rather than running past them. The counts come from each table’sTableConfigurationwhen it has one and the default row count otherwise. Folding in order is what keeps a composite key’s columns on one parent row: reduce each column by the smallest count in its chain instead and, where the counts are not multiples of each other, the columns land on different parent rows and the tuple matches nothing. A level is bounded by several tables only when a column references keys that do not imply one another, and then the smallest wins.
A key the database assigns (serial, IDENTITY) is left out of its own insert, so there is no seed
to share with it. A table filled from empty is assigned 1..N in insertion order, which the child
reproduces by counting through the same range — and, because a class carries one set of values, every
column of that class counts rather than only the ones referencing the generated key.
flowchart TD
A[For each column in table] --> B{"In a foreign key?"}
B -->|no| D["seed = columnSeed(this column)"]
B -->|yes| C{"Is the referenced key database-generated?"}
C -->|yes| I["count 1..N of the referenced table"]
C -->|no| J["seed = columnSeed(seed source of its class)<br/>wrap at the smallest referenced row count"]
D --> E[resolve generator by precedence]
I --> E
J --> E
E --> F["generate value -> bind to PreparedStatement"]
F --> G{rowCounter % batchSize == 0?}
G -->|yes| H[executeBatch]
G -->|no| A
The inner loop is the hot path (it runs rowCount × columnCount times), so TableFiller resolves
every column’s generator, seed, and FK wrap limit once into positional arrays indexed by
column ordinal — the loop then does array reads instead of hashing the Column on every cell.
Commit strategy
Section titled “Commit strategy”By default the engine leaves the connection’s autocommit state untouched — an autocommit connection
commits per executeBatch(). A CommitStrategy
on DatabaseConfiguration lets you cut that overhead: perTable() disables autocommit and commits
once when the table is filled, and everyNBatches(n) commits every n batches to bound the open
transaction. When a strategy is set, TableFiller owns the transaction (autocommit off, commit at the
configured cadence, rollback on error, prior autocommit restored afterward); the default
connectionDefault() preserves the original behavior. The parallel path’s per-table commit is the
same mechanism — its workers run with an effective perTable() strategy.
Intra-table partitioning — seeking to a row range
Section titled “Intra-table partitioning — seeking to a row range”Parallel fill (topological levels) parallelizes across tables, which doesn’t
help when a single huge table dominates — it fills alone in its level. Set partitions on that
table’s TableConfiguration and, on the parallel path, DatabaseFiller splits its [0, rowCount)
rows into that many contiguous ranges filled concurrently, one Connection per range.
The challenge is reproducibility: a worker starting at absolute row N must produce the value the
sequential fill produces at row N, without replaying rows 0..N. The engine solves this with
per-index derivation: each column’s random source is an IndexedRandom that TableFiller
repositions to the absolute row index before every cell, so every generated value is a pure function
of (columnSeed, rowIndex). Seeking to any row is O(1) — there is nothing to replay. A foreign-key
column positions at the parent’s index (rowIndex % parentRows) instead, which replays the parent’s
exact value and makes key-space wraparound a formula. Counter/cursor generators additionally implement
IndexedDataGenerator,
whose seek(rowIndex) positions their counters directly (the composite/sequence/permutation
generators are closed-form O(1); the variable-cardinality child-key generator locates the owning
parent via the ChildCardinality cumulative).
| Column kind | Positioning | Result vs. sequential |
|---|---|---|
| Positionable (all built-in generators) | IndexedRandom.position(rowIndex) before every cell | byte-identical, any partition count |
| Foreign-key on a positionable generator | position(rowIndex % parentRows) | byte-identical (FK-valid), any partition count |
Positional counters (IndexedDataGenerator — keys, sequences, permutations, prefixes) | seek(start) at the partition boundary | byte-identical |
| Non-positionable (opt-out custom generators, datafaker) | legacy: per-partition reseed; foreign keys replay start % parentRows draws | deterministic per partition count |
Because every built-in column — not just keys — is a pure function of the row index,
foreign-key validity always holds and a partitioned fill is byte-identical to the sequential
fill of the same seed, for any partition count. Only non-positionable columns (a custom
DataGenerator that opts out of positionable(), such as the datafaker integration whose values
come from an internal Faker RNG) fall back to the legacy per-partition behavior. One edge remains
unsupported: partitioning a parent whose primary key comes from a non-positionable custom generator
referenced by a foreign key (partition the child instead, or use the positional key generators). A
custom generator with internal positional state must implement IndexedDataGenerator to stay aligned
under partitioning.
Bulk load — unordered fill with constraints disabled
Section titled “Bulk load — unordered fill with constraints disabled”Parallel fill barriers between topological levels, which costs
the most on a deep, narrow foreign-key chain: each level holds few tables, so little fans out and
the fill effectively serializes down the chain.
BulkLoadStrategy.unorderedBulk()
collapses every level into a single wave — one task per table (or partition), submitted at once with
no barrier — and disables foreign-key enforcement for the duration.
This is only sound because of foreign-key fidelity: an FK column is seeded from its referenced PK column, so the data is referentially consistent regardless of insert order. Disabling enforcement removes the per-row check and the ordering requirement without changing a single generated value — so for the same seed the result has the same row content as an ordered fill across every deterministic column (physical row order aside).
flowchart TD
P["probe once: disable + re-enable on a borrowed connection"] -->|privilege ok| W
P -->|BulkLoadUnsupportedException| FB["fall back to ordered level-parallel fill"]
subgraph W["single wave — every table at once, no barrier"]
direction LR
T1["worker: disable → fill table → re-enable (finally)"]
T2["worker: disable → fill table → re-enable (finally)"]
T3["worker: disable → fill partition → re-enable (finally)"]
end
The mechanism is database-specific and lives behind two
DatabaseSupport
SPI methods, disableConstraints / enableConstraints, guarded by supportsBulkLoad():
| Support | Mechanism (per session) | Notes |
|---|---|---|
| PostgreSQL | SET session_replication_role = replica → origin | needs a superuser/rds_superuser role; privilege failure raises BulkLoadUnsupportedException |
| MySQL | SET FOREIGN_KEY_CHECKS=0/UNIQUE_CHECKS=0 → 1 | no special privilege |
| CockroachDB | unsupported (supportsBulkLoad() is false) | no session_replication_role; falls back to ordered |
Because these settings are per connection and each worker borrows its own from the pool, the
disable/enable runs inside every worker task, wrapped in a try/finally that restores the session
before the connection returns to the pool — so a constraint-disabled connection never leaks to
other pool users, even if a fill throws. Privilege is probed once up front on a throwaway
connection; if it fails, the engine logs a warning and runs the ordered level-parallel path instead
of fanning out half-disabled. Bulk mode only applies to the DataSource + threads > 1 path; it is
ignored with a warning elsewhere. The default BulkLoadStrategy.ordered() keeps the dependency-ordered
behavior with constraints always enforced.
Reproducibility — deterministic seeds from schema identity
Section titled “Reproducibility — deterministic seeds from schema identity”Bloviate datasets are reproducible across JVM runs, machines, and time — run it twice against
the same schema with the same base seed and you get byte-for-byte identical data. This is not done
by seeding one global Random; it’s done per column.
DatabaseUtils.columnSeed(Column, baseSeed)
derives a stable seed by hashing the column’s identity — its name, table, schema, catalog, JDBC
type name, and ordinal position — and mixing it with the base seed:
int identity = Objects.hash( column.name(), column.tableName(), column.schema(), column.catalog(), column.jdbcType() == null ? null : column.jdbcType().getName(), column.ordinalPosition());return baseSeed * 1_000_003L + identity;Two important properties fall out of this design:
- Run-independence. The seed depends only on schema identity, never on iteration order, hash-map
ordering, or wall-clock time. The JDBC type’s name is hashed (not the enum’s
ordinal()), so the seed is stable even if the enum changes. - FK alignment for free. Because the seed is a pure function of the column, a foreign key column and the primary key it references resolve to the same seed — which is exactly what makes the FK-fidelity trick in the table-fill section work.
Every generator — built-in, registry-supplied, or per-column override — is constructed with this engine-managed seed, so reproducibility holds no matter how a column’s generator was chosen.
The one other input a factory can receive is the fill’s GenerationContext, created once when
DatabaseFiller.fill() starts and handed to every TableFiller (so to every table, partition and worker):
it carries the asOf anchor that relative date windows resolve against. ColumnGeneratorFactory and
GeneratorFactory gained a default create(..., GenerationContext) overload that ignores it, so existing
factories are untouched. A pinned asOf keeps the data a pure function of seed and anchor; wall-clock
time is only read, once, when none is pinned (see
Relative date ranges and asOf).
The per-column seed feeds an IndexedRandom —
a repositionable SplitMix64 stream (the construction behind java.util.SplittableRandom) that the
engine positions to the absolute row index before every cell, so each value is a pure function of
(columnSeed, rowIndex). The seeding architecture is one isolated, deterministically-seeded random
source per column, so output is reproducible and order-independent. (Generators used outside the
fill engine — flat files, caller-constructed — draw from
RandomGenerators.create(seed),
the JDK general-purpose default L64X128MixRandom.)
Because a value depends only on its column’s seed and row index — never on timing, prior rows, or which tables fill alongside it — reproducibility survives concurrency. A parallel table fill (parallel fill) yields the same row content as a sequential one across every deterministic column. An intra-table partitioned fill (intra-table partitioning) is byte-identical to the sequential fill for any partition count across every built-in generator; only non-positionable custom generators vary with the partition count.
Database support — the Strategy pattern
Section titled “Database support — the Strategy pattern”Different databases expose different types. DatabaseSupport
is the strategy interface that maps a Column to a generator; AbstractDatabaseSupport
holds an EnumMap<JDBCType, GeneratorFactory> of cross-database defaults and exposes a configure()
hook that subclasses override to add or replace entries for vendor-specific types.
classDiagram
class DatabaseSupport {
<<interface>>
+getDataGenerator(Column, RandomGenerator) DataGenerator
+forConnection(Connection)$ DatabaseSupport
+forProduct(String)$ DatabaseSupport
}
class AbstractDatabaseSupport {
<<abstract>>
-EnumMap~JDBCType,GeneratorFactory~ registry
#configure(Map) void
}
DatabaseSupport <|.. AbstractDatabaseSupport
AbstractDatabaseSupport <|-- DefaultSupport
AbstractDatabaseSupport <|-- PostgresSupport
AbstractDatabaseSupport <|-- MySQLSupport
AbstractDatabaseSupport <|-- H2Support
AbstractDatabaseSupport <|-- SQLiteSupport
PostgresSupport <|-- CockroachDBSupport
MySQLSupport <|-- MariaDBSupport
| Implementation | Adds on top of the JDBC defaults |
|---|---|
DefaultSupport | Nothing — cross-database JDBC types only |
PostgresSupport | uuid, json/jsonb, inet, cidr, macaddr/macaddr8, interval, bit/varbit, xml, and text/int arrays |
MySQLSupport | JSON columns generate valid JSON instead of arbitrary text; BIT(n) generates an n-bit number, not a bit string |
MariaDBSupport | Extends MySQLSupport — MariaDB columns surface through JDBC essentially as MySQL’s |
CockroachDBSupport | Extends PostgresSupport (CockroachDB is PG wire-compatible), including its CHECK/ENUM constraint reading |
H2Support | UUID (reported as BINARY) and valid JSON |
SQLiteSupport | Nothing — SQLite’s affinity types collapse onto INTEGER/FLOAT/VARCHAR, already covered by the defaults |
You don’t have to pick manually. DatabaseSupport.forConnection(connection) reads
DatabaseMetaData.getDatabaseProductName() and selects the right strategy by substring match,
falling back to DefaultSupport. (Note: CockroachDB reached via the PG driver reports as
PostgreSQL and resolves to PostgresSupport — equivalent, since CockroachDBSupport adds no extra
behavior.)
DatabaseSupport also reads value constraints for a table — DatabaseSupport.readConstraints(...).
The default is none; PostgresSupport queries pg_constraint/pg_enum, parses the common CHECK
forms (IN, BETWEEN, comparisons, first-of-month date_trunc/EXTRACT on dates) and enum labels
into a ColumnConstraint, and TableFiller then prefers a constraint-satisfying generator
(categorical, bounded numeric, or TruncatedDateGenerator) over the type default — so generated
values conform instead of being rejected. The parser recognises only exact forms (a bare, optionally
cast, column against literals; never an expression containing a function call other than the
first-of-month spellings); anything else is logged and skipped.
Pluggable generators — Registry + ServiceLoader
Section titled “Pluggable generators — Registry + ServiceLoader”Sometimes the type-based default isn’t what you want — you want every column named email to look
like an email, regardless of its SQL type. The GeneratorRegistry
lets you override generation without subclassing DatabaseSupport, and external jars can
contribute rules automatically via Java’s ServiceLoader.
A registry supports three matcher kinds, and the fill engine resolves each column through a fixed precedence chain:
flowchart TD
Col[Column to fill] --> CC{per-column<br/>ColumnConfiguration?}
CC -->|yes| Use1[use it]
CC -->|no| NP{registry name<br/>pattern match?}
NP -->|yes| Use2[use it]
NP -->|no| TN{registry vendor<br/>typeName match?}
TN -->|yes| Use3[use it]
TN -->|no| JT{registry<br/>JDBCType match?}
JT -->|yes| Use4[use it]
JT -->|no| Def["DatabaseSupport default"]
Plugins implement the single-method GeneratorPlugin
SPI and declare themselves in META-INF/services/io.bloviate.ext.GeneratorPlugin. Calling
GeneratorRegistry.Builder.discover() loads every plugin on the classpath:
GeneratorRegistry registry = new GeneratorRegistry.Builder() .registerColumnNamePattern(".*email", (column, random) -> new EmailGenerator.Builder(random).build()) .registerTypeName("uuid", (column, random) -> new UUIDGenerator.Builder(random).build()) .discover() // pick up GeneratorPlugin services from the classpath .build();Crucially, registry- and plugin-supplied generators are still constructed with the engine’s seeded
RandomGenerator, so they remain just as reproducible as the built-ins. The optional
bloviate-datafaker module is exactly this pattern in practice: one
GeneratorPlugin that maps column names (email, first_name, phone, …) to realistic
Datafaker values, seeded from the engine’s column seed for
reproducibility — keeping Datafaker out of the core. It also offers referential realism via a
RowContext: correlated columns project fields of one per-row entity — a Person whose email/username
derive from its name, or a Geo tuple (city/state/zip/area-code that agree) drawn from a bundled
reference dataset — computed as a pure function of (seed, rowIndex) so consistency survives parallel
and partitioned fills.
The generator library — Builder pattern
Section titled “The generator library — Builder pattern”Every generator implements DataGenerator<T>,
which can generate() a typed value, generateAsString() for flat files, bind itself to a
PreparedStatement (generateAndSet), and read a value back from a ResultSet (get). Each is
constructed through a static inner Builder seeded with a RandomGenerator:
new SimpleStringGenerator.Builder(random).size(100).build();new BigDecimalGenerator.Builder(random).precision(10).digits(2).build();The library ships ~50 generators in io.bloviate.gen,
covering everything from primitives and dates to uuid, jsonb, inet/cidr, MAC addresses,
intervals, arrays, and XML. A few are worth calling out:
- Referential-fidelity generators —
CompositeKeyComponentGenerator,ChildKeyComponentGenerator, andChildCountGeneratorproduce collision-free composite keys and variable parent/child cardinalities that keep foreign keys consistent. GroupedPermutationGenerator— emits a deterministic permutation per group (e.g. TPC-C’s shuffledo_c_id) using a Feistel network with cycle-walking, achieving a unique pseudo-random permutation in O(1) memory without materializing or shuffling an array.- TPC-C generators —
io.bloviate.gen.tpccprovides benchmark-faithful fields (customer last names, zip codes, credit, delivery dates). - Distribution generators —
WeightedCategoricalGenerator,NormalDoubleGenerator/NormalIntegerGenerator,ZipfianIntegerGenerator, andSkewedTimestampGeneratoremit non-uniform values (categorical, bounded-Gaussian, power-law, recency-skewed). A column opts in through theDistributionsconvenience without writing a factory. These are specified distributions, not learned from data — so they stay deterministic by seed and compose with FK reseeding and parallel fills like any other generator.
Flat-file generation
Section titled “Flat-file generation”The same generators power FlatFileGenerator,
which needs no database at all. You describe columns with ColumnDefinition
(name + generator), pick a FileType
(CSV, TDV, or PIPE), and generate(). Output is written via
Apache Commons CSV, with headers derived from
column names:
new FlatFileGenerator.Builder("output/users") .add(new ColumnDefinition("id", new IntegerGenerator.Builder(random).build())) .add(new ColumnDefinition("email", new SimpleStringGenerator.Builder(random).build())) .rows(1000) .build() .generate();Multi-module layout
Section titled “Multi-module layout”Bloviate is a Maven reactor. The engine is self-contained; the integration modules pull in their
framework as a provided dependency so you bring your own version.
graph TD
core["bloviate-core<br/><i>self-contained engine</i><br/>db · ext · gen · file · util"]
junit["bloviate-junit<br/><i>@FillDatabase / @FillSource</i><br/>BloviateExtension"]
tc["bloviate-testcontainers<br/><i>BloviateContainers</i>"]
df["bloviate-datafaker<br/><i>DatafakerGeneratorPlugin</i>"]
junit -->|depends on| core
tc -->|depends on| core
df -->|depends on| core
junit -.->|provided| j[JUnit Jupiter]
tc -.->|provided| t[Testcontainers]
df -->|depends on| d[Datafaker]
bloviate-core— everything above: introspection, the DAG, generators, database support, flat files.bloviate-junit— declarative test-data filling. Annotate a test (class or method) with@FillDatabaseand mark aDataSource/Connectionfield with@FillSource;BloviateExtension(a JUnitBeforeEachCallback) auto-detects the rightDatabaseSupportand fills before each test.bloviate-testcontainers— fill a startedJdbcDatabaseContainerin one fluent call viaBloviateContainers.forContainer(...).bloviate-datafaker— optional semantic/realistic values by column name, via aGeneratorPluginover Datafaker. Unlike the others, Datafaker is a normal (notprovided) dependency — it only reaches your classpath if you add this module.
Design patterns at a glance
Section titled “Design patterns at a glance”Bloviate leans deliberately on a small, well-understood set of design patterns. They aren’t applied for their own sake — each one buys a concrete property (extensibility, immutability, testability) and they compose cleanly. If you’ve read the Gang of Four, this table is a fast map of where each pattern lives and what it’s doing for us.
| Pattern | Where it lives | What it buys |
|---|---|---|
| Builder | Nearly everything: DatabaseFiller.Builder, TableFiller.Builder, every *Generator.Builder, GeneratorRegistry.Builder, FlatFileGenerator.Builder, BloviateContainers.Builder | Readable construction of objects with many optional, defaulted parameters; immutable results with no telescoping constructors |
| Strategy | DatabaseSupport + per-database implementations | Database-specific behavior is swappable at runtime; adding a database means adding a class, not editing the engine |
| Template Method | AbstractDatabaseSupport seeds defaults, then calls the configure() hook | Subclasses customize only the vendor-specific slice; the invariant default registry is defined once |
| Factory (functional) | GeneratorFactory, ColumnGeneratorFactory | Defers generator creation until the engine can supply a column-seeded RandomGenerator, preserving reproducibility |
| Registry | GeneratorRegistry with documented precedence | Override generation by name/type without subclassing; rules resolved in a fixed, predictable order |
| Service Provider (SPI) | GeneratorPlugin via ServiceLoader | Third-party jars contribute generators by dropping a file on the classpath — zero engine coupling |
| Strategy / polymorphism | The DataGenerator<T> hierarchy; CommitStrategy for transaction cadence | One uniform interface (generate / bind / read-back) over ~50 type-specific implementations; swappable commit behavior |
| Seekable cursor | IndexedDataGenerator.seek(long) on the positional generators | Position a generator at an absolute row index without replaying — the basis for intra-table partitioning |
| Iterator | JGraphT’s TopologicalOrderIterator in DatabaseFiller | Fill order is expressed as a traversal, decoupled from graph construction |
| Adapter | bloviate-junit and bloviate-testcontainers | Wrap the same core engine behind framework-native front-ends (@FillDatabase, BloviateContainers) |
The payoff is that the two extension axes you actually care about — “support a new database” and “generate a new kind of value” — are both open for extension without modifying a line of the core engine (Open/Closed Principle). Strategy handles the first; Registry + SPI handle the second.
Design principles
Section titled “Design principles”A few themes recur throughout the codebase and explain most of the “why” — they’re the result of deliberate design effort, not incidental:
- Read, don’t configure. Schema is introspected from JDBC metadata, not hand-described.
- Immutability. The metadata and configuration models are Java records; registries are built once and copied defensively.
- Determinism by construction. Seeds derive from schema identity, never runtime state — so data is reproducible and foreign keys align without bookkeeping.
- Open for extension, closed for modification. The Strategy pattern (
DatabaseSupport) handles new databases; the Registry + ServiceLoader (GeneratorPlugin) handles new generation rules — neither requires touching the engine. - One engine, many front-ends. Database filling, flat files, JUnit, and Testcontainers all sit
on the same
bloviate-coregenerators.
What this buys in practice
Section titled “What this buys in practice”These principles aren’t abstract — they show up as concrete quality properties you can rely on:
- A lean core.
bloviate-corekeeps its dependency surface small and pushes JUnit and Testcontainers toprovidedscope, so integrating Bloviate doesn’t drag a testing framework into your runtime classpath, and you bring your own versions. - Tested against real databases. Integration tests run against actual PostgreSQL, MySQL, MariaDB,
and CockroachDB instances via Testcontainers — plus embedded H2 and SQLite — not mocks — over real benchmark schemas (TPC-C,
AuctionMark, Wikipedia). Behavior is verified end-to-end, including the FK ordering and round-trip
read-back through each generator’s
get(ResultSet, ...). - Reproducibility as a guarantee, not a hope. Because seeds are pure functions of schema identity, “it worked on my machine” datasets are byte-for-byte portable — which is exactly what you want from test fixtures and benchmark data.
- Modern Java, used deliberately. Records for the immutable model, sealed extension points,
functional SPIs (
@FunctionalInterface), andEnumMap-backed registries reflect a codebase built on Java 25 idioms rather than retrofitted onto them. - Thorough documentation. The public types carry real Javadoc — including the precedence rules for generator resolution and the CockroachDB/PostgreSQL driver caveat — so the contracts are written down, not folklore.
The throughline: Bloviate is a small library that takes its own design seriously. The patterns and principles above are what let it stay simple to use while remaining genuinely extensible underneath.
See also: Quick Start to start filling databases.