Configuration
By default Bloviate inspects the schema and picks a generator for every column based on its JDBC type. This guide shows how to take progressively more control — from overriding row counts, to overriding individual columns, to non-uniform distributions, reproducible seeds, and the parallel fill path.
Configuration is layered:
DatabaseConfiguration— global defaults: batch size, default row count, database support, and an optional set of per-table overrides.TableConfiguration— overrides the row count for one table, and optionally carries per-column overrides.ColumnConfiguration— overrides how a single column is generated, via aColumnGeneratorFactory(aRandom -> DataGenerator<?>lambda). The engine hands the factory a column-seededRandomso output stays reproducible.
Per-table row counts
Section titled “Per-table row counts”Generate different numbers of rows for specific tables while every other table uses the default:
import io.bloviate.db.*;import io.bloviate.ext.PostgresSupport;import java.util.Set;
// "users" gets 50 rows; all other tables fall back to the default (100)Set<TableConfiguration> tableConfigs = Set.of( new TableConfiguration("users", 50));
DatabaseConfiguration config = new DatabaseConfiguration( 10, // batch size 100, // default rows per table new PostgresSupport(), // database support tableConfigs // table-specific overrides);
new DatabaseFiller.Builder(connection, config) .build() .fill();Per-column generation overrides
Section titled “Per-column generation overrides”Pin a specific column to a custom generator. Here the status_code column on orders is
constrained to integers in [1, 10) instead of the type’s default:
import io.bloviate.db.*;import io.bloviate.ext.PostgresSupport;import io.bloviate.gen.IntegerGenerator;import java.util.Set;
// ColumnGeneratorFactory is a `RandomGenerator -> DataGenerator<?>` lambda.// The engine supplies a column-seeded RandomGenerator for reproducible output.Set<ColumnConfiguration> columnConfigs = Set.of( new ColumnConfiguration("status_code", random -> new IntegerGenerator.Builder(random).start(1).end(10).build()));
// 1,000 rows for "orders", with the column override appliedSet<TableConfiguration> tableConfigs = Set.of( new TableConfiguration("orders", 1000, columnConfigs));
DatabaseConfiguration config = new DatabaseConfiguration( 128, 100, new PostgresSupport(), tableConfigs);
new DatabaseFiller.Builder(connection, config) .build() .fill();Column names are matched case-insensitively. Any column without an override keeps its default, type-based generator.
Value distributions
Section titled “Value distributions”Real columns are rarely uniform — a status is mostly ACTIVE, a rating clusters around its
mean, a referenced product_id follows a popularity curve, and a created_at bunches toward the
present. The Distributions helper returns ready-made ColumnGeneratorFactory values so a column
can opt into a non-uniform distribution without writing a factory:
import io.bloviate.db.*;import io.bloviate.ext.PostgresSupport;import java.util.Map;import java.util.Set;
Set<ColumnConfiguration> columnConfigs = Set.of( // 70% NEW, 25% SHIPPED, 5% CANCELLED (weights need not sum to 1) new ColumnConfiguration("status", Distributions.weighted(Map.of("NEW", 0.7, "SHIPPED", 0.25, "CANCELLED", 0.05))), // normal(mean=4, sd=1) rounded and clamped to [1, 5] new ColumnConfiguration("rating", Distributions.normalInt(4, 1, 1, 5)), // Zipfian (power-law) over [1, 10000] — a few hot ids, a long thin tail new ColumnConfiguration("product_id", Distributions.zipfian(10_000)), // timestamps skewed toward the recent end of the window new ColumnConfiguration("created_at", Distributions.recentTimestamps()));
DatabaseConfiguration config = new DatabaseConfiguration( 128, 100, new PostgresSupport(), Set.of(new TableConfiguration("orders", 100_000, columnConfigs)));Available shapes: weighted(...) (categorical), normal(...) / normalInt(...) (bounded
Gaussian), zipfian(...) (power-law), and recentTimestamps(...) (recency-skewed). Each is built
from the engine’s column seed, so output stays reproducible and composes with foreign-key
reseeding and parallel fills like any other generator. These are specified distributions, not
distributions learned from real data.
Constraint conformance
Section titled “Constraint conformance”On PostgreSQL and CockroachDB, Bloviate reads each table’s CHECK constraints and ENUM
types and generates values that satisfy them — automatically, no configuration. So given:
CREATE TYPE order_status AS ENUM ('NEW', 'PAID', 'SHIPPED', 'CANCELLED');
CREATE TABLE orders ( status order_status NOT NULL, rating integer CHECK (rating BETWEEN 1 AND 5), priority integer CHECK (priority IN (1, 2, 3)), amount numeric(8,2) CHECK (amount >= 0 AND amount <= 9999.99));status only gets one of its enum labels, rating lands in [1, 5], priority is one of
1/2/3, and amount stays in range — instead of random values an insert would reject. The common
forms are honored: IN (...), BETWEEN, and >=/<=/>/< comparisons, for integer, numeric,
floating, and text columns, plus enum/domain allowed values.
Dates that must fall on the first of a period are honored too, on DATE, TIMESTAMP and
TIMESTAMP WITH TIME ZONE columns (since 3.5.0):
CREATE TABLE invoices ( billing_month date CHECK (date_trunc('month', billing_month) = billing_month), period_start timestamptz CHECK (EXTRACT(day FROM period_start) = 1));date_trunc('month' | 'quarter' | 'year', col) = col and EXTRACT(day FROM col) = 1 get a
TruncatedDateGenerator: the first day of a random month in
a fixed window (2015 up to 2025 by default; the same seed gives the same data), or midnight on that
day for a timestamp. A timestamptz value is midnight in the session time zone, which is the zone
the check is evaluated in, so the fill works whatever the connection’s zone is.
Notes:
- The parser only recognises the exact forms above: a bare column, optionally cast, compared with
literals. A
CHECKthat calls a function on the column (lower(status) IN (...),length(name) >= 1), does arithmetic, or takes any other form is skipped with a warning, and the column falls back to its type default. That includes negation,OR,LIKEpatterns, a one-sided bound, otherdate_truncunits (day,week, …), andCHECKs over more than one column. - Before 3.5.0 a quoted function argument was read as an allowed value, so
date_trunc('month', d) = dmade the fill fail withinvalid input syntax for type date: "month". - Before 3.9.0 no
CHECKorENUMwas read on CockroachDB, so those columns were filled from their type default. Most such fills failed outright — an enum column got an arbitrary string, and a range orINcheck rejected the value — but a column whose check the type default happened to satisfy did fill, and the values it produces change in 3.9.0, because they now come from the constraint. Re-pin any CockroachDB fixture you compare byte-for-byte. - A per-column override or a registry rule always wins, so you can still take full control of a constrained column.
- Open the connection with
stringtype=unspecified(already required for PostgreSQL’s extension types) so enum/INvalues bind. This applies to CockroachDB too, which is reached through the same driver. - CockroachDB is covered by the same reader (since 3.9.0): it serves the same
pg_catalogqueries, and the definitions it stores differ only in spellings the parser accepts on both — it keepsBETWEENverbatim where PostgreSQL expands it into two comparisons, writesextract('day', col)with a comma rather thanEXTRACT(day FROM col), and casts literals as::STRING. Every other database reads noCHECKs, so give those columns an explicit generator.
Foreign keys and the values they reference
Section titled “Foreign keys and the values they reference”A foreign-key column is not checked against the parent’s rows — it replays them. Bloviate generates it from the same seed at the same row index as the key it points at, so the values match without reading anything back. Three things follow, and they are worth knowing before you read generated data.
It follows the columns the constraint names. A key to a UNIQUE target works, not just one to the
parent’s primary key. The tenant pattern is the usual case:
create table accounts ( id int primary key, tenant_id int not null references tenants (id), unique (tenant_id, id));
create table documents ( tenant_id int not null, account_id int not null, foreign key (tenant_id, account_id) references accounts (tenant_id, id), foreign key (tenant_id) references tenants (id));A column in several keys makes those keys equal. documents.tenant_id above belongs to two keys at
once, so its value has to exist in both. Bloviate guarantees that by generating every column linked by
a foreign key — in either direction — from one shared seed. For an ordinary parent/child chain nothing
changes. But point one column at two unrelated keys and those two key columns come out carrying
identical values, row for row. That is not a quirk to work around: a value cannot be in two key spaces
unless the key spaces overlap, and equality is the only overlap reproducible without reading rows back.
If you did not mean the keys to be interchangeable, the schema is telling you something.
Row counts bound it. A foreign-key column cycles through its parent’s keys once it runs past them,
so a child with more rows than its parent is fine — it reuses parents. Row i of a child reads row
i % parentRows of its parent, which in turn reads row (i % parentRows) % grandparentRows of the
grandparent, all the way up. The counts come from each table’s TableConfiguration where it has one
and from the default row count otherwise.
Keys the database generates. A serial or IDENTITY key column is left out of its own insert, so
the database assigns it and there is no seed to share. Bloviate counts instead: a table filled from
empty is given 1..N in insertion order, and the child counts through the same range. This assumes the
parent’s sequence starts at 1 — filling a table that already holds rows, or whose sequence has been
advanced, leaves the child pointing at keys the parent was not given. Fill into an empty schema, or
reset the sequence first (TRUNCATE ... RESTART IDENTITY). Because the class carries one set of
values, a key linked to a generated one is filled by counting too, so both ends match.
One shape this cannot cover: a generated column inside a composite key whose parent is filled with intra-table partitions. The database hands out identities in completion order across the workers, while the key’s other columns are generated from the logical row index, so the two need not describe the same parent row. Give such a parent an ordinary key column, or fill it without partitions.
Foreign-key cycles
Section titled “Foreign-key cycles”Two or more tables that reference each other — directly, or through a chain — cannot be filled in any order: whichever goes first, its foreign key has no parent row to point at. A fill fails before writing anything, naming every cycle it found:
cannot fill: the tables [invoice, payment] reference each other, so no fill order satisfies them: ...Three ways forward:
- break the cycle in the schema;
- leave every table of the cycle out with
excludeTables(...)— dropping only one side leaves the others referencing a table that is not being filled, which fails for that reason instead; - fill with
BulkLoadStrategy.unorderedBulk(), which fills every table at once without enforcing constraints, so it needs no order. Two things have to hold for it to be chosen: aDataSourcewiththreads(n)above 1 (on a single connection the fill is sequential and the strategy is only warned about), and a support that implements bulk loading — PostgreSQL, MySQL, MariaDB and BigQuery. Anywhere else it falls back to the ordered path and fails the same way.
A table referencing itself is not a cycle: it constrains the order of rows within that one table, not the order of tables. Those fill normally, with a warning, and making a row’s parent exist before its child is up to you.
Reproducible data with seeds
Section titled “Reproducible data with seeds”DatabaseConfiguration takes a base seed. The same schema filled with the same seed always
produces identical data, so test fixtures are deterministic; change the seed for a different — but
still reproducible — dataset. Per-column seeds are derived from stable column identity, and foreign
keys share the seed of the key they reference (see
foreign keys and the values they reference), so
referential fidelity holds for any seed.
import io.bloviate.db.*;import io.bloviate.ext.PostgresSupport;
// batch size, rows/table, support, table configs, seedDatabaseConfiguration config = new DatabaseConfiguration( 128, 100, new PostgresSupport(), null, 42L);
new DatabaseFiller.Builder(connection, config).build().fill();The seed defaults to 0 when you use the four-argument constructor, so existing code keeps a
single, stable dataset without changes.
Time zones do not enter into it. A java.sql.Date, Time or Timestamp is an instant, and
turning one into the value a column holds needs a zone. Bloviate binds every temporal through an
explicit UTC calendar, so the stored value is the same on every machine: a DATE, TIME,
TIMESTAMP or DATETIME column gets the instant’s UTC wall clock, and a column that carries a zone
(TIMESTAMP WITH TIME ZONE, timestamptz) gets the instant itself. Before 3.9.1 the zone was left
to the driver, which used the JVM’s default, so the same seed produced values five hours apart on a
machine in America/New_York and one in UTC (issue #640). If your JVM was not already running in
UTC, the values a seed produces change in 3.9.1 — they stop following the machine. A JVM already on
UTC, which is the usual container and CI setup, is unaffected.
One thing a client cannot pin: MySQL’s TIMESTAMP is an instant that the server converts from the
session time zone on write and back to it on read. Its stored value therefore follows the session
zone by definition, whatever a client sends. DATE, TIME and DATETIME carry no zone and do not
move. Set the session zone explicitly (a before SQL hook running
SET time_zone = '+00:00') if you need MySQL TIMESTAMP columns pinned too.
A custom generator that binds a java.sql temporal itself should go through
io.bloviate.gen.TemporalBinding, which is the one place this zone rule lives; the
java.time-based generators (TruncatedDateGenerator) are already zone-independent.
Relative date ranges and asOf
Section titled “Relative date ranges and asOf”Date and timestamp generators default to a fixed window around 2020-01-01, so on their own they miss
“this quarter” partitions and range CHECKs. A relative window says where the data should fall in
terms of a moving anchor, asOf, instead of absolute dates: “within the last 90 days”, “from 30 days
ago to a week from now”. (Absolute bounds, start(...)/end(...), still work and are unchanged.)
import io.bloviate.db.*;import io.bloviate.ext.PostgresSupport;import io.bloviate.gen.*;import java.time.Instant;import java.util.Set;
// orders may only be placed in Q1 2026: CHECK (placed_at >= '2026-01-01' AND placed_at < '2026-04-01')Set<ColumnConfiguration> columns = Set.of( ColumnConfiguration.relative("placed_at", RelativeWindow.withinLast("90d"), (random, window) -> new SqlTimestampGenerator.Builder(random).window(window).build()), ColumnConfiguration.relative("due_on", RelativeWindow.between("-30d", "+7d"), (random, window) -> new SqlDateGenerator.Builder(random).window(window).build()));
DatabaseConfiguration config = new DatabaseConfiguration.Builder(128, 1_000, new PostgresSupport()) .seed(42L) .tableConfigurations(Set.of(new TableConfiguration("orders", 1_000, columns))) .build();
new DatabaseFiller.Builder(connection, config) .asOf(Instant.parse("2026-04-01T00:00:00Z")) // pin the anchor: same seed + same asOf = same rows .build() .fill();withinLast("90d") with asOf 2026-04-01 is the 90 days from 2026-01-01 up to (not including)
2026-04-01, so every placed_at satisfies the CHECK above, in the same rows every time.
The reproducibility rule
Section titled “The reproducibility rule”Wall-clock time is never an input to generation. The one place a date can enter is asOf, and it is
explicit:
- Pinned (
DatabaseFiller.Builder.asOf(Instant)): the same seed and the sameasOfproduce byte-identical output on every run and every JDK. This is the sanctioned, reproducible way to get “current” data. - Not pinned: when the fill starts,
asOfis read once from the clock and truncated to the start of the current UTC day (00:00Z). It is logged once at INFO the first time a relative window uses it (resolved asOf 2026-06-10T00:00:00Z ...; pin it with asOf(...) for reproducible output) and is available fromDatabaseFiller.asOf(). Two unpinned fills on the same UTC day agree; across midnight they do not. To reproduce such a fill, pin the instant it logged.
One asOf serves the whole fill: every table, every intra-table partition and every worker thread of a
parallel fill uses the same instant, so the tables of one dataset are consistent with each other. A
fill that uses no relative window is unaffected by asOf: its output is exactly what it was, and
the same seed still gives the same data with or without one.
Offsets and windows
Section titled “Offsets and windows”An offset is an optional sign, digits and one unit letter (case-sensitive):
| Unit | Meaning | Arithmetic (all in UTC) |
|---|---|---|
h | hours | exact 3,600 seconds |
d | days | exact 24 hours (UTC has no daylight saving) |
w | weeks | exact 7 days |
M | months | calendar: the day of month is clamped, so 2026-03-31 minus 1M is 2026-02-28 |
y | years | calendar: 2024-02-29 plus 1y is 2025-02-28 |
Examples: 90d, -30d, +7d, 12w, 6M, 1y, 36h. There is no minute or second unit; a lower-case
m is rejected (months are M). Whitespace inside an offset is an error.
| Window | Meaning | Resolved for asOf = 2026-04-01 |
|---|---|---|
RelativeWindow.withinLast("90d") | [asOf - 90d, asOf); the span must be positive | [2026-01-01, 2026-04-01) |
RelativeWindow.between("-30d", "+7d") | [asOf + start, asOf + end) | [2026-03-02, 2026-04-08) |
RelativeWindow.withinLast("6M") | six calendar months back | [2025-10-01, 2026-04-01) |
A window is half-open: the start is inclusive and the end exclusive, the same convention as
SqlTimestampGenerator, SqlDateGenerator, DateGenerator, InstantGenerator and
TruncatedDateGenerator. So withinLast never produces the anchor instant itself; use
between("-90d", "1d") to include asOf’s whole day. Bad input fails fast with an
IllegalArgumentException that names the problem: an unknown unit, a zero or negative withinLast, an
end that is not after the start. RelativeOffset.parse(...), RelativeWindow.withinLast(String)
and RelativeWindow.between(String, String) take the text forms, so a configuration file or command
line can hand them the same strings.
Applying a window
Section titled “Applying a window”ColumnConfiguration.relative(column, window, factory) resolves the window against the fill’s asOf
and hands the concrete bounds to your factory. Those generators accept one through
window(RelativeWindow.Resolved): SqlTimestampGenerator, SqlDateGenerator, DateGenerator,
InstantGenerator, SkewedTimestampGenerator and TruncatedDateGenerator. Distributions.recentTimestamps(window, skew)
is the recency-skewed shorthand. Anything else that needs the anchor takes the fill’s
GenerationContext: ColumnGeneratorFactory.contextual((random, context) -> ...), or, for a
registry rule, GeneratorFactory.contextual((column, random, context) -> ...),
where context.asOf() is the anchor. Existing factories (plain lambdas) are unchanged and ignore the
context.
- A relative window is the way to keep the key of a partitioned table inside its existing partitions; configure it on the partitioned parent, which is what Bloviate fills (see Keeping the partition key inside existing partitions).
- A per-column relative configuration takes precedence over a
CHECKconstraint, a registry rule and the support default, like any other per-column override. SqlDateGeneratorreads its window as whole UTC calendar days, so aDATEcolumn holds exactly the dates in the window whatever the JVM or session time zone. A window a few hours wide inside one day contains no whole date and is rejected.TruncatedDateGeneratordraws the period starts in the window (its bounds rounded up to whole dates).- For a zone-less
TIMESTAMPcolumn the JDBC driver renders the instant in the JVM’s default time zone, as it does for absolute bounds. If a hard bound (aCHECKor a partition edge) must hold and the JVM may run in another zone, useTIMESTAMP WITH TIME ZONEor start the JVM with-Duser.timezone=UTC. SqlTimeGeneratorhas no window: aTIMEcolumn has no date to be relative to.
Parallel table fill
Section titled “Parallel table fill”For large, wide schemas the fill can run in parallel. Construct the filler from a pooled
DataSource instead of a single Connection and ask for more than one worker thread:
import io.bloviate.db.*;import io.bloviate.ext.PostgresSupport;import javax.sql.DataSource;
DataSource dataSource = /* a pooled DataSource, e.g. HikariCP */;
DatabaseConfiguration config = new DatabaseConfiguration( 1000, 100_000, new PostgresSupport(), null, 42L);
new DatabaseFiller.Builder(dataSource, config) .threads(8) // fill independent tables concurrently .build() .fill();Bloviate groups tables into topological levels by their foreign keys and fills the independent tables within each level concurrently, one connection per worker, barriering between levels so a child table is never filled before its parent. Each worker fills its table in a single transaction (commit once per table). The fill stays fully reproducible: a table’s data depends only on its own seed and row order, never on which tables fill alongside it, so the same config and seed produce the same row content as a sequential fill across every deterministic column (physical row order and wall-clock columns aside).
How much this helps depends on the schema. A wide schema of independent tables sees a large speedup (~3× with 8 workers on a 10-table, 1M-row fixture); a deep, narrow foreign-key chain (each table depending on the previous) has little to parallelize. See the benchmarks for numbers.
The single-Connection constructor is unchanged and remains the default sequential path — threads
only applies to the DataSource form.
Intra-table partitioning
Section titled “Intra-table partitioning”When a single large table dominates the fill, between-table parallelism can’t help it — it sits
alone in its topological level. Set partitions on that table’s TableConfiguration to split its
rows into that many contiguous ranges filled concurrently, one connection per range, on the
parallel (DataSource + threads) path:
// split the one big table into 8 row ranges; ignored on the single-Connection pathSet<TableConfiguration> tables = Set.of( new TableConfiguration("events", 50_000_000L, 8 /* partitions */));
DatabaseConfiguration config = new DatabaseConfiguration( 1000, 0, new PostgresSupport(), tables, 42L);
new DatabaseFiller.Builder(dataSource, config).threads(8).build().fill();Partitioning is byte-identical to a sequential fill of the same seed, for any partition count:
every built-in generator derives each value as a pure function of its column seed and the absolute
row index (per-index derivation), so keys, foreign keys, and plain random columns alike land on
exactly the values the sequential fill produces, foreign-key validity always holds, and seeking a
worker to its starting row is O(1) no matter how large the table or its parents are. The only
exceptions are custom generators that opt out of per-row positioning (DataGenerator.positionable()
returning false, as the datafaker integration does because its values come from an internal Faker
RNG) — those stay deterministic for a given partition count but may differ across partition counts.
Size the connection pool for the total concurrent demand (threads, where a partitioned table
counts as partitions units). One case is unsupported: partitioning a parent table whose
primary key comes from a non-positionable custom generator referenced by a foreign key can orphan
those references — partition the child table instead, or use the positional key generators (as the
bundled TPC-C/TPC-H configurations do). A custom generator with internal positional state must
implement io.bloviate.gen.IndexedDataGenerator to stay aligned under partitioning.
Commit strategy
Section titled “Commit strategy”By default the engine leaves the connection’s autocommit untouched (a typical autocommit connection
commits per executeBatch()). Disabling autocommit and committing less often cuts overhead. Pass a
CommitStrategy to DatabaseConfiguration for the sequential path:
import io.bloviate.db.*;
// commit once per table (autocommit off for the fill, restored afterward)DatabaseConfiguration perTable = new DatabaseConfiguration( 1000, 100_000, new PostgresSupport(), null, 42L, CommitStrategy.perTable());
// or bound the open transaction: commit every 50 JDBC batchesDatabaseConfiguration everyN = new DatabaseConfiguration( 1000, 100_000, new PostgresSupport(), null, 42L, CommitStrategy.everyNBatches(50));The default, CommitStrategy.connectionDefault(), preserves today’s behavior (the engine never
touches autocommit). The parallel path already commits once per table; a configured strategy
applies there too.
Tip — driver batch rewrite. Bloviate inserts in JDBC batches, but most drivers only collapse a batch into a single multi-row
INSERTwhen you opt in via the JDBC URL: PostgreSQLreWriteBatchedInserts=true, MySQLrewriteBatchedStatements=true. Enabling it is often the single biggest fill speedup, sequential or parallel. Bloviate logs a warning at fill time when the parameter is missing, andio.bloviate.util.JdbcUrls.withBatchRewrite(url, support.batchRewriteUrlParameter())builds a correctly-parameterized URL if you construct theDataSourceyourself. CockroachDB ignores the parameter, so no warning is emitted there.
Tip — BigQuery. Rewriting is unconditional in tbc-bq-jdbc, so there is no parameter to set and no warning to emit. Instead, size
batchSizeagainst BigQuery’s limit of 10,000 query parameters per query: the effective rows per job is10_000 / columnCount, so the default 128 leaves most of a job unused. Start frombatchSize = 5000, and set the driver’sbatchLoadThresholdto match to move large batches onto its NDJSON load-job path. LeaveCommitStrategyat its default —BigQuerySupportopts out of engine-managed transactions becausesetAutoCommit(false)starts a BigQuery session and silently disables that load path, and eachexecuteBatchis already one atomic job. Bloviate logs a warning if a commit strategy is configured anyway.
Bulk load (unordered fill)
Section titled “Bulk load (unordered fill)”The parallel path normally barriers between topological levels, so a deep, narrow foreign-key
chain (each table depending on the previous) serializes — there is little within any one level to
run concurrently. BulkLoadStrategy.unorderedBulk() removes that barrier: it disables foreign-key
enforcement, fills every table at once, then re-enables enforcement.
import io.bloviate.db.*;import io.bloviate.ext.PostgresSupport;
DatabaseConfiguration config = new DatabaseConfiguration( 1000, 100_000, new PostgresSupport(), null, 42L, null, // CommitStrategy (null = default) BulkLoadStrategy.unorderedBulk()); // disable constraints, fill barrier-free, re-enable
new DatabaseFiller.Builder(dataSource, config).threads(8).build().fill();This is safe because Bloviate’s data is referentially consistent by construction: a foreign-key
column is seeded from its referenced primary-key column, so child and parent generate identical key
values regardless of insert order. Disabling enforcement therefore changes nothing about
correctness — for the same seed the result has the same row content as an ordered fill across every
deterministic column (physical row order aside) — it only
removes the ordering constraint and the per-row foreign-key checks. The win is largest on deep
chains (e.g. TPC-C’s warehouse → district → customer → open_order → order_line); wide, FK-free
schemas already saturate their workers in one level and see little change.
Requirements and fallback:
- Only effective on the parallel path (a
DataSourcewiththreads > 1); it is ignored with a warning on the single-Connectionand single-thread paths, which fill in dependency order. - Supported on PostgreSQL (
SET session_replication_role = replica, which needs a superuser/rds_superuserrole) and MySQL (SET FOREIGN_KEY_CHECKS=0/UNIQUE_CHECKS=0, no special privilege). CockroachDB does not support it and transparently falls back to the ordered level-parallel path. - Each worker disables enforcement on its own pooled connection and restores it in a
finallybefore returning the connection to the pool, so no connection ever leaks back with checks suppressed. If enforcement cannot be disabled (e.g. the role lacks privilege), the engine logs a warning and falls back to the ordered path rather than running half-disabled.
The default, BulkLoadStrategy.ordered(), preserves today’s behavior (dependency-ordered, constraints
always enforced).
SQL hooks
Section titled “SQL hooks”DatabaseFiller.Builder can run SQL scripts around the fill: before hooks run before any table
is filled, after hooks run once every table is filled. The classic use is a derived table
computed from the generated rows (a rollup, a summary), so it cannot disagree with its source;
another is emptying tables so a fill can be re-run.
import io.bloviate.db.*;import java.nio.file.Path;import java.util.Map;
new DatabaseFiller.Builder(connection, config) .before(SqlScript.resource("sql/reset.sql")) // classpath resource .after(SqlScript.file(Path.of("sql/rollup.sql")) // file on disk .withTokens(Map.of("schema", "reporting"))) // ${schema} in the script .after(SqlScript.inline("analyze", "ANALYZE orders_by_day")) // inline text .build() .fill();-- sql/rollup.sqlDELETE FROM ${schema}.orders_by_day;INSERT INTO ${schema}.orders_by_day (day, orders, revenue)SELECT order_date, count(*), sum(total) FROM orders GROUP BY order_date;before(...) and after(...) are additive and ordered: call them as often as you like and the
scripts run in the order added. A SqlScript is a name (used in log lines and error messages) plus
its text, from SqlScript.inline(name, sql), SqlScript.file(path) or SqlScript.resource(name).
SqlScriptRunner.run(connection, script) runs one on its own, for example to load a schema.
What a script may contain. Statements are separated by ;. A ; inside a 'string' (with ''
as the escape, and \' inside E'...'), a "quoted" or `backtick` identifier, a -- or
/* */ comment (which nests, as in PostgreSQL), or a PostgreSQL dollar-quoted body ($$...$$,
$tag$...$tag$, so CREATE FUNCTION bodies work) does not end a statement. Empty statements are
skipped and the last statement needs no ;. ${name} is replaced from the script’s tokens,
everywhere except inside comments; a token with no value fails the script before any statement runs.
MySQL’s DELIMITER command is not supported and fails with a clear message rather than
mis-splitting the script, and backslash escapes are honoured only in E'...' strings.
Where hooks run. Before hooks run first, ahead of the schema read, so they can create the tables
about to be filled. After hooks run last. On the single-Connection path both run on the connection
you supplied. On a DataSource without threads(n) (or with threads(1)) Bloviate borrows one
connection and uses it for the before hooks, the fill and the after hooks, so session state a hook
sets (SET search_path, a temporary table) carries through, and the after hooks see the fill’s own
uncommitted rows even if the pool hands out autoCommit=false connections. With threads(n) greater
than 1 each phase borrows a connection, runs its hooks and returns it before the workers start or
after they have all finished, so a pool as small as threads cannot be starved; there, session state
a hook sets does not carry into the fill or into the other phase.
Transactions. Each script runs on the connection as it is; Bloviate never changes its autocommit setting.
- Autocommit on (the default for most drivers and pools): each statement commits as it runs. A failure leaves the statements before it applied.
- Autocommit off: the script commits once, after its last statement succeeds, and rolls back if
any statement fails. There is no separate transaction for the script: it runs in the connection’s
open transaction, so the commit or rollback applies to everything pending on that connection,
not only the script’s own statements. That includes work the caller had run and not yet committed
and rows from a fill left uncommitted by
CommitStrategy.connectionDefault(). A failingbeforeorafterscript therefore rolls those back too, and a successful one commits them. - Databases that commit DDL implicitly (MySQL, for one) commit it regardless of the mode.
Failures. Any failing statement fails fill() with a SQLException whose message names the
script and the statement’s number, line and first line of text (for example
SQL script [rollup.sql] failed at statement 2 (line 4) [INSERT INTO ...]: ...); the driver’s
exception is the cause and its SQL state is kept. A failing before hook stops the fill before any
table is written. A failing after hook is thrown, not swallowed. after hooks do not run if the
fill itself failed. There is no cross-table rollback: as with any failed fill, tables filled before
the failure stay filled, and a failing after hook does not undo the fill.
A fill is not idempotent. Running it again against tables that already hold its rows collides on
the primary keys. The supported way to re-run is a before hook that empties the tables first:
.before(SqlScript.inline("reset", "TRUNCATE orders, customers CASCADE"))Derived tables. The tables to fill are read after the before hooks and before the after hooks.
A table an after hook creates (CREATE TABLE summary AS SELECT ...) is therefore never filled. A
derived table that already exists is filled like any other, so exclude it
(excludeTables("summary"), see Selecting tables and schema) and let
the after hook compute it; otherwise it is filled with random rows that the hook then has to delete.
Selecting tables and schema
Section titled “Selecting tables and schema”By default Bloviate fills every table (views are not filled) in the connection’s current catalog and
schema. DatabaseFiller.Builder narrows that:
new DatabaseFiller.Builder(connection, config) .schema("reporting") // fill this schema, not the connection's current one .includeTables("orders", "order_*") // only these (default: all tables) .excludeTables("order_stats", "tmp_*") // ...except these .after(SqlScript.inline("rollup", """ INSERT INTO order_stats (day, orders, revenue) SELECT order_date, count(*), sum(total) FROM orders GROUP BY order_date""")) .build() .fill();includeTables and excludeTables take varargs or a Collection<String>, and are additive across
calls.
Patterns. A pattern is an unqualified table name, matched case-insensitively (like
TableConfiguration names). * matches any run of characters (including none) and ? matches exactly
one; every other character, including ., [ and %, is literal, and there is no escape. Patterns
match table names within the selected schema; they are never schema-qualified.
Rules.
- With no
includeTables, every table is a candidate; with it, only tables matching at least one include pattern are.excludeTablesis applied after that, so an excluded table is out even if an include pattern also names it. - An include pattern that matches no table fails
fill()with anIllegalArgumentExceptionnaming the pattern and listing the tables found: an empty selection is nearly always a typo. So does a selection that leaves no table at all. - An exclude pattern that matches no table only logs a warning, so an exclude list can outlive a dropped table.
- These checks run when the schema is read, after any
beforehooks, and before any row is written.
Foreign keys to a table that is not filled. Values in a foreign-key column are generated from the
parent’s primary key, so a selected table cannot reference a table that is left out. fill() fails
with an IllegalArgumentException before writing any row, naming every offending child table, its
foreign-key column(s) and the missing parent:
cannot fill: table [orders] foreign key on column(s) [customer_id] references table [customers], whichis not among the tables being filled (left out by includeTables/excludeTables, or in another schema).Add the referenced table(s) to includeTables (or remove the excludeTables pattern that drops them), orexclude the referencing table(s) as well. Nothing was written.Excluding a table that nothing references (a leaf, such as a derived rollup) always works, and a table that references itself is not an excluded parent. A foreign key into another schema (or catalog) is reported the same way, naming the parent’s schema, even when the selected schema has a table of the same name: the foreign key does not reference that one. A table in another schema cannot be included, so exclude the referencing table.
Table configurations that do not apply. A TableConfiguration naming a table that does not exist
in the selected schema is still ignored, but fill() now logs one warning listing those names. A second
warning lists configurations for tables that exist in the schema but that includeTables/excludeTables
left out, since they have no effect.
The derived-table pattern. A rollup or summary table exists in the schema but must be computed from
the generated detail rows, not filled with random data. Exclude it, and populate it in an after
hook (as in the example above). Hooks run in the selected schema, so their unqualified names resolve
there. Because the hook reads the rows the fill wrote, the derived table always agrees with them.
Schema and catalog selection. schema(...) and catalog(...) call Connection.setSchema /
setCatalog on every connection the fill uses: the metadata read, each parallel worker, and the
hook phases. The previous values are put back afterwards:
- On a
Connectionyou supply, whenfill()returns or throws, so your connection is left as it was. - On a
DataSourceconnection, before it goes back to the pool (withthreads(1)the single borrowed connection is scoped for the whole run; withthreads(n)each worker’s connection is scoped as it is borrowed). If a pooled connection cannot be restored it is aborted so the pool discards it.
The name is passed to the driver as given, so its case must match the database’s ("reporting" and
"REPORTING" differ on PostgreSQL, and H2 folds unquoted names to upper case). Bloviate checks that the
connection reports the requested value after setting it, so a schema that does not exist, or a driver
that ignores the request, is a SQLException from fill() (“cannot select schema [x]…”) rather than a
silent fill of the wrong schema. The name is a literal, never a pattern: JDBC’s catalog calls take
the schema and table name as LIKE patterns, so Bloviate escapes _ and % before passing them on and
a schema named tenant_1 matches tenant_1 alone, not tenantx1. Databases differ:
- PostgreSQL, H2, CockroachDB: use
schema(...). PostgreSQL cannot change database on a connection, socatalog(...)there fails unless it names the current database. - MySQL, MariaDB: a database is a catalog and there are no schemas, so use
catalog("db_name");schema(...)fails. - SQLite: neither is supported;
schema(...)/catalog(...)fail.
Two details worth knowing. PostgreSQL’s driver implements setSchema by replacing the whole
search_path with the one schema, so inside the fill (hooks included) types and functions living in
other schemas, such as public, need qualifying, and the restore puts back the single schema the
connection reported, not a multi-entry search_path.
On a connection with autocommit off, PostgreSQL’s setSchema is part of your open transaction, as
is anything you have pending. fill() leaves that transaction open and the connection back on its
original schema (getSchema() and getAutoCommit() are as you left them), so you can commit or roll
back afterwards. The selection is applied and restored for each phase (before hooks, the fill, after
hooks), and when a phase committed (hooks always do; the fill does with an engine-managed
commit strategy) the restore is committed as well, so a later rollback() cannot
put the connection back on the selected schema. With the default connection-default strategy the
fill’s own rows stay uncommitted until you commit, and a rollback() discards them together with the
selection. If a fill fails in a way that aborts the transaction (a constraint violation, say) the
restore cannot run until you rollback(); the connection is then back on its original schema, and
you get the fill’s own error, with the failed restore attached as a suppressed exception.
Selecting the connection’s current schema explicitly changes nothing, including the generated data: the seed depends on the table’s real schema and catalog names, never on how they were chosen.
Partitioned tables
Section titled “Partitioned tables”On PostgreSQL, Bloviate fills a declaratively partitioned table (PARTITION BY RANGE, LIST or
HASH, at any depth) through its parent: it discovers the partitioned table, leaves every one of
its partitions out, and inserts rows into the parent so the database routes each one to the partition
that accepts it. A partition is a table in its own right that only accepts rows within its bounds, so
filling one directly with unconstrained values fails; that is why Bloviate never does.
CREATE TABLE orders ( id bigint NOT NULL, tenant_id integer NOT NULL, placed_at timestamp NOT NULL, total numeric(10,2) NOT NULL, PRIMARY KEY (id, placed_at)) PARTITION BY RANGE (placed_at);CREATE TABLE orders_2024_01 PARTITION OF orders FOR VALUES FROM ('2024-01-01') TO ('2024-02-01');CREATE TABLE orders_2024_02 PARTITION OF orders FOR VALUES FROM ('2024-02-01') TO ('2024-03-01');Here Bloviate fills orders; orders_2024_01 and orders_2024_02 are not tables to fill. Keys,
foreign keys and constraints work on the parent as on any table: a primary key that includes the
partition key, a foreign key from a partitioned table to a plain table, and a foreign key from a plain
table to a partitioned table (PostgreSQL clones that key onto every partition; Bloviate reads it once,
against the parent).
Constrain the partition key. Generated values default to a window around 2020 for timestamps and
dates, and to arbitrary text or numbers otherwise, none of which is likely to fall in your partitions.
Give the partition key a ColumnConfiguration whose range the partitions cover. For date and timestamp
keys whose partitions follow the calendar, prefer a relative window
(see below) over hard-coded dates:
import io.bloviate.gen.SqlTimestampGenerator;import java.sql.Timestamp;import java.time.LocalDateTime;
ColumnConfiguration placedAt = new ColumnConfiguration("placed_at", random -> new SqlTimestampGenerator.Builder(random) .start(Timestamp.valueOf(LocalDateTime.of(2024, 1, 1, 0, 0))) // inclusive .end(Timestamp.valueOf(LocalDateTime.of(2024, 3, 1, 0, 0))) // exclusive .build());
Set<TableConfiguration> tables = Set.of(new TableConfiguration("orders", 10_000, Set.of(placedAt)));For a LIST key use a generator that only produces the listed values (for example
Distributions.weighted(Map.of("eu", 1, "us", 1))); a HASH key needs nothing, because every value hashes
to some partition. A column that references a partitioned table’s key is generated by its own column’s
generator, seeded from the parent’s key, so give it the same configuration as the key it references
(here, placed_at of the referencing table).
A DEFAULT partition. If the table has one, a value that no other partition accepts lands there, so
an unconfigured fill succeeds, and everything ends up in the default partition. Configure the key as
above to spread rows over the named partitions.
When a value fits no partition. With no DEFAULT partition the insert fails, and fill() throws
the driver’s SQLException for the failed batch, which names the table and the offending key:
Batch entry 0 insert into "shop"."orders" ("id","placed_at",...) values (...) was aborted:ERROR: no partition of relation "orders" found for row Detail: Partition key of the failing row contains (placed_at) = (2020-03-19 09:46:32.763).Nothing is skipped silently. Under the default CommitStrategy.connectionDefault() the engine leaves
your connection’s autocommit alone, so what survives the failure is up to it: on an autocommit
connection, the batches already executed (earlier tables, and earlier batches of the failing one) stay
committed, as for any other failure; on a connection with autocommit off, they are still uncommitted and
you can rollback(). A commit strategy of perTable() rolls the failing table back.
Selecting and configuring. Use the partitioned table’s name everywhere: includeTables("orders"),
excludeTables("orders") and new TableConfiguration("orders", ...). Bloviate warns, and otherwise
ignores, a name that belongs to a partition:
- a
TableConfigurationfororders_2024_01logstable configuration for [orders_2024_01] is ignored: it is a partition of [orders] ... configure [orders] instead; - an
includeTables/excludeTablespattern that matches only partitions logs a warning naming them and their parent (an include pattern that then leaves nothing to fill fails as usual). A wildcard such asorders*that matches the parent as well is fine: partitions are just never filled.
Composition. Partitioned tables work with everything else, verified against PostgreSQL: parallel
fills (threads(n)), unorderedBulk(), schema(...) and the table selection above. The intra-table
partitions setting of TableConfiguration (see Intra-table partitioning)
is a different thing that shares the word: it splits one table’s rows across workers and has nothing to
do with SQL partitioning. The two combine: each worker inserts its row range through the partitioned parent, and
the rows equal those of a sequential fill.
What is not supported.
- Only PostgreSQL (with
PostgresSupport) discovers partitions. On other databases nothing changes: MySQL, MariaDB, CockroachDB, H2, SQLite, DuckDB and BigQuery do not expose partitions as separate tables through JDBC, so a partitioned table is one table there (or, for CockroachDB, not SQL-partitioned at all). A PostgreSQL fill configured withDefaultSupportdoes not know about partitions, so it keeps the behaviour described in the note below. - Legacy inheritance partitioning (
CREATE TABLE ... INHERITS) is not declarative partitioning: those tables are filled as the ordinary tables they are. - A foreign key that references one partition directly cannot be honoured, since a partition is never filled: Bloviate logs a warning and the fill fails, before writing anything, naming the referenced partition as a table that is not being filled. Reference the partitioned table instead. (A direct foreign key that exactly mirrors one to the parent, the same columns against the same-named column of a partition, cannot be told apart from the copies PostgreSQL makes and is treated as one.)
- Bloviate does not create partitions or choose a partition key range for you. To keep the key inside
the partitions you have, give it a window relative to a pinned
asOf(see below).
Behaviour change (3.6.0). Before partitioned tables were supported, Bloviate skipped the parent and filled each partition as an independent table with unconstrained values, which fails unless the partitions happen to accept them. A schema whose partitions do accept the default values (for example a table partitioned to cover 2020) used to fill leaf-by-leaf and now fills through the parent: the same seed produces different rows there. Schemas without partitioned tables are unaffected.
Keeping the partition key inside existing partitions
Section titled “Keeping the partition key inside existing partitions”Partitions are usually cut by the calendar: this month, this quarter. A key drawn from fixed dates
misses them as time passes. A relative window resolved against a
pinned asOf keeps it inside, and the same seed and asOf give the same rows. The partitions below
cover 2025-10-01 up to 2026-04-01; the fill goes through the parent ledger, never a partition:
// ledger PARTITION BY RANGE (booked_at); partitions from 2025-10-01 to 2026-04-01ColumnConfiguration bookedAt = ColumnConfiguration.relative("booked_at", RelativeWindow.withinLast("90d"), (random, window) -> new SqlTimestampGenerator.Builder(random).window(window).build());
new DatabaseFiller.Builder(connection, new DatabaseConfiguration.Builder(128, 10_000, new PostgresSupport()) .tableConfigurations(Set.of(new TableConfiguration("ledger", 10_000, Set.of(bookedAt)))) .build()) .asOf(Instant.parse("2026-04-01T00:00:00Z")) // [2026-01-01, 2026-04-01): every row routes to ledger_2026_q1 .build() .fill();Pin asOf so the window sits inside the partitions that exist. If you leave it unpinned it is today’s
UTC date, and a window that reaches a period with no partition (and no DEFAULT) fails with the
“no partition of relation … found for row” error described above. A window that spans several
partitions spreads the rows over them.
Configuration options reference
Section titled “Configuration options reference”Database configuration options
Section titled “Database configuration options”- Batch Size: Number of records inserted in each batch operation
- Record Count: Default number of records to generate per table
- Database Support: Database-specific implementation for optimal compatibility
- Table Configurations: Override the row count for specific tables, and the intra-table
partitionscount for splitting a large table across workers on the parallel path - Column Configurations: Override the generator for specific columns (case-insensitive, reproducible)
- Seed: Base seed for reproducible generation; the same schema and seed always produce the same
data (defaults to
0) - asOf (
DatabaseFiller.Builder.asOf(Instant)): the anchor relative date windows are measured from; pin it for reproducible output (defaults to the start of the current UTC day) - Commit Strategy: How the engine commits — leave autocommit alone (default), commit once per table, or commit every N batches
- Bulk Load Strategy: Fill in foreign-key dependency order (default), or
unorderedBulk()to disable constraint enforcement and fill every table at once with no topological barrier (parallel path only; PostgreSQL/MySQL, with CockroachDB falling back)
Parallelism (worker threads for concurrent table fill) is configured on the
DatabaseFiller.Builder via threads(n) with the DataSource constructor, and so are the
before(...)/after(...) SQL hooks and the schema(...)/catalog(...)/
includeTables(...)/excludeTables(...) table and schema selection.
File generation options
Section titled “File generation options”- Output Format: CSV, TSV, or pipe-delimited
- Row Count: Number of rows to generate
- Custom Column Definitions: Full control over data generation