Skip to content

Configuration

By default Bloviate inspects the schema and picks a generator for every column based on its JDBC type. This guide shows how to take progressively more control — from overriding row counts, to overriding individual columns, to non-uniform distributions, reproducible seeds, and the parallel fill path.

Configuration is layered:

  • DatabaseConfiguration — global defaults: batch size, default row count, database support, and an optional set of per-table overrides.
  • TableConfiguration — overrides the row count for one table, and optionally carries per-column overrides.
  • ColumnConfiguration — overrides how a single column is generated, via a ColumnGeneratorFactory (a Random -> DataGenerator<?> lambda). The engine hands the factory a column-seeded Random so output stays reproducible.

Generate different numbers of rows for specific tables while every other table uses the default:

import io.bloviate.db.*;
import io.bloviate.ext.PostgresSupport;
import java.util.Set;
// "users" gets 50 rows; all other tables fall back to the default (100)
Set<TableConfiguration> tableConfigs = Set.of(
new TableConfiguration("users", 50)
);
DatabaseConfiguration config = new DatabaseConfiguration(
10, // batch size
100, // default rows per table
new PostgresSupport(), // database support
tableConfigs // table-specific overrides
);
new DatabaseFiller.Builder(connection, config)
.build()
.fill();

Pin a specific column to a custom generator. Here the status_code column on orders is constrained to integers in [1, 10) instead of the type’s default:

import io.bloviate.db.*;
import io.bloviate.ext.PostgresSupport;
import io.bloviate.gen.IntegerGenerator;
import java.util.Set;
// ColumnGeneratorFactory is a `RandomGenerator -> DataGenerator<?>` lambda.
// The engine supplies a column-seeded RandomGenerator for reproducible output.
Set<ColumnConfiguration> columnConfigs = Set.of(
new ColumnConfiguration("status_code",
random -> new IntegerGenerator.Builder(random).start(1).end(10).build())
);
// 1,000 rows for "orders", with the column override applied
Set<TableConfiguration> tableConfigs = Set.of(
new TableConfiguration("orders", 1000, columnConfigs)
);
DatabaseConfiguration config = new DatabaseConfiguration(
128, 100, new PostgresSupport(), tableConfigs);
new DatabaseFiller.Builder(connection, config)
.build()
.fill();

Column names are matched case-insensitively. Any column without an override keeps its default, type-based generator.

Real columns are rarely uniform — a status is mostly ACTIVE, a rating clusters around its mean, a referenced product_id follows a popularity curve, and a created_at bunches toward the present. The Distributions helper returns ready-made ColumnGeneratorFactory values so a column can opt into a non-uniform distribution without writing a factory:

import io.bloviate.db.*;
import io.bloviate.ext.PostgresSupport;
import java.util.Map;
import java.util.Set;
Set<ColumnConfiguration> columnConfigs = Set.of(
// 70% NEW, 25% SHIPPED, 5% CANCELLED (weights need not sum to 1)
new ColumnConfiguration("status", Distributions.weighted(Map.of("NEW", 0.7, "SHIPPED", 0.25, "CANCELLED", 0.05))),
// normal(mean=4, sd=1) rounded and clamped to [1, 5]
new ColumnConfiguration("rating", Distributions.normalInt(4, 1, 1, 5)),
// Zipfian (power-law) over [1, 10000] — a few hot ids, a long thin tail
new ColumnConfiguration("product_id", Distributions.zipfian(10_000)),
// timestamps skewed toward the recent end of the window
new ColumnConfiguration("created_at", Distributions.recentTimestamps())
);
DatabaseConfiguration config = new DatabaseConfiguration(
128, 100, new PostgresSupport(),
Set.of(new TableConfiguration("orders", 100_000, columnConfigs)));

Available shapes: weighted(...) (categorical), normal(...) / normalInt(...) (bounded Gaussian), zipfian(...) (power-law), and recentTimestamps(...) (recency-skewed). Each is built from the engine’s column seed, so output stays reproducible and composes with foreign-key reseeding and parallel fills like any other generator. These are specified distributions, not distributions learned from real data.

On PostgreSQL and CockroachDB, Bloviate reads each table’s CHECK constraints and ENUM types and generates values that satisfy them — automatically, no configuration. So given:

CREATE TYPE order_status AS ENUM ('NEW', 'PAID', 'SHIPPED', 'CANCELLED');
CREATE TABLE orders (
status order_status NOT NULL,
rating integer CHECK (rating BETWEEN 1 AND 5),
priority integer CHECK (priority IN (1, 2, 3)),
amount numeric(8,2) CHECK (amount >= 0 AND amount <= 9999.99)
);

status only gets one of its enum labels, rating lands in [1, 5], priority is one of 1/2/3, and amount stays in range — instead of random values an insert would reject. The common forms are honored: IN (...), BETWEEN, and >=/<=/>/< comparisons, for integer, numeric, floating, and text columns, plus enum/domain allowed values.

Dates that must fall on the first of a period are honored too, on DATE, TIMESTAMP and TIMESTAMP WITH TIME ZONE columns (since 3.5.0):

CREATE TABLE invoices (
billing_month date CHECK (date_trunc('month', billing_month) = billing_month),
period_start timestamptz CHECK (EXTRACT(day FROM period_start) = 1)
);

date_trunc('month' | 'quarter' | 'year', col) = col and EXTRACT(day FROM col) = 1 get a TruncatedDateGenerator: the first day of a random month in a fixed window (2015 up to 2025 by default; the same seed gives the same data), or midnight on that day for a timestamp. A timestamptz value is midnight in the session time zone, which is the zone the check is evaluated in, so the fill works whatever the connection’s zone is.

Notes:

  • The parser only recognises the exact forms above: a bare column, optionally cast, compared with literals. A CHECK that calls a function on the column (lower(status) IN (...), length(name) >= 1), does arithmetic, or takes any other form is skipped with a warning, and the column falls back to its type default. That includes negation, OR, LIKE patterns, a one-sided bound, other date_trunc units (day, week, …), and CHECKs over more than one column.
  • Before 3.5.0 a quoted function argument was read as an allowed value, so date_trunc('month', d) = d made the fill fail with invalid input syntax for type date: "month".
  • Before 3.9.0 no CHECK or ENUM was read on CockroachDB, so those columns were filled from their type default. Most such fills failed outright — an enum column got an arbitrary string, and a range or IN check rejected the value — but a column whose check the type default happened to satisfy did fill, and the values it produces change in 3.9.0, because they now come from the constraint. Re-pin any CockroachDB fixture you compare byte-for-byte.
  • A per-column override or a registry rule always wins, so you can still take full control of a constrained column.
  • Open the connection with stringtype=unspecified (already required for PostgreSQL’s extension types) so enum/IN values bind. This applies to CockroachDB too, which is reached through the same driver.
  • CockroachDB is covered by the same reader (since 3.9.0): it serves the same pg_catalog queries, and the definitions it stores differ only in spellings the parser accepts on both — it keeps BETWEEN verbatim where PostgreSQL expands it into two comparisons, writes extract('day', col) with a comma rather than EXTRACT(day FROM col), and casts literals as ::STRING. Every other database reads no CHECKs, so give those columns an explicit generator.

Foreign keys and the values they reference

Section titled “Foreign keys and the values they reference”

A foreign-key column is not checked against the parent’s rows — it replays them. Bloviate generates it from the same seed at the same row index as the key it points at, so the values match without reading anything back. Three things follow, and they are worth knowing before you read generated data.

It follows the columns the constraint names. A key to a UNIQUE target works, not just one to the parent’s primary key. The tenant pattern is the usual case:

create table accounts (
id int primary key,
tenant_id int not null references tenants (id),
unique (tenant_id, id));
create table documents (
tenant_id int not null,
account_id int not null,
foreign key (tenant_id, account_id) references accounts (tenant_id, id),
foreign key (tenant_id) references tenants (id));

A column in several keys makes those keys equal. documents.tenant_id above belongs to two keys at once, so its value has to exist in both. Bloviate guarantees that by generating every column linked by a foreign key — in either direction — from one shared seed. For an ordinary parent/child chain nothing changes. But point one column at two unrelated keys and those two key columns come out carrying identical values, row for row. That is not a quirk to work around: a value cannot be in two key spaces unless the key spaces overlap, and equality is the only overlap reproducible without reading rows back. If you did not mean the keys to be interchangeable, the schema is telling you something.

Row counts bound it. A foreign-key column cycles through its parent’s keys once it runs past them, so a child with more rows than its parent is fine — it reuses parents. Row i of a child reads row i % parentRows of its parent, which in turn reads row (i % parentRows) % grandparentRows of the grandparent, all the way up. The counts come from each table’s TableConfiguration where it has one and from the default row count otherwise.

Keys the database generates. A serial or IDENTITY key column is left out of its own insert, so the database assigns it and there is no seed to share. Bloviate counts instead: a table filled from empty is given 1..N in insertion order, and the child counts through the same range. This assumes the parent’s sequence starts at 1 — filling a table that already holds rows, or whose sequence has been advanced, leaves the child pointing at keys the parent was not given. Fill into an empty schema, or reset the sequence first (TRUNCATE ... RESTART IDENTITY). Because the class carries one set of values, a key linked to a generated one is filled by counting too, so both ends match.

One shape this cannot cover: a generated column inside a composite key whose parent is filled with intra-table partitions. The database hands out identities in completion order across the workers, while the key’s other columns are generated from the logical row index, so the two need not describe the same parent row. Give such a parent an ordinary key column, or fill it without partitions.

Two or more tables that reference each other — directly, or through a chain — cannot be filled in any order: whichever goes first, its foreign key has no parent row to point at. A fill fails before writing anything, naming every cycle it found:

cannot fill: the tables [invoice, payment] reference each other, so no fill order satisfies them: ...

Three ways forward:

  • break the cycle in the schema;
  • leave every table of the cycle out with excludeTables(...) — dropping only one side leaves the others referencing a table that is not being filled, which fails for that reason instead;
  • fill with BulkLoadStrategy.unorderedBulk(), which fills every table at once without enforcing constraints, so it needs no order. Two things have to hold for it to be chosen: a DataSource with threads(n) above 1 (on a single connection the fill is sequential and the strategy is only warned about), and a support that implements bulk loading — PostgreSQL, MySQL, MariaDB and BigQuery. Anywhere else it falls back to the ordered path and fails the same way.

A table referencing itself is not a cycle: it constrains the order of rows within that one table, not the order of tables. Those fill normally, with a warning, and making a row’s parent exist before its child is up to you.

DatabaseConfiguration takes a base seed. The same schema filled with the same seed always produces identical data, so test fixtures are deterministic; change the seed for a different — but still reproducible — dataset. Per-column seeds are derived from stable column identity, and foreign keys share the seed of the key they reference (see foreign keys and the values they reference), so referential fidelity holds for any seed.

import io.bloviate.db.*;
import io.bloviate.ext.PostgresSupport;
// batch size, rows/table, support, table configs, seed
DatabaseConfiguration config = new DatabaseConfiguration(
128, 100, new PostgresSupport(), null, 42L);
new DatabaseFiller.Builder(connection, config).build().fill();

The seed defaults to 0 when you use the four-argument constructor, so existing code keeps a single, stable dataset without changes.

Time zones do not enter into it. A java.sql.Date, Time or Timestamp is an instant, and turning one into the value a column holds needs a zone. Bloviate binds every temporal through an explicit UTC calendar, so the stored value is the same on every machine: a DATE, TIME, TIMESTAMP or DATETIME column gets the instant’s UTC wall clock, and a column that carries a zone (TIMESTAMP WITH TIME ZONE, timestamptz) gets the instant itself. Before 3.9.1 the zone was left to the driver, which used the JVM’s default, so the same seed produced values five hours apart on a machine in America/New_York and one in UTC (issue #640). If your JVM was not already running in UTC, the values a seed produces change in 3.9.1 — they stop following the machine. A JVM already on UTC, which is the usual container and CI setup, is unaffected.

One thing a client cannot pin: MySQL’s TIMESTAMP is an instant that the server converts from the session time zone on write and back to it on read. Its stored value therefore follows the session zone by definition, whatever a client sends. DATE, TIME and DATETIME carry no zone and do not move. Set the session zone explicitly (a before SQL hook running SET time_zone = '+00:00') if you need MySQL TIMESTAMP columns pinned too.

A custom generator that binds a java.sql temporal itself should go through io.bloviate.gen.TemporalBinding, which is the one place this zone rule lives; the java.time-based generators (TruncatedDateGenerator) are already zone-independent.

Date and timestamp generators default to a fixed window around 2020-01-01, so on their own they miss “this quarter” partitions and range CHECKs. A relative window says where the data should fall in terms of a moving anchor, asOf, instead of absolute dates: “within the last 90 days”, “from 30 days ago to a week from now”. (Absolute bounds, start(...)/end(...), still work and are unchanged.)

import io.bloviate.db.*;
import io.bloviate.ext.PostgresSupport;
import io.bloviate.gen.*;
import java.time.Instant;
import java.util.Set;
// orders may only be placed in Q1 2026: CHECK (placed_at >= '2026-01-01' AND placed_at < '2026-04-01')
Set<ColumnConfiguration> columns = Set.of(
ColumnConfiguration.relative("placed_at", RelativeWindow.withinLast("90d"),
(random, window) -> new SqlTimestampGenerator.Builder(random).window(window).build()),
ColumnConfiguration.relative("due_on", RelativeWindow.between("-30d", "+7d"),
(random, window) -> new SqlDateGenerator.Builder(random).window(window).build()));
DatabaseConfiguration config = new DatabaseConfiguration.Builder(128, 1_000, new PostgresSupport())
.seed(42L)
.tableConfigurations(Set.of(new TableConfiguration("orders", 1_000, columns)))
.build();
new DatabaseFiller.Builder(connection, config)
.asOf(Instant.parse("2026-04-01T00:00:00Z")) // pin the anchor: same seed + same asOf = same rows
.build()
.fill();

withinLast("90d") with asOf 2026-04-01 is the 90 days from 2026-01-01 up to (not including) 2026-04-01, so every placed_at satisfies the CHECK above, in the same rows every time.

Wall-clock time is never an input to generation. The one place a date can enter is asOf, and it is explicit:

  • Pinned (DatabaseFiller.Builder.asOf(Instant)): the same seed and the same asOf produce byte-identical output on every run and every JDK. This is the sanctioned, reproducible way to get “current” data.
  • Not pinned: when the fill starts, asOf is read once from the clock and truncated to the start of the current UTC day (00:00Z). It is logged once at INFO the first time a relative window uses it (resolved asOf 2026-06-10T00:00:00Z ...; pin it with asOf(...) for reproducible output) and is available from DatabaseFiller.asOf(). Two unpinned fills on the same UTC day agree; across midnight they do not. To reproduce such a fill, pin the instant it logged.

One asOf serves the whole fill: every table, every intra-table partition and every worker thread of a parallel fill uses the same instant, so the tables of one dataset are consistent with each other. A fill that uses no relative window is unaffected by asOf: its output is exactly what it was, and the same seed still gives the same data with or without one.

An offset is an optional sign, digits and one unit letter (case-sensitive):

UnitMeaningArithmetic (all in UTC)
hhoursexact 3,600 seconds
ddaysexact 24 hours (UTC has no daylight saving)
wweeksexact 7 days
Mmonthscalendar: the day of month is clamped, so 2026-03-31 minus 1M is 2026-02-28
yyearscalendar: 2024-02-29 plus 1y is 2025-02-28

Examples: 90d, -30d, +7d, 12w, 6M, 1y, 36h. There is no minute or second unit; a lower-case m is rejected (months are M). Whitespace inside an offset is an error.

WindowMeaningResolved for asOf = 2026-04-01
RelativeWindow.withinLast("90d")[asOf - 90d, asOf); the span must be positive[2026-01-01, 2026-04-01)
RelativeWindow.between("-30d", "+7d")[asOf + start, asOf + end)[2026-03-02, 2026-04-08)
RelativeWindow.withinLast("6M")six calendar months back[2025-10-01, 2026-04-01)

A window is half-open: the start is inclusive and the end exclusive, the same convention as SqlTimestampGenerator, SqlDateGenerator, DateGenerator, InstantGenerator and TruncatedDateGenerator. So withinLast never produces the anchor instant itself; use between("-90d", "1d") to include asOf’s whole day. Bad input fails fast with an IllegalArgumentException that names the problem: an unknown unit, a zero or negative withinLast, an end that is not after the start. RelativeOffset.parse(...), RelativeWindow.withinLast(String) and RelativeWindow.between(String, String) take the text forms, so a configuration file or command line can hand them the same strings.

ColumnConfiguration.relative(column, window, factory) resolves the window against the fill’s asOf and hands the concrete bounds to your factory. Those generators accept one through window(RelativeWindow.Resolved): SqlTimestampGenerator, SqlDateGenerator, DateGenerator, InstantGenerator, SkewedTimestampGenerator and TruncatedDateGenerator. Distributions.recentTimestamps(window, skew) is the recency-skewed shorthand. Anything else that needs the anchor takes the fill’s GenerationContext: ColumnGeneratorFactory.contextual((random, context) -> ...), or, for a registry rule, GeneratorFactory.contextual((column, random, context) -> ...), where context.asOf() is the anchor. Existing factories (plain lambdas) are unchanged and ignore the context.

  • A relative window is the way to keep the key of a partitioned table inside its existing partitions; configure it on the partitioned parent, which is what Bloviate fills (see Keeping the partition key inside existing partitions).
  • A per-column relative configuration takes precedence over a CHECK constraint, a registry rule and the support default, like any other per-column override.
  • SqlDateGenerator reads its window as whole UTC calendar days, so a DATE column holds exactly the dates in the window whatever the JVM or session time zone. A window a few hours wide inside one day contains no whole date and is rejected. TruncatedDateGenerator draws the period starts in the window (its bounds rounded up to whole dates).
  • For a zone-less TIMESTAMP column the JDBC driver renders the instant in the JVM’s default time zone, as it does for absolute bounds. If a hard bound (a CHECK or a partition edge) must hold and the JVM may run in another zone, use TIMESTAMP WITH TIME ZONE or start the JVM with -Duser.timezone=UTC.
  • SqlTimeGenerator has no window: a TIME column has no date to be relative to.

For large, wide schemas the fill can run in parallel. Construct the filler from a pooled DataSource instead of a single Connection and ask for more than one worker thread:

import io.bloviate.db.*;
import io.bloviate.ext.PostgresSupport;
import javax.sql.DataSource;
DataSource dataSource = /* a pooled DataSource, e.g. HikariCP */;
DatabaseConfiguration config = new DatabaseConfiguration(
1000, 100_000, new PostgresSupport(), null, 42L);
new DatabaseFiller.Builder(dataSource, config)
.threads(8) // fill independent tables concurrently
.build()
.fill();

Bloviate groups tables into topological levels by their foreign keys and fills the independent tables within each level concurrently, one connection per worker, barriering between levels so a child table is never filled before its parent. Each worker fills its table in a single transaction (commit once per table). The fill stays fully reproducible: a table’s data depends only on its own seed and row order, never on which tables fill alongside it, so the same config and seed produce the same row content as a sequential fill across every deterministic column (physical row order and wall-clock columns aside).

How much this helps depends on the schema. A wide schema of independent tables sees a large speedup (~3× with 8 workers on a 10-table, 1M-row fixture); a deep, narrow foreign-key chain (each table depending on the previous) has little to parallelize. See the benchmarks for numbers.

The single-Connection constructor is unchanged and remains the default sequential path — threads only applies to the DataSource form.

When a single large table dominates the fill, between-table parallelism can’t help it — it sits alone in its topological level. Set partitions on that table’s TableConfiguration to split its rows into that many contiguous ranges filled concurrently, one connection per range, on the parallel (DataSource + threads) path:

// split the one big table into 8 row ranges; ignored on the single-Connection path
Set<TableConfiguration> tables = Set.of(
new TableConfiguration("events", 50_000_000L, 8 /* partitions */));
DatabaseConfiguration config = new DatabaseConfiguration(
1000, 0, new PostgresSupport(), tables, 42L);
new DatabaseFiller.Builder(dataSource, config).threads(8).build().fill();

Partitioning is byte-identical to a sequential fill of the same seed, for any partition count: every built-in generator derives each value as a pure function of its column seed and the absolute row index (per-index derivation), so keys, foreign keys, and plain random columns alike land on exactly the values the sequential fill produces, foreign-key validity always holds, and seeking a worker to its starting row is O(1) no matter how large the table or its parents are. The only exceptions are custom generators that opt out of per-row positioning (DataGenerator.positionable() returning false, as the datafaker integration does because its values come from an internal Faker RNG) — those stay deterministic for a given partition count but may differ across partition counts.

Size the connection pool for the total concurrent demand (threads, where a partitioned table counts as partitions units). One case is unsupported: partitioning a parent table whose primary key comes from a non-positionable custom generator referenced by a foreign key can orphan those references — partition the child table instead, or use the positional key generators (as the bundled TPC-C/TPC-H configurations do). A custom generator with internal positional state must implement io.bloviate.gen.IndexedDataGenerator to stay aligned under partitioning.

By default the engine leaves the connection’s autocommit untouched (a typical autocommit connection commits per executeBatch()). Disabling autocommit and committing less often cuts overhead. Pass a CommitStrategy to DatabaseConfiguration for the sequential path:

import io.bloviate.db.*;
// commit once per table (autocommit off for the fill, restored afterward)
DatabaseConfiguration perTable = new DatabaseConfiguration(
1000, 100_000, new PostgresSupport(), null, 42L, CommitStrategy.perTable());
// or bound the open transaction: commit every 50 JDBC batches
DatabaseConfiguration everyN = new DatabaseConfiguration(
1000, 100_000, new PostgresSupport(), null, 42L, CommitStrategy.everyNBatches(50));

The default, CommitStrategy.connectionDefault(), preserves today’s behavior (the engine never touches autocommit). The parallel path already commits once per table; a configured strategy applies there too.

Tip — driver batch rewrite. Bloviate inserts in JDBC batches, but most drivers only collapse a batch into a single multi-row INSERT when you opt in via the JDBC URL: PostgreSQL reWriteBatchedInserts=true, MySQL rewriteBatchedStatements=true. Enabling it is often the single biggest fill speedup, sequential or parallel. Bloviate logs a warning at fill time when the parameter is missing, and io.bloviate.util.JdbcUrls.withBatchRewrite(url, support.batchRewriteUrlParameter()) builds a correctly-parameterized URL if you construct the DataSource yourself. CockroachDB ignores the parameter, so no warning is emitted there.

Tip — BigQuery. Rewriting is unconditional in tbc-bq-jdbc, so there is no parameter to set and no warning to emit. Instead, size batchSize against BigQuery’s limit of 10,000 query parameters per query: the effective rows per job is 10_000 / columnCount, so the default 128 leaves most of a job unused. Start from batchSize = 5000, and set the driver’s batchLoadThreshold to match to move large batches onto its NDJSON load-job path. Leave CommitStrategy at its default — BigQuerySupport opts out of engine-managed transactions because setAutoCommit(false) starts a BigQuery session and silently disables that load path, and each executeBatch is already one atomic job. Bloviate logs a warning if a commit strategy is configured anyway.

The parallel path normally barriers between topological levels, so a deep, narrow foreign-key chain (each table depending on the previous) serializes — there is little within any one level to run concurrently. BulkLoadStrategy.unorderedBulk() removes that barrier: it disables foreign-key enforcement, fills every table at once, then re-enables enforcement.

import io.bloviate.db.*;
import io.bloviate.ext.PostgresSupport;
DatabaseConfiguration config = new DatabaseConfiguration(
1000, 100_000, new PostgresSupport(), null, 42L,
null, // CommitStrategy (null = default)
BulkLoadStrategy.unorderedBulk()); // disable constraints, fill barrier-free, re-enable
new DatabaseFiller.Builder(dataSource, config).threads(8).build().fill();

This is safe because Bloviate’s data is referentially consistent by construction: a foreign-key column is seeded from its referenced primary-key column, so child and parent generate identical key values regardless of insert order. Disabling enforcement therefore changes nothing about correctness — for the same seed the result has the same row content as an ordered fill across every deterministic column (physical row order aside) — it only removes the ordering constraint and the per-row foreign-key checks. The win is largest on deep chains (e.g. TPC-C’s warehouse → district → customer → open_order → order_line); wide, FK-free schemas already saturate their workers in one level and see little change.

Requirements and fallback:

  • Only effective on the parallel path (a DataSource with threads > 1); it is ignored with a warning on the single-Connection and single-thread paths, which fill in dependency order.
  • Supported on PostgreSQL (SET session_replication_role = replica, which needs a superuser/rds_superuser role) and MySQL (SET FOREIGN_KEY_CHECKS=0/UNIQUE_CHECKS=0, no special privilege). CockroachDB does not support it and transparently falls back to the ordered level-parallel path.
  • Each worker disables enforcement on its own pooled connection and restores it in a finally before returning the connection to the pool, so no connection ever leaks back with checks suppressed. If enforcement cannot be disabled (e.g. the role lacks privilege), the engine logs a warning and falls back to the ordered path rather than running half-disabled.

The default, BulkLoadStrategy.ordered(), preserves today’s behavior (dependency-ordered, constraints always enforced).

DatabaseFiller.Builder can run SQL scripts around the fill: before hooks run before any table is filled, after hooks run once every table is filled. The classic use is a derived table computed from the generated rows (a rollup, a summary), so it cannot disagree with its source; another is emptying tables so a fill can be re-run.

import io.bloviate.db.*;
import java.nio.file.Path;
import java.util.Map;
new DatabaseFiller.Builder(connection, config)
.before(SqlScript.resource("sql/reset.sql")) // classpath resource
.after(SqlScript.file(Path.of("sql/rollup.sql")) // file on disk
.withTokens(Map.of("schema", "reporting"))) // ${schema} in the script
.after(SqlScript.inline("analyze", "ANALYZE orders_by_day")) // inline text
.build()
.fill();
-- sql/rollup.sql
DELETE FROM ${schema}.orders_by_day;
INSERT INTO ${schema}.orders_by_day (day, orders, revenue)
SELECT order_date, count(*), sum(total) FROM orders GROUP BY order_date;

before(...) and after(...) are additive and ordered: call them as often as you like and the scripts run in the order added. A SqlScript is a name (used in log lines and error messages) plus its text, from SqlScript.inline(name, sql), SqlScript.file(path) or SqlScript.resource(name). SqlScriptRunner.run(connection, script) runs one on its own, for example to load a schema.

What a script may contain. Statements are separated by ;. A ; inside a 'string' (with '' as the escape, and \' inside E'...'), a "quoted" or `backtick` identifier, a -- or /* */ comment (which nests, as in PostgreSQL), or a PostgreSQL dollar-quoted body ($$...$$, $tag$...$tag$, so CREATE FUNCTION bodies work) does not end a statement. Empty statements are skipped and the last statement needs no ;. ${name} is replaced from the script’s tokens, everywhere except inside comments; a token with no value fails the script before any statement runs. MySQL’s DELIMITER command is not supported and fails with a clear message rather than mis-splitting the script, and backslash escapes are honoured only in E'...' strings.

Where hooks run. Before hooks run first, ahead of the schema read, so they can create the tables about to be filled. After hooks run last. On the single-Connection path both run on the connection you supplied. On a DataSource without threads(n) (or with threads(1)) Bloviate borrows one connection and uses it for the before hooks, the fill and the after hooks, so session state a hook sets (SET search_path, a temporary table) carries through, and the after hooks see the fill’s own uncommitted rows even if the pool hands out autoCommit=false connections. With threads(n) greater than 1 each phase borrows a connection, runs its hooks and returns it before the workers start or after they have all finished, so a pool as small as threads cannot be starved; there, session state a hook sets does not carry into the fill or into the other phase.

Transactions. Each script runs on the connection as it is; Bloviate never changes its autocommit setting.

  • Autocommit on (the default for most drivers and pools): each statement commits as it runs. A failure leaves the statements before it applied.
  • Autocommit off: the script commits once, after its last statement succeeds, and rolls back if any statement fails. There is no separate transaction for the script: it runs in the connection’s open transaction, so the commit or rollback applies to everything pending on that connection, not only the script’s own statements. That includes work the caller had run and not yet committed and rows from a fill left uncommitted by CommitStrategy.connectionDefault(). A failing before or after script therefore rolls those back too, and a successful one commits them.
  • Databases that commit DDL implicitly (MySQL, for one) commit it regardless of the mode.

Failures. Any failing statement fails fill() with a SQLException whose message names the script and the statement’s number, line and first line of text (for example SQL script [rollup.sql] failed at statement 2 (line 4) [INSERT INTO ...]: ...); the driver’s exception is the cause and its SQL state is kept. A failing before hook stops the fill before any table is written. A failing after hook is thrown, not swallowed. after hooks do not run if the fill itself failed. There is no cross-table rollback: as with any failed fill, tables filled before the failure stay filled, and a failing after hook does not undo the fill.

A fill is not idempotent. Running it again against tables that already hold its rows collides on the primary keys. The supported way to re-run is a before hook that empties the tables first:

.before(SqlScript.inline("reset", "TRUNCATE orders, customers CASCADE"))

Derived tables. The tables to fill are read after the before hooks and before the after hooks. A table an after hook creates (CREATE TABLE summary AS SELECT ...) is therefore never filled. A derived table that already exists is filled like any other, so exclude it (excludeTables("summary"), see Selecting tables and schema) and let the after hook compute it; otherwise it is filled with random rows that the hook then has to delete.

By default Bloviate fills every table (views are not filled) in the connection’s current catalog and schema. DatabaseFiller.Builder narrows that:

new DatabaseFiller.Builder(connection, config)
.schema("reporting") // fill this schema, not the connection's current one
.includeTables("orders", "order_*") // only these (default: all tables)
.excludeTables("order_stats", "tmp_*") // ...except these
.after(SqlScript.inline("rollup", """
INSERT INTO order_stats (day, orders, revenue)
SELECT order_date, count(*), sum(total) FROM orders GROUP BY order_date"""))
.build()
.fill();

includeTables and excludeTables take varargs or a Collection<String>, and are additive across calls.

Patterns. A pattern is an unqualified table name, matched case-insensitively (like TableConfiguration names). * matches any run of characters (including none) and ? matches exactly one; every other character, including ., [ and %, is literal, and there is no escape. Patterns match table names within the selected schema; they are never schema-qualified.

Rules.

  • With no includeTables, every table is a candidate; with it, only tables matching at least one include pattern are. excludeTables is applied after that, so an excluded table is out even if an include pattern also names it.
  • An include pattern that matches no table fails fill() with an IllegalArgumentException naming the pattern and listing the tables found: an empty selection is nearly always a typo. So does a selection that leaves no table at all.
  • An exclude pattern that matches no table only logs a warning, so an exclude list can outlive a dropped table.
  • These checks run when the schema is read, after any before hooks, and before any row is written.

Foreign keys to a table that is not filled. Values in a foreign-key column are generated from the parent’s primary key, so a selected table cannot reference a table that is left out. fill() fails with an IllegalArgumentException before writing any row, naming every offending child table, its foreign-key column(s) and the missing parent:

cannot fill: table [orders] foreign key on column(s) [customer_id] references table [customers], which
is not among the tables being filled (left out by includeTables/excludeTables, or in another schema).
Add the referenced table(s) to includeTables (or remove the excludeTables pattern that drops them), or
exclude the referencing table(s) as well. Nothing was written.

Excluding a table that nothing references (a leaf, such as a derived rollup) always works, and a table that references itself is not an excluded parent. A foreign key into another schema (or catalog) is reported the same way, naming the parent’s schema, even when the selected schema has a table of the same name: the foreign key does not reference that one. A table in another schema cannot be included, so exclude the referencing table.

Table configurations that do not apply. A TableConfiguration naming a table that does not exist in the selected schema is still ignored, but fill() now logs one warning listing those names. A second warning lists configurations for tables that exist in the schema but that includeTables/excludeTables left out, since they have no effect.

The derived-table pattern. A rollup or summary table exists in the schema but must be computed from the generated detail rows, not filled with random data. Exclude it, and populate it in an after hook (as in the example above). Hooks run in the selected schema, so their unqualified names resolve there. Because the hook reads the rows the fill wrote, the derived table always agrees with them.

Schema and catalog selection. schema(...) and catalog(...) call Connection.setSchema / setCatalog on every connection the fill uses: the metadata read, each parallel worker, and the hook phases. The previous values are put back afterwards:

  • On a Connection you supply, when fill() returns or throws, so your connection is left as it was.
  • On a DataSource connection, before it goes back to the pool (with threads(1) the single borrowed connection is scoped for the whole run; with threads(n) each worker’s connection is scoped as it is borrowed). If a pooled connection cannot be restored it is aborted so the pool discards it.

The name is passed to the driver as given, so its case must match the database’s ("reporting" and "REPORTING" differ on PostgreSQL, and H2 folds unquoted names to upper case). Bloviate checks that the connection reports the requested value after setting it, so a schema that does not exist, or a driver that ignores the request, is a SQLException from fill() (“cannot select schema [x]…”) rather than a silent fill of the wrong schema. The name is a literal, never a pattern: JDBC’s catalog calls take the schema and table name as LIKE patterns, so Bloviate escapes _ and % before passing them on and a schema named tenant_1 matches tenant_1 alone, not tenantx1. Databases differ:

  • PostgreSQL, H2, CockroachDB: use schema(...). PostgreSQL cannot change database on a connection, so catalog(...) there fails unless it names the current database.
  • MySQL, MariaDB: a database is a catalog and there are no schemas, so use catalog("db_name"); schema(...) fails.
  • SQLite: neither is supported; schema(...)/catalog(...) fail.

Two details worth knowing. PostgreSQL’s driver implements setSchema by replacing the whole search_path with the one schema, so inside the fill (hooks included) types and functions living in other schemas, such as public, need qualifying, and the restore puts back the single schema the connection reported, not a multi-entry search_path.

On a connection with autocommit off, PostgreSQL’s setSchema is part of your open transaction, as is anything you have pending. fill() leaves that transaction open and the connection back on its original schema (getSchema() and getAutoCommit() are as you left them), so you can commit or roll back afterwards. The selection is applied and restored for each phase (before hooks, the fill, after hooks), and when a phase committed (hooks always do; the fill does with an engine-managed commit strategy) the restore is committed as well, so a later rollback() cannot put the connection back on the selected schema. With the default connection-default strategy the fill’s own rows stay uncommitted until you commit, and a rollback() discards them together with the selection. If a fill fails in a way that aborts the transaction (a constraint violation, say) the restore cannot run until you rollback(); the connection is then back on its original schema, and you get the fill’s own error, with the failed restore attached as a suppressed exception.

Selecting the connection’s current schema explicitly changes nothing, including the generated data: the seed depends on the table’s real schema and catalog names, never on how they were chosen.

On PostgreSQL, Bloviate fills a declaratively partitioned table (PARTITION BY RANGE, LIST or HASH, at any depth) through its parent: it discovers the partitioned table, leaves every one of its partitions out, and inserts rows into the parent so the database routes each one to the partition that accepts it. A partition is a table in its own right that only accepts rows within its bounds, so filling one directly with unconstrained values fails; that is why Bloviate never does.

CREATE TABLE orders (
id bigint NOT NULL, tenant_id integer NOT NULL, placed_at timestamp NOT NULL, total numeric(10,2) NOT NULL,
PRIMARY KEY (id, placed_at)
) PARTITION BY RANGE (placed_at);
CREATE TABLE orders_2024_01 PARTITION OF orders FOR VALUES FROM ('2024-01-01') TO ('2024-02-01');
CREATE TABLE orders_2024_02 PARTITION OF orders FOR VALUES FROM ('2024-02-01') TO ('2024-03-01');

Here Bloviate fills orders; orders_2024_01 and orders_2024_02 are not tables to fill. Keys, foreign keys and constraints work on the parent as on any table: a primary key that includes the partition key, a foreign key from a partitioned table to a plain table, and a foreign key from a plain table to a partitioned table (PostgreSQL clones that key onto every partition; Bloviate reads it once, against the parent).

Constrain the partition key. Generated values default to a window around 2020 for timestamps and dates, and to arbitrary text or numbers otherwise, none of which is likely to fall in your partitions. Give the partition key a ColumnConfiguration whose range the partitions cover. For date and timestamp keys whose partitions follow the calendar, prefer a relative window (see below) over hard-coded dates:

import io.bloviate.gen.SqlTimestampGenerator;
import java.sql.Timestamp;
import java.time.LocalDateTime;
ColumnConfiguration placedAt = new ColumnConfiguration("placed_at", random ->
new SqlTimestampGenerator.Builder(random)
.start(Timestamp.valueOf(LocalDateTime.of(2024, 1, 1, 0, 0))) // inclusive
.end(Timestamp.valueOf(LocalDateTime.of(2024, 3, 1, 0, 0))) // exclusive
.build());
Set<TableConfiguration> tables = Set.of(new TableConfiguration("orders", 10_000, Set.of(placedAt)));

For a LIST key use a generator that only produces the listed values (for example Distributions.weighted(Map.of("eu", 1, "us", 1))); a HASH key needs nothing, because every value hashes to some partition. A column that references a partitioned table’s key is generated by its own column’s generator, seeded from the parent’s key, so give it the same configuration as the key it references (here, placed_at of the referencing table).

A DEFAULT partition. If the table has one, a value that no other partition accepts lands there, so an unconfigured fill succeeds, and everything ends up in the default partition. Configure the key as above to spread rows over the named partitions.

When a value fits no partition. With no DEFAULT partition the insert fails, and fill() throws the driver’s SQLException for the failed batch, which names the table and the offending key:

Batch entry 0 insert into "shop"."orders" ("id","placed_at",...) values (...) was aborted:
ERROR: no partition of relation "orders" found for row
Detail: Partition key of the failing row contains (placed_at) = (2020-03-19 09:46:32.763).

Nothing is skipped silently. Under the default CommitStrategy.connectionDefault() the engine leaves your connection’s autocommit alone, so what survives the failure is up to it: on an autocommit connection, the batches already executed (earlier tables, and earlier batches of the failing one) stay committed, as for any other failure; on a connection with autocommit off, they are still uncommitted and you can rollback(). A commit strategy of perTable() rolls the failing table back.

Selecting and configuring. Use the partitioned table’s name everywhere: includeTables("orders"), excludeTables("orders") and new TableConfiguration("orders", ...). Bloviate warns, and otherwise ignores, a name that belongs to a partition:

  • a TableConfiguration for orders_2024_01 logs table configuration for [orders_2024_01] is ignored: it is a partition of [orders] ... configure [orders] instead;
  • an includeTables/excludeTables pattern that matches only partitions logs a warning naming them and their parent (an include pattern that then leaves nothing to fill fails as usual). A wildcard such as orders* that matches the parent as well is fine: partitions are just never filled.

Composition. Partitioned tables work with everything else, verified against PostgreSQL: parallel fills (threads(n)), unorderedBulk(), schema(...) and the table selection above. The intra-table partitions setting of TableConfiguration (see Intra-table partitioning) is a different thing that shares the word: it splits one table’s rows across workers and has nothing to do with SQL partitioning. The two combine: each worker inserts its row range through the partitioned parent, and the rows equal those of a sequential fill.

What is not supported.

  • Only PostgreSQL (with PostgresSupport) discovers partitions. On other databases nothing changes: MySQL, MariaDB, CockroachDB, H2, SQLite, DuckDB and BigQuery do not expose partitions as separate tables through JDBC, so a partitioned table is one table there (or, for CockroachDB, not SQL-partitioned at all). A PostgreSQL fill configured with DefaultSupport does not know about partitions, so it keeps the behaviour described in the note below.
  • Legacy inheritance partitioning (CREATE TABLE ... INHERITS) is not declarative partitioning: those tables are filled as the ordinary tables they are.
  • A foreign key that references one partition directly cannot be honoured, since a partition is never filled: Bloviate logs a warning and the fill fails, before writing anything, naming the referenced partition as a table that is not being filled. Reference the partitioned table instead. (A direct foreign key that exactly mirrors one to the parent, the same columns against the same-named column of a partition, cannot be told apart from the copies PostgreSQL makes and is treated as one.)
  • Bloviate does not create partitions or choose a partition key range for you. To keep the key inside the partitions you have, give it a window relative to a pinned asOf (see below).

Behaviour change (3.6.0). Before partitioned tables were supported, Bloviate skipped the parent and filled each partition as an independent table with unconstrained values, which fails unless the partitions happen to accept them. A schema whose partitions do accept the default values (for example a table partitioned to cover 2020) used to fill leaf-by-leaf and now fills through the parent: the same seed produces different rows there. Schemas without partitioned tables are unaffected.

Keeping the partition key inside existing partitions

Section titled “Keeping the partition key inside existing partitions”

Partitions are usually cut by the calendar: this month, this quarter. A key drawn from fixed dates misses them as time passes. A relative window resolved against a pinned asOf keeps it inside, and the same seed and asOf give the same rows. The partitions below cover 2025-10-01 up to 2026-04-01; the fill goes through the parent ledger, never a partition:

// ledger PARTITION BY RANGE (booked_at); partitions from 2025-10-01 to 2026-04-01
ColumnConfiguration bookedAt = ColumnConfiguration.relative("booked_at", RelativeWindow.withinLast("90d"),
(random, window) -> new SqlTimestampGenerator.Builder(random).window(window).build());
new DatabaseFiller.Builder(connection, new DatabaseConfiguration.Builder(128, 10_000, new PostgresSupport())
.tableConfigurations(Set.of(new TableConfiguration("ledger", 10_000, Set.of(bookedAt))))
.build())
.asOf(Instant.parse("2026-04-01T00:00:00Z")) // [2026-01-01, 2026-04-01): every row routes to ledger_2026_q1
.build()
.fill();

Pin asOf so the window sits inside the partitions that exist. If you leave it unpinned it is today’s UTC date, and a window that reaches a period with no partition (and no DEFAULT) fails with the “no partition of relation … found for row” error described above. A window that spans several partitions spreads the rows over them.

  • Batch Size: Number of records inserted in each batch operation
  • Record Count: Default number of records to generate per table
  • Database Support: Database-specific implementation for optimal compatibility
  • Table Configurations: Override the row count for specific tables, and the intra-table partitions count for splitting a large table across workers on the parallel path
  • Column Configurations: Override the generator for specific columns (case-insensitive, reproducible)
  • Seed: Base seed for reproducible generation; the same schema and seed always produce the same data (defaults to 0)
  • asOf (DatabaseFiller.Builder.asOf(Instant)): the anchor relative date windows are measured from; pin it for reproducible output (defaults to the start of the current UTC day)
  • Commit Strategy: How the engine commits — leave autocommit alone (default), commit once per table, or commit every N batches
  • Bulk Load Strategy: Fill in foreign-key dependency order (default), or unorderedBulk() to disable constraint enforcement and fill every table at once with no topological barrier (parallel path only; PostgreSQL/MySQL, with CockroachDB falling back)

Parallelism (worker threads for concurrent table fill) is configured on the DatabaseFiller.Builder via threads(n) with the DataSource constructor, and so are the before(...)/after(...) SQL hooks and the schema(...)/catalog(...)/ includeTables(...)/excludeTables(...) table and schema selection.

  • Output Format: CSV, TSV, or pipe-delimited
  • Row Count: Number of rows to generate
  • Custom Column Definitions: Full control over data generation