Symptom
A Postgres container built from the official image (or one based on it, like pgvector/pgvector) comes up healthy, but unrelated features answer HTTP 500, with errors of this form behind them (the objects named are the ones that were missing in this incident):
ERROR: relation "invite_requests" does not exist
ERROR: function witan_bought_unit(uuid, uuid) does not exist
Some of the newest migration is there (its first new columns exist), and everything after that point is missing. Restarting the container doesn't help. The log of the second start says only:
PostgreSQL Database directory appears to contain a database; Skipping initialization
The real error is in the log of the very first start, and it's easy to miss:
/usr/local/bin/docker-entrypoint.sh: running /docker-entrypoint-initdb.d/02-feature.sql
psql:/docker-entrypoint-initdb.d/02-feature.sql:4: ERROR: column o.wallet_address does not exist
When it happens
- The schema is applied by mounting a directory of numbered files into the init directory:
db:
image: pgvector/pgvector:pg17
restart: unless-stopped
volumes:
- db-data:/var/lib/postgresql/data
- ./db/init:/docker-entrypoint-initdb.d:ro
- One file contains a statement that fails (a typo, or a reference to a column that another branch renamed).
- The volume is new: a fresh dev stack, a CI database, or a throwaway test copy. Long-lived databases migrated by a separate tool never run these files, so they don't show it.
Cause
On an empty data directory the entrypoint runs initdb, then runs each *.sql file in name order with psql -v ON_ERROR_STOP=1. Each statement autocommits. When a statement fails:
- Everything before it in that file stays applied. The file isn't wrapped in a transaction.
- The rest of that file and every later file are skipped. The entrypoint exits with an error.
restart: unless-stopped (or your orchestrator) starts the container again. PGDATA is no longer empty, so the entrypoint skips initialization and starts the server on the partial schema.
From then on the stack looks healthy (pg_isready passes) and the failure shows up only as missing objects, wherever they're first used.
Reproduced on postgres:17-alpine (PostgreSQL 17.11) with three files:
- File 01 created a table.
- File 02 added a column, then failed on a function, then had one more
CREATE TABLE. - File 03 created another table.
Result: the column from 02 was present. The function, the table after the error in 02 and the table from 03 were all missing. The container was up again after one automatic restart.
Fix
Repair the broken file, then re-initialize the affected databases. Restarting won't re-run init:
docker compose down -v # drops the volume: only for dev/CI data you can lose
docker compose up -d
For a database you can't drop, apply the fixed file and every later one by hand (or with your migration tool). Write migrations so that a file that stopped half-way can run again:
IF NOT EXISTS and CREATE OR REPLACE;- constraints added in
DO $$ … EXCEPTION WHEN duplicate_object …; - updates limited to rows not yet migrated.
Then stop it from happening again. Apply the init directory to a fresh database in CI on every pull request, the same way the entrypoint does:
for f in $(ls db/init | grep -E '^[0-9]{2}-.*\.sql$' | sort); do
psql -d fresh -q -v ON_ERROR_STOP=1 -f "db/init/$f" > /dev/null \
|| { echo "::error file=db/init/$f::$f fails on a fresh database"; exit 1; }
done
Use the image from your compose file as the CI service container. If deploys use a separate migration tool, also check the upgrade path: apply the last release's files, then run your migrator. Compare the two schemas (columns from information_schema.columns plus function signatures from pg_proc). A migration can work on one path and fail on the other, and two pull requests that each pass alone can break together.
Verify
- Run
docker compose logs db | grep -E 'ERROR|Skipping initialization' on a stack you suspect. If an ERROR appears before the first Skipping initialization, the schema is partial. - Check that the last migration's last object exists:
SELECT to_regclass('public.<last_table>'), to_regproc('<last_function>');. NULL means init stopped early. - In CI, the fresh-apply job fails on the pull request that adds the broken file, and names the file.
Notes
- A healthcheck on
pg_isready can't catch this. If you want the stack to refuse to start, check for a sentinel object from your newest migration. - PL/pgSQL function bodies are only checked when they run, so a fresh apply can pass while a plpgsql function is still broken. Call the important ones in CI too. SQL-language functions are checked at
CREATE time (see the related unit on column … does not exist at CREATE FUNCTION). - The CI job in the reference project takes about a minute. On the pull request that added it, it applied 55 files fresh and upgraded from the previous release (adding files 55 to 59). The two paths agreed on 673 columns and functions.
The full body — free, open to anyone, no key.
Source: Diagnosed and fixed in WITAN's own stack on 2026-10-01, during the v0.20.0 release candidates. Migration 58 failed twice on fresh databases: first a missing space before ON in a CREATE INDEX, then a column that did not exist. Test copies started on a partial schema, and 14 suites failed with missing tables and functions. The two fixes and a CI check that applies db/init fresh and as an upgrade were merged on 2026-10-01 and shipped in v0.20.0. Production never ran the broken migration. The image used was pgvector/pgvector:pg17. The init-skip behaviour was reproduced on 2026-10-01 with postgres:17-alpine (PostgreSQL 17.11) on Docker Desktop (Engine 29.7.2, WSL2).