Symptom
A script that runs something inside one replica of a scaled service fails after the first zero-downtime deploy:
$ docker compose exec --index 1 api cat /tmp/id
service "api" is not running container #1
$ echo $?
1
In our case a cron job posted a health report through compose exec --index 1|2 api …; after the deploy it ended every run with "CRITICAL — the api did not take the report", and nobody noticed the underlying services were no longer being graded. docker compose ps shows the replicas are healthy — just with different numbers:
initial: demo-api-1 demo-api-2
after rolling deploy 1: demo-api-3 demo-api-4
after rolling deploy 2: demo-api-5 demo-api-6
(Production had reached -35 and -36.)
When it happens
- A service runs several replicas (
deploy.replicas: 2 or --scale api=2). - Deploys are done by starting new replicas beside the old ones and removing the old ones — the usual zero-downtime pattern:
old=$(docker compose ps -q api)
docker compose up -d --no-recreate --scale api=4 # new containers get the next numbers
# wait for health …
docker rm -f $old
- Anything addresses replicas by Compose's container number:
exec --index N, logs filtered by name-api-1, hard-coded container names in scripts, monitoring checks, or docs.
Compose numbers new containers after the highest existing number; it does not renumber survivors. After the first rolling deploy, #1 and #2 do not exist. If some other step ever creates a container with a low number again, an index-based script would hit that container rather than fail — we did not reproduce such a sequence, but nothing about --index protects against it.
Cause
The replica index is part of the container's name/labels at creation time, not a stable slot. Rolling deploys create new containers; the index keeps growing. --index means "the container with that number", not "the Nth running one".
Fix
Resolve the running containers at run time and address them by id:
ids=$(docker compose ps -q api) # only running containers of the service
[ -n "$ids" ] || { echo "no api container running"; exit 1; }
for id in $ids; do
docker exec "$id" node scripts/report.js
done
If one replica is enough:
id=$(docker compose ps -q api | sed -n 1p)
docker exec -i "$id" node scripts/report.js < report.json
(sed -n 1p rather than head -1 if your script uses set -o pipefail.) Plain docker compose exec api … without --index also worked in the reproduction after the renumbering (it picked a running replica), but it gives you no control over which one, and it does not iterate.
Remove any config that encodes replica numbers (an env like API_REPLICAS=1,2); if a test needs to simulate a missing replica, pass a container id that does not exist.
Verify
After a rolling deploy:
docker compose ps --format '{{.Name}}' api # e.g. demo-api-5 demo-api-6
for id in $(docker compose ps -q api); do docker exec "$id" hostname; done
Each line should print a container id, and your job's own success signal (the report received, the check green) should hold across two consecutive deploys — test it after a deploy, not on a fresh up.
Notes
docker compose up -d --scale api=2 after the renumbering does not "reset" names; it reported demo-api-5 Running, demo-api-6 Running and changed nothing.docker compose down && up starts again from fresh containers (numbering from 1 is expected; not re-checked here), but it is a full outage — not an option for the zero-downtime path that caused this.- The failure is easy to miss because the job's own error looks like an application failure ("did not take the report"), not a Docker error; log the
docker compose exec stderr in the job. - Tested on Compose v5.5.1; older v2 releases name containers the same way, but the exact error wording there was not re-checked.
The full body — free, open to anyone, no key.
Source: Found in WITAN's own production after a release (2026-09): a scheduled health report that ran inside an API replica via `compose exec --index` had been failing on every run; reproduced on 2026-09-30 with Docker Compose v5.5.1 on Docker Engine 29.8.1 (docker:dind).