Symptom
Two separate problems in one semantic-search API that embeds each query on the CPU with @huggingface/transformers (ONNX Runtime) and the quantized Xenova/multilingual-e5-small:
- Under load, semantic search managed about 20 requests a second, with a median of 4.3 s at 100 users. The model itself is small.
- A tempting fix, collecting concurrent queries into one batched inference, made search results depend on which other queries happened to be in the batch.
Measurements
Threads. ONNX Runtime by default starts intra-op threads to match the core count, for every inference. On a 72-core development host that cost 62 ms per query embedding. With one or two threads it took 7–10 ms (9.5 ms with one thread, 7.4 ms with two). On a 12-CPU Docker Desktop VM, measured for this unit (40 queries, median / p90):
intraOpNumThreads | median | p90 |
|---|
| default | 7.6 ms | 21.3 ms |
| 1 | 7.6 ms | 9.6 ms |
| 2 | 6.2 ms | 8.5 ms |
| 4 | 7.1 ms | 12.3 ms |
The default's cost grows with the core count: little difference in the median on 12 CPUs (but a much worse tail), and 6–8× on 72 cores.
Semantic search through the whole stack, before → after fixing two threads (same load test, 72-core host):
| users | req/s | median | answered by keyword fallback |
|---|
| 10 | 9 → 111 | 796 → 66 ms | none |
| 50 | 18 → 200 | 2,405 → 102 ms | 14% |
| 100 | 20 → 332 | 4,306 → 105 ms | 41% |
(The fallback is the API's own overload path: past a short queue it answers by keyword instead of waiting.)
Batching. The same 16 queries embedded one at a time vs together in one call (2026-10-02, q8 model):
- cosine between a query's "alone" and "in a batch" vectors: 0.9959–0.9983, never 1;
- the order of the top 3 (out of 6 passages) changed for 6 of the 16 queries;
- the same query embedded alone twice: bit-for-bit identical.
The project's earlier measurement found the same thing: cosine down to 0.995, and the top three reordered for 2 queries in 10. A query embedded alone while 15 others were running concurrently (separate calls) was bit-for-bit the vector it gets alone.
Cause
- Threads: with many cores and many concurrent requests, each inference's thread pool fights the others for the machine. Few threads per inference, many inferences in parallel, wins for small models.
- Batching: the project's explanation is that the q8 model's dynamic quantization is computed over the whole batch, so a row's result depends on its neighbours. This wasn't checked against an fp32 model here, because only the quantized file was cached.
Fix
import { pipeline } from "@huggingface/transformers";
const THREADS = Math.max(1, Number(process.env.EMBED_THREADS ?? 2));
const extractor = await pipeline("feature-extraction", "Xenova/multilingual-e5-small", {
dtype: "q8",
session_options: { intraOpNumThreads: THREADS, interOpNumThreads: 1 },
});
// one query per call, never batched — a batched query is a different vector
const out = await extractor([`query: ${text}`], { pooling: "mean", normalize: true });
Then: a small, fixed number of concurrent embedding slots with a short queue (2 slots and 32 waiting in the reference API), a fallback when the queue is full (keyword search), and a cache of recent query vectors keyed by the trimmed text.
Use the same settings for stored vectors (the indexing worker) as for queries, and check it. In the reference project, vectors from the two-thread extractor had cosine 0.9999993 against the stored ones.
Verify
- Embed the same text alone twice: identical. Alone vs inside a batch of 16: if the cosine isn't 1, don't batch queries whose ranking matters.
- Measure per-inference latency with 1, 2, 4 and default threads on the production core count. A laptop hides the problem.
- Load-test the search endpoint and count fallback answers, not only latency.
Notes
- Batching is still fine for offline indexing if every passage is always embedded the same way. The problem is mixing alone and batched vectors, or rankings that change with traffic.
- The 72-core numbers were measured on 2026-09-27 on a development host that has since been retired. The 12-CPU numbers are from a Windows PC under Docker Desktop. Re-measure on your own hardware.
The full body — free, open to anyone, no key.
Source: First-hand from WITAN's semantic search API. The thread-count change, the 72-core latencies, the load-test before/after table and the first batching measurement were made on 2026-09-27 on the project's 72-core development host and shipped in v0.10.0 (numbers from that commit and the project's load baseline). The 12-CPU thread table and the 16-query batching measurement were made for this unit on 2026-10-02 on a Windows PC (Docker Desktop VM, 12 CPUs) in the project's api image: transformers.js 3.8.1, Node 22.23.3, Xenova/multilingual-e5-small q8 from the local cache, passages being WITAN's six published units. The quantization explanation is the project's; it was not checked against an fp32 model.