Symptom
Semantic search over a small corpus never says "nothing found". Typing asdkjhqwe returns results whose similarities look respectable, and sometimes higher than a real question's best match. A test that asserts "a nonsense query finds nothing close" passes when run alone and fails in the full regression.
Measurements
Model Xenova/multilingual-e5-small (384 dimensions, q8 ONNX) in transformers.js 3.8.1, mean pooling and normalization, query: / passage: prefixes as the model card asks. Passages are title + "\n" + body[:2000].
1. Six technical articles (Postgres, nginx, Docker, OAuth, payments), measured on 2026-10-02:
| query kind | best cosine | worst cosine |
|---|
| on-topic (6 questions, each about one article) | 0.884–0.909 | 0.767–0.817 |
nonsense (asdkjhqwe, qwpoeiruty lkjhg, zzzz 9981 kkk, …) | 0.802–0.824 | 0.763–0.788 |
| off-topic real questions (pasta, Everest, 1998 World Cup, knitting, Korean restaurants) | 0.729–0.778 | 0.701–0.752 |
Every query sits between about 0.70 and 0.91 against every passage. Nonsense ranks above sensible off-topic questions. The gap that separates a real answer is the lead of the best over the second: 0.026–0.097 for on-topic queries, at most 0.013 for nonsense and off-topic.
2. On the project's 18 published units (2026-10-01, from its search code): 28 off-topic or nonsense questions had a best match of at most 0.8423, never more than 0.013 above the second. 38 questions with an answer on the market had a best of 0.79–0.91. Those below 0.845 that still ranked the right unit first stood 0.018–0.04 above the second.
A floor that worked on that corpus
// close: at the floor; or the best, near the floor and clearly ahead of the next. Rows far below the best drop.
export function closeEnough(rows) { // sorted best first; fetch one more row than you show
if (!rows.length) return [];
const best = rows[0].similarity;
const ahead = rows.length > 1 && best - rows[1].similarity >= 0.015;
if (!(best >= 0.845 || (best >= 0.82 && ahead))) return [];
return rows.filter((r, i) => i === 0 || (r.similarity >= 0.845 && r.similarity >= best - 0.05));
}
On the 18-unit corpus: all 28 off-topic questions got nothing, and 33 of the 38 answerable ones kept their answer. These thresholds are corpus-dependent. With more passages, an off-topic query finds closer neighbours. Make each value configurable and re-measure as the corpus grows.
When nothing passes, return an empty result and say what to do next (the reference API points the caller to a requests board), rather than the nearest rows.
The test trap: digits in the "nonsense" query
The full regression's test database holds every suite's fixtures, and many are titled with Unix-time stamps. The failing assertion asked asdkjhqwe sw<epoch>. Measured on 2026-10-02 against three fixture-like passages titled with other epoch numbers, plus one normal article:
| query | best fixture | normal article |
|---|
asdkjhqwe zxqvmplk | 0.784 | 0.798 |
asdkjhqwe sw1759403999 | 0.847 | 0.793 |
asdkjhqwe 1759403999 | 0.849 | 0.795 |
A shared run of digits alone, not the same number, lifted a nonsense query over a 0.845 floor. Run alone, the suite passed. In the full regression, with every other suite's stamped fixtures in the same database, one of them came close enough.
What the reference project changed:
- the nonsense query has no digits and none of the run's stamp words (
asdkjhqwe zxqvmplk); - the rule is tested on its own with made-up similarity lists (under the floor and level → none; at it → one; near it and ahead → one; two at it → two; far below → none; empty → none);
- the end-to-end "nothing close" answer is asserted only when the database's own unfiltered ranking, run through the same rule, finds nothing close. Otherwise the test asserts what the rule gives.
Notes
- A probability-looking cosine from e5 isn't a confidence. Compare against measured distributions for your corpus, not against 0.5 or 0.8 by intuition.
- An assertion about global content ("nothing in the index matches X") on a shared test database is coupled to every other suite's fixtures. Test the rule with synthetic inputs, and check the content only under conditions you've measured.
The full body — free, open to anyone, no key.
Source: First-hand from WITAN's semantic search (pgvector, multilingual-e5-small via transformers.js). The 18-unit measurements and the floor rule come from the change that added the floor, merged on 2026-10-01 and shipped in v0.20.0. The test failure was in the v0.20.3 release gate on 2026-10-02 (the search suite's nothing-close check failed in the full regression and passed alone), fixed the same day. The six-article table and the digit table were measured for this unit on 2026-10-02 on a Windows PC (Docker Desktop, 12 CPUs) in the project's api image with the cached q8 model, transformers.js 3.8.1, Node 22.23.3. The six articles were WITAN's six published error-to-fix units, and the fixture passages were written for the measurement. These are small samples; the numbers describe this model and these texts, not every corpus.