One reference image or twenty? What similarity search needs

If you've built a classifier before, you know the drill: collect a hundred labeled chips, split them into train and test, fine-tune, check your confusion matrix, repeat until the numbers look acceptable. That workflow trains the instinct that any image search task needs a pile of examples before it's worth running. For similarity search on satellite imagery, that instinct is wrong more often than it's right, and it's worth understanding why before you waste a week building a reference set nobody needed.

One tile versus a training set

A classifier learns a decision boundary. It needs enough positive and negative examples to figure out which pixels, textures, and shapes separate "target" from "not target" across all the variation those categories contain. One photo of a runway doesn't teach a model what every runway looks like, so classifiers are hungry for volume almost by definition.

A similarity search built on image embeddings does something different. A foundation model has already learned a general representation of what's in overhead imagery, structure, texture, shadow, material, layout, from training on a broad swath of imagery before you ever touch it. When you mark one tile, you're not teaching the model anything new. You're asking it: of everything in your search area, what else sits close to this point in the space the model already built? That's a lookup, not a training run. One clean reference image is enough to run it.

This is the real answer to the minimum training examples question for image retrieval: for a classifier, plan on dozens to hundreds depending on how much visual variation the class contains. For embedding-based similarity retrieval, the floor is one.

When one example isn't the right call

One example being enough isn't the same as one example always being the best choice. A single query tile does a clean job when the thing you're hunting is visually consistent: a specific compound layout, a distinctive structure footprint, a storage configuration with a signature shape. The target doesn't vary much from instance to instance, so one good chip gives the model a sharp anchor point.

Problems start when the category you care about has more than one visual form. Say you're looking for improvised river crossings, and some are plank bridges while others are pontoon rafts. A single reference tile of the plank bridge will pull back more plank bridges and miss the rafts entirely. The search isn't broken; you asked a narrower question than you meant to. Add two or three reference tiles, one per visual variant, and run them as separate queries or merge the results into one pass. That's the real shape of the one example versus many question: it's not about statistical confidence, it's about how many distinct looks your target has.

A second common reason to add a reference image is tightening a noisy result set. If your first query pulls back a lot of near-misses, a known reference of a look-alike you want excluded can sharpen the ranking. That's refinement, not training. You're still not building a labeled dataset, you're giving the search a second or third coordinate to triangulate from.

What this means for how you work

The practical upshot is that you don't need to stockpile chips before you start. Mark the one clear example you already have, whether it's from a prior report, an open-source image, or a tile you pulled yourself, and see what comes back across your search area. If the target has one visual form, that single tile is your whole workflow. If it has two or three, add a tile for each variant and treat it as refinement, not as reaching some minimum threshold.

That's the gap between how analysts have been forced to search, tile by tile across an archive too large to browse by eye, and what embedding-based retrieval on high-res imagery makes possible once you give it a starting point. If you've got one example and a search area that would otherwise take weeks to scan manually, get in touch about early access to Broad Area Search and see what else in your area looks like it.