ImageTextRemoverA PictureEditor.com tool

Flat ground, busy ground

Before you draw a single box, look at what is underneath the words. That one glance predicts the result better than the choice of fill, the size of the box or anything else you can change.

Three grounds, one caption bar

Picture the same wide caption strip in three frames. In the first it lies across a pale evening sky. In the second it lies across a bank of grass. In the third it lies across a bookshelf, with the shelf edge running straight through the middle of it.

What comes back, and why
GroundDiffusion fillTrained fill
Sky, a wall, a blurred backdropInvisibleInvisible
Grass, gravel, fabric, foliageVisible smearConvincing
A shelf, a railing, a face, small typeVisible smearConfident and wrong

The first row is the reason the diffusion fill is the default here. It downloads nothing, finishes in under a fifth of a second, and on a smooth ground there is nothing a trained model can add — both answers are simply correct.

Why texture and structure are different problems

Texture is detail without a plan: blades of grass, chips of gravel, the weave of a jumper. A fill that averages its way inward from the border turns texture into soup, because averaging is exactly the wrong operation for detail whose whole character is that neighbouring pixels disagree. A trained network has seen a great deal of grass and will happily produce more of it, and nobody can tell the new grass from the old, because there was never anything to match.

Structure is detail with a plan. A shelf edge, a window frame, a course of bricks, a horizon, the line of a jaw. It has to continue across the box in the one particular way that agrees with both sides, and there is exactly one right answer that neither fill can know. A trained network will still produce something confident — that is what it does — and confident and wrong is the worst result available, because it looks fine at a glance and falls apart when somebody looks properly.

A crisp fiction is worse than a visible smear on a picture you are about to print. That is the reason the box list marks regions with structure running across their border, and suggests the plainer fill for them.

How the site decides to warn you

It is a crude test and it is deliberately crude. The site walks the four borders just outside your box and asks how often a strong edge appears on both of two facing sides at the same offset. Two opposite borders that break in the same places usually means something runs between them. It will occasionally warn you about a busy but structureless background, and it will occasionally miss a faint line, which is why the warning is a note beside a thumbnail rather than a block on the run.

What to do about a bad ground

Crop instead, if the lettering is near an edge and the frame can spare it — losing a strip of a photograph costs nothing next to inventing one badly. Pick another frame, if this is a still from something with more than one. Shrink the box, if part of it sits over plain ground and only part of it crosses the shelf: two smaller boxes with a strip of untouched picture between them is often cleaner than one large one.

And if none of those work, take the smear. The diffusion fill over a busy ground looks like what it is, and there is a kind of honesty in a result that does not pretend the picture was always that way.

Back to the picture