The upscale broke text that was already correct
Generative upscalers re-draw letterforms and will respell a word that was already right. Mask text out of the pass, or drop denoise on text-bearing images.
The word was right. You proofed it. Then you ran the upscale so the file would hold at print size, or on a 4K plate, and the same word came back as a neighbour of itself: a swapped vowel, a collapsed stem, a kerning that no longer matches the brief. That is not a soft file. That is a generative model re-drawing letterforms, and it will happily respell a string it was given correctly.
An upscaler does not enlarge pixels. Traditional resizing spreads the pixels you have, which is why it looks soft. A generative pass predicts the missing high-frequency detail and draws it in. Texture (skin, cloth, foliage, paint) is what that prediction is good at. Information (a specific sequence of characters) is what it fabricates. Both halves of that sentence matter, and they are why "the upscale looks sharper" is not the same claim as "the upscale is faithful."
If the detail that matters in the frame is a word, the upscaler is the wrong tool to run over that word.
Why a correct letter gets re-spelled
A letterform is a structure of strokes. At the source resolution those strokes occupy a handful of pixels. The upscaler's job is to invent the pixels between them. For a pore or a weave there is a strong prior for what should be there, and the invention is usually believable. For a letter there is also a strong prior, and it is the prior for letter-shaped marks, not for your string.
Three things make the fabrication more likely:
The source did not actually encode the glyph. At typical draft resolutions a regular-weight letter is a handful of pixels of stroke. The upscaler is not recovering a G. It is looking at a G-shaped blob and sampling a G-like shape. Neighbours of G are in that sample.
The pass is tiled, or regional. Many generative upscalers denoise tiles independently. A letter that straddles a tile boundary is solved twice. That is a seam problem on texture and a spelling problem on type: the left half of an R and the right half of a K can agree they are "a letter" and still not be the same letter.
Sharpening is not spelling. A high-frequency boost makes stems crisper and also makes the wrong stem look more confident. Over-sharpening a video upscale is already known to amplify flicker. On still type it amplifies a misspelling you would have caught when the letter was still a bit soft.
This is the same reason the upscale-versus-native decision treats a product macro with legible text as a generate-native shot. The rest of this post is what to do when you already have a correct frame and still need the extra resolution.
Protect the type, upscale the rest
The reliable repair is not "pick a better upscaler." Topaz Upscale Image is a detail-faithful enlarger. Clarity Crystal is documented as regenerating fine texture rather than stretching pixels, which is exactly the behaviour that respells. Either can be the right pass for the photograph and the wrong pass for the word. Treat the lettering as a plate you protect.
- Duplicate the source. One copy is the photograph. One copy is the type lockup.
- Mask the text out of the upscale. Cover every glyph you have proofed, plus a small margin of the surface they sit on so you are not cutting through a stem. Run the upscaler on the unmasked picture. The image upscale path is that pass.
- Scale the original lettering to the new pixel size with a non-generative resample (bicubic or similar, in any layer editor). The letters stay the letters. They will look slightly softer than the regenerated photograph around them. That is the honest trade: a slightly soft correct word beats a sharp wrong one.
- Seat the type back on the upscaled plate. If the join shows, a light inpaint over the edge of the mask only, not over the glyphs. You are blending paper into paper, not asking the sampler to re-draw the ink.
You now have generated texture at the delivery size and original spelling at the delivery size. Proof the string at 100 percent against the brief before you export. The whole point of this procedure is that the upscaler does not get a vote on the characters.
If the type is a headline you set yourself, the cleaner version is: upscale a clean plate, then set the type as a layer at the delivery resolution. Do not bake display type into a frame you intend to enlarge.
If you cannot mask: drop the denoise, or skip the pass
Not every upscale surface gives you a mask. Some hosted passes are a single control: pick a target size and run. In that case you have two remaining levers.
Drop denoising strength (or the control that is doing the same job). Denoise, creativity, resemblance, "recover detail": whatever the slider is called, it decides how much of the source gets thrown away before regeneration. High values are a new picture that happens to share a silhouette. Low values are a correction. Text-bearing images want the correction. Bracket it if the control exists: same file at low, medium, and high, and keep the one that still spells. Do not raise the value because the photograph looks a bit soft. Soft and correct is the deliverable; sharp and wrong is a reprint.
Hosted models do not always expose that dial. When they do not, do not look for it. Switch to the mask-and-composite path above, or skip the upscale.
Skip the upscale and generate at the delivery size. If the frame does not exist yet, this is the default for anything with copy you will be held to. If the frame exists and the type is already correct, a non-generative enlarge plus a sharpening pass that does not re-synthesise is often enough for web delivery. Reserve the generative pass for the photograph around the words, or for frames that have no words.
A working rule, compressed:
| Frame | Generative upscale | What you do with the type |
|---|---|---|
| No lettering, or lettering nobody can check | Yes | Nothing |
| Short hero words, already proofed | Yes, with type masked out | Non-generative scale of the original glyphs, composite back |
| Logo, price, claim, legal line | No | Generate native, or composite real type onto a clean plate |
| Text-bearing frame, no mask control on the upscaler | Only at low denoise, if a denoise control exists | If it respells, revert and composite |
Two passes of 2x usually beat one 4x pass for control on photographs. That advice does not extend to type. Two generative passes are two chances to respell. One protected pass is the ceiling.
Check at the delivery size, in the real colour space
Thumbnail review is how this defect ships. A swapped vowel is invisible in a grid and obvious on a phone, a packing carton, or a six-sheet. Read the string out loud against the brief at 100 percent of the output size, not at the size of the source.
On video the same fabrication happens per frame, and it does not agree with itself. A label that reads BAKERY for most of a clip and BAKREY for a handful of frames is the thing a viewer's eye locks onto. Do not run a still-image upscaler over every frame of a clip that contains copy. Play the result at full speed and watch the letters. If the word crawls, keep the type off the generative pass: clean plate, then real captions or a real label composite. On-screen text in generated video covers the shot-design version of that rule.
Upscale last, after grading and after any text repair. Enhancing a misspelling, then fixing the word, then enhancing again is how you pay for the defect twice.
FAQ
Why did only some of the letters change?
Because each glyph is a small independent structure. The upscaler samples them locally. A stem that was well encoded at source (a thick I, a wide H) survives. A thin stroke or a closed counter (an e, an a, an 8) is the first to be replaced by a neighbour. Proof the whole string. Do not assume the letters you did not glance at were spared.
Can I upscale first and inpaint the word afterwards?
You can, and it is the wrong order. You have now asked a second generative model to invent the word at the larger size, with no original glyph to lock to. Inpaint the source if the source is wrong; protect the source if the source is right; then enlarge. An inpaint belongs on the file that already spells, or on a clean plate you will set real type onto.
The upscaler has no denoise control and no mask. Now what?
Do not run it on a text-bearing image. Enlarge with a non-generative resample if you only need more pixels for a layout, or generate the asset at the delivery size, or composite real type onto an upscaled clean plate. A single-knob generative enlarge of proofed copy is how GRAND becomes GRAMD at 4K.
Does a "text" or "CGI" preset fix this?
It can reduce halos and keep stems from ringing, which is a sharpening problem, not a spelling problem. A preset that is kinder to type is still synthesising type. Use it if you have nothing else and the copy is decorative. Do not trust it with a brand name. Mask the name out.