Stealth model debuts on public leaderboards
HappyHorse 1.0 topped Artificial Analysis with no lab attached to it. How codename entries work, why labs run them, and when testing one early pays off.
For a stretch of April, the top text-to-video entry on Artificial Analysis belonged to nobody. HappyHorse 1.0 appeared on the board around 7 April 2026 under that name and nothing else — no lab, no model card, no endpoint — and climbed to first place before fal's write-up and the wider round of reporting settled that Alibaba was behind it.
That sequence is not a fluke of one release. Anonymous or codenamed entries are a normal part of how model launches get staged now, and if you evaluate models for a living, the useful question is not "who is it" but "what should I do about an entry I cannot yet attribute." The answer is narrower than most people assume, and it mostly is not "test it."
What a codename entry looks like from the outside
There is no official marker for a stealth entry, so you identify one by what is missing rather than what is present. A live, attributed model release arrives with a provider page, a model card, a documented endpoint, a published license and a price. A codename entry has a rank and an output sample, and that is roughly it.
Concrete things worth checking before you decide an entry is a stealth debut:
- No provider attribution on the board row. Established rows carry a lab name. A bare product-style name with nothing behind it is the first tell.
- A name that does not match any known family. Labs version their public lines predictably — a number, a tier, a suffix. Names that read like a mascot rather than
family-version-tierare usually placeholders. - It appears on exactly one board. Attributed releases get submitted broadly. A single-arena appearance suggests a controlled test rather than a launch.
- No callable endpoint anywhere. This is the decisive one. If there is no API, no console, no waitlist, the entry is a measurement exercise, not a product.
- Vote count well below the board average. Elo moves fast at low sample sizes. A new entry sitting at number one on a fraction of the votes of the rows beneath it has a rank that has not stabilised yet — worth reading alongside what an Elo rating actually is.
That last point is the one that trips people up most often, and it deserves its own note: a first-place finish on a thin vote count and a first-place finish on a mature one are not the same claim, even though the board renders them identically.
Why a lab would ship without its name on it
None of the labs running these entries publish their reasoning, so treat what follows as the plausible motives rather than confirmed strategy. They are consistent enough across releases to plan around.
A clean preference read. Brand halo is real. Voters who know a model came from a lab they already trust score it differently from an unlabelled output. Stripping the name is the only way to get a preference signal that measures the model rather than the logo.
An exit that costs nothing. An unnamed entry that lands mid-table can be withdrawn quietly. The same model launched under a lab's flagship line and landing mid-table is a news cycle. Anonymity is optionality.
Timing against competitors. A lab that can see where its unreleased model sits against everything shipping today, before rivals can see the same thing, gets to choose its launch date on better information.
A smaller compliance surface. An entry with no public license, no commercial availability and no named operator attracts less scrutiny than a product. That gap closes the moment it becomes purchasable.
Should you test one before attribution lands?
Almost always no, and the reason is not caution about quality. It is that there is usually nothing to test. A codename entry with no endpoint offers you sample outputs on a leaderboard page and nothing else. You cannot run your own prompts through it, which means you cannot answer the only question that matters — whether it does your job better than what you already run.
Where an endpoint does exist, the calculus changes but not as much as you would hope. The cost of early testing is rarely the credits. It is switching cost: prompts tuned to a model's quirks, presets, aspect-ratio behaviour, an internal style guide written against its output. Rebuilding that around a model that might be renamed, repriced, relicensed or withdrawn is how teams end up doing the same migration twice.
The proportionate response at codename stage is to watch, not to build:
- Record the capability, not the ranking. If the samples show something your current stack cannot do — a duration ceiling, a native audio pass, a text-rendering jump — write down the specific capability. That is the thing that will still be true after the name changes.
- Do not rewrite prompts against it. Prompt tuning is the expensive part and the least portable.
- Wait for the endpoint, the price and the license together. Any one of the three arriving alone is not a launch.
- Re-check the board once votes mature. Ranks built on thin samples move. The rank three weeks later is the one worth acting on.
Where HappyHorse actually landed
Attribution arrived, and then a real product did. HappyHorse 1.1 followed on 23 June 2026 with 1080p output, clips from three to fifteen seconds, and native synchronised audio with multilingual lip-sync — an ordinary, specifiable model release with things you can check against your own requirements. Note that the language spec got vaguer, not sharper, on the way: the 1.0 entries name their seven languages outright, while 1.1 says "multilingual" and leaves it there. A codename becoming a product does not mean every claim behind it becomes checkable.
The board position settled too, and it settled lower than the debut suggested. On Arena's text-to-video leaderboard as of 14 August 2026, HappyHorse 1.0 sits seventh at 1428 Elo, across a board carrying 616,845 votes and 45 models. That is a strong result and a long way from the number-one debut on a different board months earlier. Both readings were accurate at the time. Neither was a stable description of the model.
A caution on which board you are reading, because it matters more this year than last: Arena rebranded from LMArena in January 2026, and Artificial Analysis is a separate company running its own image and video boards. They currently disagree sharply on video — Wan 3.0 leads Artificial Analysis' text-to-video-with-audio board at 1243 while not appearing in Arena's top ten at all. A vendor claiming "the number one video model" in mid-2026 has usually chosen which arena to quote. We wrote up how to read an image arena leaderboard without being fooled; the same discipline applies double to video right now.
Both HappyHorse generations are in the Versely catalog with pages you can read before you spend anything: HappyHorse 1.0 Text to Video, HappyHorse 1.1 Image to Video, and the full HappyHorse provider page. Each model page carries its current Elo, its category rank and its credit cost in the same view, which is the comparison that actually decides a batch — a rank on its own never has been. If you are tracking what has arrived recently rather than what is rumoured, new video models by year is the more useful surface than any single board.
FAQ
Is an anonymous leaderboard entry a sign a model is being hidden for bad reasons?
No. The usual reason is measurement quality: preference votes are contaminated by brand recognition, and the only way to remove that is to remove the name. Treat a codename entry as an experiment being run in public rather than as a product being concealed.
How long does attribution usually take?
There is no reliable interval, and the one example here is not a pattern. HappyHorse 1.0 appeared on Artificial Analysis around 7 April 2026 and Alibaba was identified as the lab in reporting that followed, with the 1.1 release landing 23 June. Plan around "when the endpoint and license appear," not around an expected number of weeks.
If a codename entry ranks first, should I stop my current model migration?
No. A first-place rank on an unattributed entry with no endpoint gives you nothing to migrate to. Finish the migration to something you can call, price and license today, and keep the capability note for when the stealth entry becomes a product.
Does a debut rank tell you anything at all?
It tells you the ceiling of what voters preferred on a small sample of prompts, at one moment, on one board. That is a real signal about capability existing somewhere in the field. It is not a signal about which model you should run next month, which is why the vote count next to the rank deserves as much attention as the rank.