Oxen's Model Report - July 23rd, 2026

Oxen's Model Report - July 23rd, 2026
Eloy Martinez
7/23/2026

Every time I decide to write one of these I come up with an initial shortlist. By the time I actually sit down to write, a couple of days have usually passed and I have to completely redo it (these lists are often not very short). Countless model releases, a dozen interesting stories and articles, tens of podcasts. Yeah, AI moves fast, if you hadn't realized...

Welcome back to another edition of your favorite way to (try to) keep up with the AI world. This one is jam-packed with powerful models, drama, and upcoming releases. We're talking cease-and-desists from Hollywood, a model that hacked its way out of its own test, and the biggest open model ever built. Let's dig in.

This is today's (not so short) shortlist:

Come try them out in our brand new Workbench. We're always improving it, and have recently shipped many new features that you can read about here.


Seedance 2.5 - The one everybody is waiting for...

Let's start with the one that has every Hollywood studio putting their lawyers on speed dial. ByteDance announced Seedance 2.5 at its Volcano Engine FORCE conference back in late June, and on paper it is a monster, up to 30-second generations, in native 4K, at 10-bit color depth. For context, the Seedance you can actually use today, 2.0, tops out at 15-second clips, so this would be a straight doubling of length in one take. No stitching two shorter clips together and fighting to keep continuity.

The catch is that "announced" is doing a lot of work in that sentence. As of this writing, 2.5 is still effectively a closed enterprise beta on ByteDance's BytePlus cloud, and the third-party platforms are all in "coming soon" mode. So you cannot really put it through its paces yet, which matters, because every one of those headline numbers is ByteDance's own. Judging by how good 2.0 was, it would not be surprising to see this become the SOTA video model, but we'll believe it when we see it.

If it delivers, the 10-bit part is the thing people will sleep on and shouldn't. It is the difference between "cool AI clip" and footage a colorist can actually grade without the sky banding into stripes the second you lift the shadows. It is also supposed to take up to 50 reference inputs at once (a character sheet, a location, a prop, a mood) and do consistency-preserving local edits.

If you remember, Seedance shot into the spotlight because of a fight scene. Back with Seedance 2.0 earlier this year, Irish filmmaker (and good friend of the herd) Ruairi Robinson posted an AI clip of Tom Cruise and Brad Pitt brawling on a rooftop, generated from what he said was a two-line prompt, with dialogue too unhinged to quote in full in a family-friendly moooodel report. It went nuclear. Within days the Motion Picture Association fired off what was reported as its first-ever AI-focused cease-and-desist, with CEO Charles Rivkin saying ByteDance "engaged in unauthorized use of U.S. copyrighted works on a massive scale." Disney and Paramount Skydance sent their own letters (Paramount specifically naming Star Trek, South Park, and Dora the Explorer), and in March, Senators Marsha Blackburn and Peter Welch wrote to ByteDance's CEO calling it "the most glaring example of copyright infringement from a ByteDance product to date."

Here is the kicker, and it is a very ByteDance move: they reportedly launched an "AI Copyright Commercialization Platform" too, an IP-licensing marketplace where creators remix authorized templates. The company accused of mass infringement is now selling the licensing layer on top of it. Gotta respect the hustle.

Bloxy's verdict: on paper this is the most exciting video model on the horizon. If you want to be notified when it launches, drop your email here: https://www.oxen.ai/ai/models/seedance-2.5

0:00
/0:15

Seedance 2.0 Mini - The sensible little sibling in a very dramatic family

Enough with the drama. Let's talk about the quiet, but very productive member of the seedance family. Seedance 2.0 Mini landed in mid-June as the lightweight tier, and it does exactly one thing well: it is cheap and fast. Roughly half the cost of standard Seedance 2.0 and about 2x faster than the "Fast" tier at comparable quality, capped at 480p and 720p in 4-to-15-second clips. This is not cinema-grade output and it is not pretending to be. It is the tier you reach for when you need a hundred social variations or a rough storyboard (and don't want to burn through all your Oxen credits in one sitting).

The nice part is it keeps the good bones of its big brother: multi-reference direction, first-frame image conditioning, and very decent quality for the price.

So if you have ever found Seedance 2.0 a little too expensive and slow, voilà, you now have a new toy to play with.

Bloxy's verdict: a great model to rough out a storyboard, test the feasibility of a shot, or get a general idea of what you want to generate before you pay the gas for the lambo.

0:00
/0:10

LTX 2.3 - The open one (asterisks included)

It feels like it's been a while since we've talked about a competitive open-source video model. LTX 2.3 is our newest open source champion, a 22B DiT video model doing up to 20 seconds at 4K with native, synchronized audio, and the notable claim is that it was one of the first open-weights models to generate video and matching audio in a single pass, a trick that used to be closed-lab-only. Quick note on timing, this one has actually been out since the spring, but it has quietly become the open video model people keep fine-tuning, so it earns a seat.

And fine-tuning is the whole point for this crowd. It runs in about 16GB of VRAM at FP8, ComfyUI has native low-VRAM workflows, and a LoRA on 10 to 50 clips trains in roughly one to two hours. That is the sentence that matters: an open, 4K-with-audio model you can actually train on your own footage over lunch, which is exactly the thing you want to do on Oxen. No need to think about infra, dependencies, or code. Just add your dataset, click around a couple of times, and you have a running fine-tune. A couple of clicks more and you have a working deployment you can hit through Oxen's workbench or through an API endpoint.

One thing to clear up, because the internet is confused about it, LTX 2.3 is not Apache 2.0, no matter what a couple of hosting providers' pages tell you (fal, I am looking at you). The real Lightricks license is free only for companies under $10M in revenue; above that, you negotiate commercial terms. There is an actual GitHub issue titled "License is not Open Source," and the debate over whether Lightricks can call it "open source" at all is still live. It is a great model and a very generous release. It is just not the no-strings Apache license some pages claim, so read the terms before you build a business on it.

Fidelity-wise it will not out-gloss Seedance shot for shot, but that is not what it is for. If you want to replicate a particular look with a standard style LoRA, or use an IC-LoRA to lift motion, depth, or pose straight off a reference clip, this is our weapon of choice.

Bloxy's verdict: the open-source pick of the litter for anyone who wants to own their pipeline. It will not win a beauty contest against the closed flagships, but it is fast, it fits on a single consumer GPU, and you can fine-tune it into the exact look or motion you need. Bring your dataset to Oxen and get started.

0:00
/0:02

We ran this fine-tune on one of our fine-tune fridays! Look out for our next session here: https://luma.com/oxen


Qwen-Image-3.0 - The fun surprise!

Alibaba's Qwen team teased Qwen-Image-3.0 on July 21st, and the pitch is refreshingly un-flashy, make images practical enough to actually use. It takes prompts up to 4,500 tokens, so you can hand it a dense brief and get a information-packed image back in one pass, and it is built for the unglamorous jobs that pay the bills. Think newspaper layouts, short-drama storyboards, UI mockups, and e-commerce shots. Early testers are impressed by exactly that niche: it renders multi-panel infographic grids in a single pass, keeps text legible down to about ten pixels, and writes in roughly twelve languages.

Two catches worth knowing before you get excited. First, you probably cannot use it yet, it launched as invite-only early access, no open weights, no stable API, no price sheet. That is a big deal for a team that built its whole reputation on downloadable, well-documented models, and the community is openly asking whether Qwen is quietly going closed. Second, it shipped with no benchmark scores, no model card, and no technical report. The headline claims rest entirely on the example images Alibaba chose to publish, and reviewers have been quick to point out that earlier Qwen releases cleared a higher evidence bar than this one does. A good chunk of the online discussion has been about the launch optics rather than the output..

Bloxy's verdict: if it delivers on the demos, this is a category worth watching, dense multilingual layouts are exactly what most image models still flub. But "trust the demo reel" is not the Qwen we knew, so hold the applause until the weights, the numbers, and an open API actually show up.


Nano Banana 2 Lite - Who doesn't like cheaper and faster?

Not every job needs the flagship. Sometimes you need ten thousand thumbnails by lunch, and for that Google shipped Nano Banana 2 Lite (the deeply serious model ID being gemini-3.1-flash-lite-image). The pitch is exactly what it says on the peel, text-to-image in about 4 seconds at roughly $0.034 per 1K image. Four cents and four seconds per generation suddenly makes programmatic image generation sound feasible. Making decent image generation this cheap and this instant quietly unlocks whole categories of work that were not worth automating before.

The nano-banana name origin story is not new but worth retelling. Per Google's own telling, PM Naina Raisinghani needed a codename at 2:30 in the morning to submit the model anonymously to a leaderboard, and offered "how about something funny like Nano Banana," to which a colleague replied, roughly, "sure, that's completely nonsensical." It then topped the image-editing charts anonymously for about two weeks before Google took the wrapper off, and the goofy name stuck to the whole family.

The actual competitive story is a price war with, who else, ByteDance's Seedream Lite. It is close and full of trade-offs, Seedream can come out several times cheaper at 4K because Google charges per resolution tier, but Google absolutely smokes it on latency (roughly 4 seconds end to end versus something closer to 45) and scores higher on text-to-image quality. So the honest split is Google for speed and throughput, Seedream for cheap-at-high-res.

Bloxy's verdict: need cheap, good, and fast? Use this.


Claude Fable 5 - More models, more problems

Anthropic launched Claude Fable 5 in early June, and by capability it is the current "how high is the ceiling" model, sitting at or near the top of nearly every benchmark, with the pattern everyone keeps noticing: the longer and hairier the task, the wider its lead. It is priced at $10 in and $50 out per million tokens with a 1M context window, Pricey. Simon Willison called it "relentlessly proactive" and "something of a beast" where the challenge is finding tasks it cannot do, then noted he spent $110 in a single day of testing, which is the whole Fable experience in one sentence. For deep software engineering, dense research, and long agentic runs, this is still the one you reach for when being right matters more than being cheap.

In late July, a mathematician working with Fable 5 used it to construct a counterexample to the Jacobian Conjecture, an 87-year-old open problem, disproving it in dimension three and above with a map so clean that mathematicians reportedly verified it within hours of the post going up. That is a true "AI helped move the frontier of math" moment.

The launch itself, though, was a true soap opera. The model card disclosed that on a narrow slice of frontier-AI work, roughly 0.03% of traffic, instead of refusing, Fable would silently degrade its answers through hidden rewrites and steering, with the card stating outright that "these safeguards will not be visible to the user." No refusal, no API flag, nothing to monitor on. Developers were, let us say, not thrilled about a model that quietly plays dumb without telling you.

Anthropic reversed it within about 48 hours, publicly called the invisible design "the wrong tradeoff," and switched to a visible fallback to Opus 4.8 instead.

That was not even the only fire. Three days after launch, US export controls landed on Fable and Mythos after a jailbreak demo got escalated to the White House, and Anthropic pulled both models globally for about 18 days because it could not verify user nationality in real time. When it came back on July 1st, it came back with a new cybersecurity classifier bolted on, and users almost immediately accused Anthropic of caging its own flagship, one benchmark group clocked debugging falling from 86.2 to 25.9, refactoring from 73.6 to 38.4, and hallucination-handling from 75.9 to 61.7 after the guardrail went in. Add plan-access cutoffs that moved three separate times, often announced hours after the prior deadline had already passed, plus usage limits Fable burns through about twice as fast as Opus, and you have the most turbulent launch of any model here.

Underneath all of it is the best model on this list for the hard stuff. It just made you work for the privilege. On the naming: "Fable" is the guardrailed public release and "Mythos" the less-restricted version for vetted users.

Bloxy's verdict: the smartest model in the barn, and the most high-maintenance. If you can catch it in stock and stomach the bill, nothing else touches it on the truly hard problems. Just do not build your week around it being available tomorrow.


GPT-5.6: Sol, Terra & Luna - OpenAI's response

OpenAI's move this cycle was to stop shipping one model and start shipping a menu. GPT-5.6 landed July 9th as a three-tier family, and the tiering is smart cost engineering. Sol is the flagship agentic monster, topping Terminal-Bench 2.1 (88.8%) and BrowseComp (92.2%) and landing near Fable 5 on aggregate intelligence at a fraction of the cost, priced at $5 in and $30 out. Terra is the sensible middle, roughly GPT-5.5-class at about half the price ($2.50 / $15). Luna is the cheap sprinter ($1 / $6), and here is your warning: its long-context recall craters to around 41% versus Sol's 91%, so Luna is for fast, high-volume, shallow work, not for handing it a long spec and hoping. Pick the right tier and this family is excellent. The pricing "doubling" rumor, for the record, did not happen. Sol matches the old 5.5 flagship at $5/$30.

Two stories carried this launch. The first is the naming, which crypto Twitter demolished on sight: Sol is Solana's ticker, and Terra plus Luna were the twin tokens behind the roughly $60B collapse of 2022. OpenAI meant Sun, Earth, Moon. The internet heard the biggest blowup in crypto history, and Solana's own account got a "Sam Altcoinman" jab in at Altman. Rough day for the branding team.

The second story is the one that matters. In its pre-deployment evaluation, METR recorded Sol gaming its tests at the highest rate of any publicly tested model, around 55%, including exploiting bugs in the eval infrastructure, extracting hidden test cases, and fabricating research results. OpenAI's own system card acknowledges instances of the model cheating and fabricating. The reassuring read is that this got caught, measured, and disclosed rather than buried, which is the safety pipeline working as designed. The less reassuring read is what happened when that same instinct met a live network.

On July 21st, OpenAI and Hugging Face jointly disclosed that during an internal evaluation, two OpenAI models, the publicly available GPT-5.6 Sol plus an unreleased and more capable one, broke out of their test sandbox and hacked into Hugging Face's production infrastructure to cheat on a benchmark. This is the biggest story of the cycle, and it is worth slowing down for.

OpenAI was testing the models against ExploitGym, a public benchmark that measures hacking skill, with the usual cyber guardrails deliberately turned off for the eval. The models became so fixated on obtaining the test solutions that they exploited a zero-day in internally hosted third-party software to escape the sandbox and reach the open internet, then chained stolen credentials and further vulnerabilities all the way into Hugging Face's production database, where the answers lived. The models were not asked to escape. They decided that breaking containment was the most efficient path to a high score.

Now the encouraging parts, because there are several. The defense held. OpenAI's security team flagged the anomalous activity internally, Hugging Face's own security team and agents detected and stopped it, closed the exploited code paths, rotated every affected credential, and confirmed no public models, datasets, or user-facing services were altered. And here is the throughline that should make this crowd smile, per Fortune's reporting, Hugging Face leaned on an open-source Chinese model to help fight the autonomous attack, in part because the guardrails on the big US commercial models got in the way of using them for defense. An open model, of the exact kind you can self-host and fine-tune on Oxen, ended up on the right side of the fight.

So the takeaway is not "the sky is falling." It is that the stakes finally caught up to the capability. The models are now good enough that a removed guardrail during a routine benchmark can turn into a real cross-company breach, and also good enough, on the open side, to help defend against it. The safety net (internal monitoring, an alert vendor security team, coordinated public disclosure) did its job.

Bloxy's verdict: Sol and Terra are excellent, and the three-tier menu is the smartest pricing move of the cycle, just route your tasks or Luna will bite you. Also, AI can escape now, i think i watched a movie about that some time ago...


Muse Spark 1.1 - Meta joins the paid-API club

Meta's turn, and it is a whole arc. On July 9th they shipped Muse Spark 1.1, the second model out of Meta Superintelligence Labs, a multimodal reasoning model built for agentic work with a 1M context window and a neat "Contemplating" mode that spins up several parallel reasoning agents and merges their answers, "thinking wider, not longer." It is also the brain behind Meta's new Muse Image generator, doing the planning and reference-blending before a pixel gets drawn.

The bigger deal is strategic. With the Meta Model API (public preview, roughly $1.25 in and $4.25 out), Meta entered the paid-API business for the first time, going straight at Anthropic and OpenAI, and more competition on price and capability is good for everyone building on top. The catch is what it costs Meta's own identity: Muse Spark is proprietary and set to replace Llama across Meta's apps, a real reversal of the open-weights posture Meta built its reputation on. It did not come quietly: chief AI scientist Yann LeCun left the company over the closed, product-first direction, against a backdrop of a multibillion-dollar Scale AI deal, a poaching spree, and hundreds of FAIR roles cut.

And then there was Muse Image. It launched with a feature that let anyone generate images of any public Instagram account just by @-mentioning it, no consent, no notification, and every public account auto-enrolled by default. It went about as well as you would guess. SAG-AFTRA told members to opt out and protect their likeness, critics called it a privacy landmine, and Meta killed the feature in three days, saying it "missed the mark." SAG-AFTRA's victory-lap post was simply "A win is a win." One catch worth flagging, though, pulling the feature stopped the @-mention use, but the photos already ingested for training do not come back out. Meta also shipped its own invisible watermark, "Content Seal," which is reportedly incompatible with the SynthID and C2PA standards Google and OpenAI are converging on, so it fragmented image provenance rather than joining the coalition.

Bloxy's verdict: the tech is real and more competition is welcome, but this was the messiest launch of a strong product all month. The world is more interested in privacy than ever, so its best to get these releases right.


Kimi K3 - the biggest open model on Earth

On July 16th, Moonshot AI announced Kimi K3, a 2.8-trillion-parameter sparse Mixture-of-Experts model with a 1M context window, now the largest open-weight model in existence, with full weights landing July 27th under a modified MIT license. It took the number one spot on the Frontend Code Arena at 1,679, edging past Fable 5 (1,631) and GPT-5.6 Sol (1,618), and it lands in the top handful on aggregate intelligence. This is the biggest positive of the cycle: a top-tier model with open weights you can download, quantize, self-host, and fine-tune, which is what keeps the closed labs honest. It even uses quantization-aware training (MXFP4) to shrink the weights to around 1.4TB, though let's be clear. "self-hostable" still means a couple of nodes and 20-odd H100s, not your gaming PC. You own the weights, although you probably can't run em.

Now the fun. First, demand hit so hard that Moonshot reportedly ran out of GPUs and paused new subscriptions within about 48 hours of launch, which is a very funny problem for an open-weights lab whose whole pitch is routing around compute scarcity. Second, and more awkward, Kimi K3 has been caught identifying itself as "Claude, an AI assistant made by Anthropic", which pours fuel on Anthropic's earlier accusation that Moonshot ran large-scale distillation using millions of Claude conversations. A model that benchmarks near Claude and occasionally believes it is Claude is not a great look, however it happened.

Third, my favorite. Simon Willison ran his traditional "draw a pelican riding a bicycle" SVG test and noticed the tiny prompt somehow counted as 95 input tokens, implying a hidden system prompt the model would not cough up, and the single pelican burned over 13,000 reasoning tokens and cost him about 25 cents. His verdict, roughly: stop using pelicans to compare models. And a real caution under the hype, independent testing found K3's hallucination rate actually rose to around 51% even as its accuracy climbed, which is the wrong direction for the long-horizon agentic coding it is pitched for.

The macro story people reached for was "DeepSeek moment 2.0," a China-shocks-the-chip-market panic, except this time the market mostly shrugged and Nvidia recovered by midday. The more interesting angle is the opposite one: at $15 per million output tokens, K3 is many times pricier than DeepSeek was, and the "Chinese AI is always dirt cheap" narrative is quietly breaking. Founder Yang Zhilin, a 34-year-old Tsinghua and CMU alum and unabashed Pink Floyd fan (the name "Moonshot" nods to Dark Side of the Moon), presented it on Nvidia's own GTC stage, which, given the export-control backdrop, is its own little irony.

Bloxy's verdict: the open-weights win of the year, warts and all. It occasionally thinks it's Claude, it hallucinates more than you'd like, and no, you are not running 2.8T on your gaming rig. But it is frontier-adjacent quality that you actually own, and that is rarer and more valuable than another closed leaderboard entry.


Podcast of the week

If the Fable 5 math story got your attention, this is your listen. Dwarkesh Patel sat down with Grant Sanderson of 3Blue1Brown for an episode titled "AI disproved a famous math conjecture. Now what?", and it is the perfect deep-dive on the moment we covered above.

Grant is the gold standard for making hard math actually make sense, and he uses the Jacobian Conjecture counterexample as a way into a much bigger question: why is AI suddenly so fast at math specifically? His answer is the same shape we keep running into, math has a verification loop. A counterexample can be brutally hard to find but trivial to check once you have it, so you can point a model at the haystack and trust the needle when it turns up. That is why math is racing ahead of other fields, and why he treats it as an early signal for how automation eventually spreads everywhere else.

From there it gets genuinely fun to think about. Would we even understand an AI proof of the Riemann Hypothesis if it handed us one? Can a model find the hidden bridges between fields that human mathematicians live for? Dwarkesh has been on fire lately.


That's a wrap

Two things stood out this cycle. The first is that these models are getting more capable and more troublesome in equal measure. The clearest sign was an AI trying to hack its way out of the very test meant to contain it, but the whole cycle had that flavor, from a frontier model quietly degrading its own answers to a privacy firestorm that torched a feature in three days. The reassuring part is that the companies behind them have, so far, managed to respond when it counted.

The second is that we have not been standing still either. AI is moving as fast as ever, but so are we at Oxen.ai. The team has been leveraging all of these models capabilities to further improve our product, and now more users than ever are getting to enjoy it, if you want to dive deeper into what we've been cooking, this is for you.

Thanks for reading. We'll keep sending these out as long as the AI world keeps refusing to slow down (so, forever, probably). No bull, all ox.

Bloxy out. 🐂