Circles and Qwen3.8-Max: frontier models in production speak in numbers
OpenAI publishes numbers from Circles, an MVNO operator that uses OpenAI’s API and Codex to personalize customer experience in telecommunications. The stated result: +22% ARPU and -9% churn, with improved development efficiency. These are the metrics by which a company decides whether a technology stays or goes.
The same day, Alibaba opened the weights for Qwen3.8-Max, the 2.4T-parameter model we covered arriving in US stacks. Two parallel signals: a proprietary case study with business numbers and a frontier open-source with MIT license entering the workflow of those building.
Why it matters. Models, both open and proprietary, stop being a bet and become a line item on the P&L. If your organization is evaluating AI adoption, the right question is which business indicator moves when you put it in production. Benchmarks say little about this. Circles measured ARPU and churn. Qwen3.8-Max gives those with cost constraints or data sovereignty requirements a competitive, downloadable option.
The two paths reinforce each other. The choice becomes operational.
If you want to try it: find a process where you can measure a business indicator before and after introducing AI. Without that number, you’re doing marketing.
In detail
Circles’ case study is interesting because OpenAI rarely publishes such concrete business numbers. Circles is an MVNO (Mobile Virtual Network Operator) active in Asian markets. It uses OpenAI’s API and Codex to personalize user experience, automate part of software development, and manage customer support.
The stated numbers:
- +22% ARPU: each customer generates a fifth more revenue on average.
- -9% churn: abandonment rate drops by nearly a tenth.
- Improved development efficiency: OpenAI doesn’t quantify by how much, and that’s the first limit of the source.
The data comes from the official case study page, so it should be read as marketing material from an interested party. No metric is accompanied by measurement methodology, reference period, or control group. For decision-makers, the signal is interesting but proof is missing: a +22% ARPU increase could include factors unrelated to AI.
On the other front, Qwen3.8-Max is the second Chinese open model in two weeks to enter operational US stacks, after Kimi K3 served via Telnyx API. Alibaba promised the weights for the week following the August 3 announcement, with MIT license. The model is a mixture-of-experts with 2.4T total parameters, 95B active per token, 1M-token context window. Reported benchmarks (PaperBench 93.0, CoWorkBench 74.8) place it in the frontier proprietary range, but they’re self-reported and should be taken with caution.
The convergence that matters is this: a real operator documents measurable financial impact, and simultaneously an open model with frontier-class capability becomes downloadable. For those evaluating AI adoption in their organization, that means having both a reference business case and an option that doesn’t lock you into a single provider.
Limitations remain strong on both sides. The Circles case study is single-source and non-independent: third-party follow-up would be valuable. Qwen’s benchmarks are from the vendor. Qwen3.8-Max’s size (2.4T parameters) makes it non-executable locally for nearly all organizations. You need dedicated inference infrastructure or an API provider. Alibaba offers the API at $2 per million input tokens and $6 for output, with caching at $0.25, but price alone isn’t enough to decide without testing quality on your own use case.
The operational lesson for readers: when evaluating AI’s impact on your work, ask for business numbers and ask how they were measured. A case study without methodology is an anecdote. And a 2.4T-parameter open model is a concrete opportunity only if you have the compute to use it, or if you serve it via API at a cost that makes sense for your actual volume.