Radar · 24/07/2026 · happened on 23/07/2026 · models

Mollick maps the models: two choices for real work, permissions as the first line of defense

Ethan Mollick publishes the summer update to his guide on which AI to use for which task. The center of gravity has shifted: using AI today means running agents with computer access, not just back-and-forth chatting.

Why it matters to you. Mollick divides the field into three tiers. For low-stakes questions (a recipe, a letter draft), any free model works fine. For high-stakes issues (a second medical or legal opinion), you need the most powerful models: Claude Opus or GPT-5.6 Sol with high thinking level. For real agent work, the options for those who don’t want to wrestle with configurations are two: ChatGPT or Claude, at 20 dollars a month.

Mollick tested both on a real task: preparing an MBA seminar from his Gmail inbox. In ten minutes, both searched the web, built teaching materials, and drafted responses to a colleague. ChatGPT sent the email without asking, because an old permission setting was still active. The lesson: authorization settings matter as much as model choice.

If you want to try it. Open ChatGPT Work or Claude Cowork, connect just one non-sensitive application, and give it a real but low-risk task, with all permissions set to “ask first”.

In detail

The paradigm shift Mollick describes is what we’ve seen mature in recent weeks: agents have moved from demos to daily work. Codex reached 7 million users, Claude Code updates almost daily, and the question has become “how much do I let it do?” instead of “can I trust it?”.

Three tiers, three choices.

Mollick’s framework is pragmatic. First tier: low-stakes chat. Free models are all sufficient; choose whichever interface you like. Second tier: issues where getting it wrong is costly. Here the answer is the most powerful model you can access, with high thinking level. Claude Opus on one side, GPT-5.6 Sol on the other. They cost, and the price gap per usable output can be enormous, but the error margin is lower.

Third tier: agent work, where AI takes control of tools and data. Here Mollick is clear: for those who don’t want to build their own infrastructure, ChatGPT and Claude are the only two reasonable options at 20 dollars a month. You can save money going elsewhere, but it takes expertise most people don’t have.

The extra computer.

The concrete distinction of this edition is between two modes: the AI company gives you a virtual computer (ChatGPT Work, Claude Cowork), or you give it access to yours. The first is simpler and less powerful. The second opens more possibilities but more risks.

Mollick’s MBA seminar prep experiment shows both sides. Both systems understood the task, searched the web, produced materials. But ChatGPT sent a real email without asking, because a permission granted earlier was still active. It’s the perfect illustration of why authorization is the first place to look.

Mollick also flags the risk of prompt injection: an agent reading your mail and browsing the web can run into text written by others trying to manipulate it. Models have become more resistant, but the problem isn’t solved. Another reason to limit what the agent can touch and keep permissions active on everything it sends, spends, or deletes.

Where the guide is weaker.

Mollick admits that product naming from both companies is confusing and poorly documented. Work and Cowork, Sol and Opus and Fable, the various thinking levels: names change and official explanations stay vague. The guide helps you find your way, but readers should expect to need to experiment.

Reference to open-weight models is almost absent. Mollick mentions them but discourages them for those seeking ease of use, because the infrastructure to use them as agents still requires too much work. It’s consistent with what we’ve seen on Laguna S 2.1 and Kimi K3: the quality gap closes, but the operational gap stays open.

To decide with your own data and tasks, rather than trusting the general guide, you can follow the playbook on comparing two models in fifteen minutes: same task, same test cases, results table.

Type to search across course, playbooks, skills, papers…