Someone asks us this most weeks, usually phrased as "which one should we actually pay for". The honest answer is that the leading assistants have converged on the boring middle. For drafting an email or summarising a document, you would struggle to pick them apart blind. The differences that matter now are not raw intelligence. They are the features bolted around the model, and the jobs each one has been shaped to do.
So here is the current line up, what each is genuinely good at, and where each falls over. One caveat before we start, and it applies to everything below: this field moves fast enough that specific version numbers go stale within weeks. We have deliberately named capabilities rather than model numbers, because the numbers will have changed by the time you read this.
The one to default to if you are only going to have one. Not because it wins any particular category outright, but because it has absorbed the most jobs. Memory carries context between conversations, so it gets more useful the longer you use it. Projects group related chats into a persistent workspace with their own files and their own memory, which stays separate from your main chat history. Custom GPTs let you package a repeatable task for other people to use.
Its deep research mode is the strongest of the bunch: give it a question and it will browse independently for several minutes and come back with a long, cited report. Since early 2026 you can see and edit its research plan before it starts, which is a meaningful improvement if you know your subject. Agent mode goes further and carries out multi step tasks across the web and connected apps, and the two can be combined.
Where we reach when the writing itself is the deliverable. It holds a tone across a long piece, takes direction on style rather than drifting back to house voice after two paragraphs, and needs less rewriting at the end. That last part is the whole argument: the time saved is in the editing, not the drafting.
It also produces finished files rather than blocks of text you have to reformat. Ask for a document, a spreadsheet or a deck and it builds the thing. Skills are the other reason to look: you teach Claude a task once, as a folder of instructions, and it applies that any time the task comes up in any chat. For an agency that means brand guidelines, tone rules and reporting formats stop being something you paste in every single time. Skills have been available since late 2025 across the free and paid tiers, though they need code execution switched on.
Thinner on workflow features than ChatGPT, but the place to go for anything visual. Its image models, known by the Nano Banana nickname, are exceptional at two specific things: putting legible text inside an image, and editing an existing picture from a plain language instruction while leaving everything else untouched. For mockups, posters and variant testing, that combination is worth more than raw beauty.
The video side is genuinely strong too. Veo generates picture and synchronised audio in one pass, dialogue and effects and ambience together, rather than handing you a silent clip to score afterwards. It comes in a tier range spanning roughly eight fold in price, so draft on the cheap tier and finish on the expensive one.
Moonshot's Kimi is the one that keeps embarrassing the assumption that frontier AI is a Western duopoly. It is open weight, free to use through its own chat interface, handles large batches of uploaded files comfortably, and is oddly good at building presentations. On agentic and coding benchmarks it has traded blows with the leading closed models and won some of them.
Every assistant can search the web now, so the case for a dedicated research tool is narrower than it was. It still holds, though, because Perplexity assumes you are in research mode and never stops citing. When you want a specific figure with a source attached rather than a confident paragraph, it is the fastest route.
Meta's Llama runs on your own hardware, so nothing you feed it leaves your machine. For analysing a confidential document set, or anything covered by a client NDA, that property is worth more than a few benchmark points. The trade off is the obvious one: no live internet, no real time information, and you are responsible for the hardware.
ChatGPT for everyday work and deep research. Claude for writing and finished documents. Gemini for images and video. Kimi if you want frontier capability without a subscription. Perplexity when you need a sourced fact. Llama, or another open weight model, when the data cannot leave the building.
The pattern underneath all six is the same one we keep arriving at. None of these replaces the judgement about what is worth making, who it is for, or whether it is any good. They compress the mechanical middle of the work: the first draft, the reformatting, the reading of forty tabs. That is a real gain and it is worth paying for. It is just not the thing people keep claiming it is.
Our practical advice is to pick two, not six. One daily driver and one specialist for whatever your team actually does most. Six subscriptions nobody has learned properly is worse than one everybody has.