Really Good Ads

Thinking20 July 2026 · 8 min read

Which AI chatbot should you use, and for what job?

We get asked this constantly. Six assistants, one job each, and the honest limits of every one.
Really Good AdsPerspective
A chatbot icon rendered as digital pixels
EverydayWritingVisuals Open weightsResearchPrivate

Someone asks us this most weeks, usually phrased as "which one should we actually pay for". The honest answer is that the leading assistants have converged on the boring middle. For drafting an email or summarising a document, you would struggle to pick them apart blind. The differences that matter now are not raw intelligence. They are the features bolted around the model, and the jobs each one has been shaped to do.

So here is the current line up, what each is genuinely good at, and where each falls over. One caveat before we start, and it applies to everything below: this field moves fast enough that specific version numbers go stale within weeks. We have deliberately named capabilities rather than model numbers, because the numbers will have changed by the time you read this.

01. ChatGPT, the daily driver

The one to default to if you are only going to have one. Not because it wins any particular category outright, but because it has absorbed the most jobs. Memory carries context between conversations, so it gets more useful the longer you use it. Projects group related chats into a persistent workspace with their own files and their own memory, which stays separate from your main chat history. Custom GPTs let you package a repeatable task for other people to use.

Its deep research mode is the strongest of the bunch: give it a question and it will browse independently for several minutes and come back with a long, cited report. Since early 2026 you can see and edit its research plan before it starts, which is a meaningful improvement if you know your subject. Agent mode goes further and carries out multi step tasks across the web and connected apps, and the two can be combined.

Do not trust the citations on sight. Deep research modes synthesise from what they retrieve, they do not independently verify it. A Columbia Journalism Review audit found meaningful citation hallucination rates across these tools. The report is a starting point for a human, not a finished document. If the number is going in a client deck, open the source.
A man in front of a wall of documents connected by red string
Deep research, essentially. Hundreds of sources, all connected, and still somebody has to check the string actually goes where it says.

02. Claude, the wordsmith

Where we reach when the writing itself is the deliverable. It holds a tone across a long piece, takes direction on style rather than drifting back to house voice after two paragraphs, and needs less rewriting at the end. That last part is the whole argument: the time saved is in the editing, not the drafting.

It also produces finished files rather than blocks of text you have to reformat. Ask for a document, a spreadsheet or a deck and it builds the thing. Skills are the other reason to look: you teach Claude a task once, as a folder of instructions, and it applies that any time the task comes up in any chat. For an agency that means brand guidelines, tone rules and reporting formats stop being something you paste in every single time. Skills have been available since late 2025 across the free and paid tiers, though they need code execution switched on.

Declared bias. Anthropic makes Claude, and Anthropic makes the assistant that helped write this article. Treat our enthusiasm here with the scepticism you would apply to any other recommendation from an interested party, and go and test it against the others on your own work.
An illustration of William Shakespeare at his writing desk
The wordsmith brief. Tone, register and knowing which word earns its place.

03. Gemini, the creative department

Thinner on workflow features than ChatGPT, but the place to go for anything visual. Its image models, known by the Nano Banana nickname, are exceptional at two specific things: putting legible text inside an image, and editing an existing picture from a plain language instruction while leaving everything else untouched. For mockups, posters and variant testing, that combination is worth more than raw beauty.

The video side is genuinely strong too. Veo generates picture and synchronised audio in one pass, dialogue and effects and ambience together, rather than handing you a silent clip to score afterwards. It comes in a tier range spanning roughly eight fold in price, so draft on the cheap tier and finish on the expensive one.

"Best image generator" needs qualifying. On photorealism, Google's own Imagen line still tests better, and Midjourney has more artistic range. Nano Banana Pro wins on text rendering, editing control and speed, which is a different claim. Note also that Google applies SynthID watermarking to output, which matters if you are putting AI work in front of clients who ask.
Pablo Picasso painting a canvas in his studio
A creative department of one. The tool never had the idea, and it still does not.

04. Kimi, the open weight challenger

Moonshot's Kimi is the one that keeps embarrassing the assumption that frontier AI is a Western duopoly. It is open weight, free to use through its own chat interface, handles large batches of uploaded files comfortably, and is oddly good at building presentations. On agentic and coding benchmarks it has traded blows with the leading closed models and won some of them.

The headline you may have seen is out of date. The "Kimi beats every Western model" story went round when its thinking model launched in late 2025. Moonshot has shipped at least two generations since. The current picture is more interesting and less tidy: it leads on some agentic and coding benchmarks and trails the top closed models on general intelligence, maths and reasoning indices. Very strong, not a clean sweep.

05. Perplexity, the researcher

Every assistant can search the web now, so the case for a dedicated research tool is narrower than it was. It still holds, though, because Perplexity assumes you are in research mode and never stops citing. When you want a specific figure with a source attached rather than a confident paragraph, it is the fastest route.

The free year is mostly gone. You may have seen that PayPal, Revolut and various telecoms give Perplexity Pro away as a perk. Some of that was real, but the widely shared PayPal and Venmo twelve month offer closed at the end of 2025, and the carrier deals rotate by region and expire without much warning. Student, education and some government routes are still live. Check inside your own account rather than chasing a deal you read about, and treat a free year as a bonus rather than a reason to choose it.

06. Llama, the one that stays in the building

Meta's Llama runs on your own hardware, so nothing you feed it leaves your machine. For analysing a confidential document set, or anything covered by a client NDA, that property is worth more than a few benchmark points. The trade off is the obvious one: no live internet, no real time information, and you are responsible for the hardware.

Two things to check before committing. The Llama licence has carried restrictions that have excluded companies domiciled in the EU, so if you have European entities, read it properly rather than assuming open weights means unrestricted. And Llama is no longer the only sensible local option: Qwen, DeepSeek, Gemma and Kimi all ship open weights now, and several outperform it on specific jobs. "Runs locally" is the category, not the product.

So what do we actually tell people?

ChatGPT for everyday work and deep research. Claude for writing and finished documents. Gemini for images and video. Kimi if you want frontier capability without a subscription. Perplexity when you need a sourced fact. Llama, or another open weight model, when the data cannot leave the building.

The tool does not have the taste. It never has. What it has is your afternoon back, and what you do with that is still the job.

The pattern underneath all six is the same one we keep arriving at. None of these replaces the judgement about what is worth making, who it is for, or whether it is any good. They compress the mechanical middle of the work: the first draft, the reformatting, the reading of forty tabs. That is a real gain and it is worth paying for. It is just not the thing people keep claiming it is.

Our practical advice is to pick two, not six. One daily driver and one specialist for whatever your team actually does most. Six subscriptions nobody has learned properly is worse than one everybody has.

Want a hand working out which two? Talk to us.
Checked July 2026 against vendor documentation and release notes from OpenAI, Anthropic, Google, Moonshot AI, Perplexity and Meta, plus independent benchmark reporting. Model names and capabilities change frequently, so specific version numbers have been avoided deliberately. Verify current features and pricing before making procurement decisions. Film, television and photographic stills are used editorially and remain the property of their respective rights holders.
← Back to News