Best fit
Teams, writers, developers, researchers and buyers deciding which AI subscription or API to use.
ChatGPT alternatives
By Daniel Reeve · tested and updated September 2026
Many users ask which AI is best, but the better question is which AI wins your task. A comparison workflow makes the answer visible instead of relying on vibes, leaderboards or brand loyalty.
Quick answer
Do not compare models by reading comparisons, including this one. Twenty minutes with one of your own prompts will tell you more than any ranking, and it stays true when the models change.
Decision map
Teams, writers, developers, researchers and buyers deciding which AI subscription or API to use.
People who only need a quick casual answer.
Answer this task. Then list assumptions, failure points and how another model might disagree.
Step by step
| Area | Useful for | Watch out for |
|---|---|---|
| Prompt design | Use the same prompt and same source material across models. | Changing the prompt changes the test. |
| Scoring | Score correctness, usefulness, specificity, tone, citations and format. | Do not let a confident style hide wrong facts. |
| Second-pass critique | Ask a different model to critique the best answer. | Critique also needs human review. |
| Source verification | Use search-focused tools when facts matter. | Open official or primary sources yourself. |
| Repeatability | Run the test on multiple examples from your real work. | One lucky answer does not prove a tool is best. |
The workflow
This is the method behind every recommendation on this site, reduced to something you can run yourself this afternoon. It is deliberately small — a test you actually complete beats a benchmark you admire.
Keep the results
Record the date, the model version string and the exact prompt alongside your scores. These systems change often enough that an unlabelled result is worthless within a quarter, and you will not remember which version you tested.
Re-run the same frozen prompts in three months. The change over time is often more useful than the original ranking — it tells you which vendor is actually improving on the work you do.
Sources
Every figure on this page was read from the official documentation below on 2026-09-10. Prices, limits and model names change without notice — the source is authoritative, this page is not.
FAQ
About twenty minutes for one task across two or three models. Anything much longer and you will not repeat it, which matters because these systems change often enough that a one-off comparison ages quickly.
Two or three, from different providers. Beyond that the marginal information drops sharply and the scoring gets sloppy — and sloppy scoring is worse than no comparison, because it feels like evidence.
Edit distance: minutes of your work between the answer and something you would send or merge. It correlates with real-world usefulness far better than any impression of quality formed while reading.
Because “ChatGPT” and “Gemini” are product names, not models, and what sits behind them changes. A result without a version string cannot be reproduced or fairly compared against a later run.