AI Atlas Part 1: Assistants, research and science
12 ratings across 3 categories — which tools make the shortlist for assistants, research and science, and where their limits are.
In short
For assistants, research and science, the internal AI Atlas rates 12 products in 3 categories by feature fit, controllability and workflow integration — based on vendor documentation, not on an own product test.
At a glance
- 3 categories and 12 rated products in this part.
- Rating: feature fit 50%, controllability 25%, workflow integration 25%, 1–5 points each.
- Method: review of official vendor documentation, not an own lab test.
- Survey date: 19 September 2026; 45 categories, 95 products, 123 ratings in total.
- Every category names a concrete comparison test readers can run themselves.
Part 1 of the AI Atlas series sorts out which tools belong on the shortlist for assistants, research and science. The basis is an internal comparison dated 19 September 2026: 45 categories, 95 products, 123 ratings. Suitability was rated from the vendors' official feature documentation — feature fit 50 percent, controllability 25 percent, workflow integration 25 percent, 1 to 5 points each. This is explicitly not an own product test and not a measurement of output quality.
01 · General AI Assistants & LLMs
Take ChatGPT and Claude into the shortlist together; Gemini in particular where Google workflows are already in place. No intelligence winner was measured here.
| Product | Rating | Best suited for | Limitation | Cost driver |
|---|---|---|---|---|
| ChatGPT | ★★★★★ 5 | Mixed knowledge work and finished work files | Mode and the tools you have subscribed to affect the result. | Plan and usage limits |
| Claude | ★★★★★ 5 | Long texts, documents and multi-step knowledge work | Sources, permissions and results have to be checked. | Plan and usage limits |
| Gemini / Gemini API | ★★★★½ 4.5 | Multimodal applications and Google-adjacent use | Check the specific model, the modalities and access separately. | Model-dependent API usage / app plan |
| Mistral Vibe, vormals Le Chat | ★★★★ 4 | Alternative chat and coding environment | Do not equate it with earlier Le Chat plans without checking. | Plan / API separately |
| Grok | ★★★★ 4 | Conversation and timely topics as a research entry point | Current information is not automatically reliable. | Plan / API separately |
| DeepSeek API | ★★★★ 4 | Custom API applications | The application has to be built in-house; check quality and data path separately. | Usage-based API |
| Qwen-Modellfamilie | ★★★★ 4 | Choosing a suitable open model yourself | Check weights, license and hardware requirements for each specific model. | Hosting and compute / API |
Comparison test: Give every candidate an identical brief with sources, contradictions and the desired output file.
02 · Web Research & Sourcing
Perplexity for searching the open web; Gemini Notebook for a defined set of your own sources. The recommendations address different tasks.
| Product | Rating | Best suited for | Limitation | Cost driver |
|---|---|---|---|---|
| Perplexity | ★★★★½ 4.5 | Fast entry point with linked original sources | Citations have to actually support the specific claim. | Plan and research quota |
| Gemini Notebook / NotebookLM | ★★★★★ 5 | Questions about a defined set of sources | Completeness and quality depend on the underlying set of sources. | Quota / workspace access |
| ChatGPT | ★★★★ 4 | Research followed by a written work-up | Mode and the tools you have subscribed to affect the result. | Plan and usage limits |
Comparison test: Require ten verifiable statements; open every original source and count the correct supporting passages.
03 · Science & Literature Review
Elicit for structured extraction and review workflows; Consensus for literature questions and citation trails.
| Product | Rating | Best suited for | Limitation | Cost driver |
|---|---|---|---|---|
| Elicit | ★★★★★ 5 | Screening studies systematically and extracting them into tables | Study selection and extractions need expert review. | Plan / usage pool |
| Consensus | ★★★★½ 4.5 | Finding literature and tracing connections | A summary is no substitute for assessing a study's methodology. | Plan / search quotas |
Comparison test: Provide a known set of studies; count missing studies, incorrect figures and unsupported conclusions.
The series has nine parts; the remaining parts carry the numbers 2, 3, 4, 5, 6, 7, 8, 9. Stars stand for documented suitability, not for a winner: in several categories the comparison deliberately recommends two candidates side by side because they solve different tasks.
FAQ
Is this an own product test?
No. It rates documented suitability taken from official vendor documentation. Output quality, speed and value for money in daily use are not measured here.
How are the stars calculated?
From three sub-scores of 1 to 5: feature fit counts 50 percent, controllability and workflow integration 25 percent each. The result is rounded to half stars.
Why are two products sometimes tied?
Because they solve different tasks. In those cases the comparison names both on purpose instead of inventing a winner.
Sources
- ChatGPT — Hersteller
- Claude — Hersteller
- Gemini / Gemini API — Hersteller
- Mistral Vibe, vormals Le Chat — Hersteller
- Grok — Hersteller
- DeepSeek API — Hersteller
- Qwen-Modellfamilie — Hersteller
- Perplexity — Hersteller
- Gemini Notebook / NotebookLM — Hersteller
- Elicit — Hersteller
- Consensus — Hersteller