AI Atlas Part 2: Coding, apps and automation
12 ratings across 4 categories — which tools make the shortlist for coding, apps and automation, and where their limits are.
In short
For coding, apps and automation, the internal AI Atlas rates 12 products in 4 categories by feature fit, controllability and workflow integration — based on vendor documentation, not on an own product test.
At a glance
- 4 categories and 12 rated products in this part.
- Rating: feature fit 50%, controllability 25%, workflow integration 25%, 1–5 points each.
- Method: review of official vendor documentation, not an own lab test.
- Survey date: 19 September 2026; 45 categories, 95 products, 123 ratings in total.
- Every category names a concrete comparison test readers can run themselves.
Part 2 of the AI Atlas series sorts out which tools belong on the shortlist for coding, apps and automation. The basis is an internal comparison dated 19 September 2026: 45 categories, 95 products, 123 ratings. Suitability was rated from the vendors' official feature documentation — feature fit 50 percent, controllability 25 percent, workflow integration 25 percent, 1 to 5 points each. This is explicitly not an own product test and not a measurement of output quality.
04 · Programming & Repository Work
Test Codex and Claude Code on well-scoped project tasks; Cursor for work in an AI editor; Copilot for a GitHub-centric team.
| Product | Rating | Best suited for | Limitation | Cost driver |
|---|---|---|---|---|
| OpenAI Codex | ★★★★★ 5 | Repository tasks through to reviewed changes | Production readiness only comes from real tests and review. | Plan / usage limits |
| Claude Code | ★★★★★ 5 | Terminal and repository workflows | Environment access, tests and approvals have to be set up. | Plan / API usage |
| Cursor | ★★★★★ 5 | Interactive development in an AI editor | Switching models and agent loops change cost and behavior. | Plan / model consumption |
| GitHub Copilot | ★★★★★ 5 | Existing developer teams working in GitHub | Repository rules and human code review remain necessary. | Plan / premium usage |
Comparison test: Have three real issues solved with hidden regression tests; only reviewed changes count.
05 · Websites, Apps & Prototypes
v0 for visually controlled web interfaces; Lovable or Replit for an integrated entry into app building. No automatic seal of production readiness.
| Product | Rating | Best suited for | Limitation | Cost driver |
|---|---|---|---|---|
| v0 | ★★★★★ 5 | Web UI with visible editing and deployment | Test authentication, data access and payment logic separately. | Generation plus hosting |
| Lovable | ★★★★½ 4.5 | Fast full-stack prototype | A working preview does not prove a secure production setup. | Generation plus cloud services |
| Replit Agent | ★★★★★ 5 | Building and running an app in one environment | Ongoing hosting costs and access rules are part of the project. | Agent usage plus hosting |
Comparison test: Test login, role separation, mobile view, export and restore on the same mini project.
21 · Automation & Cross-System Agents
n8n for technically supervised custom flows, Make for visual integration, Zapier Agents for connected SaaS tasks. The model is only one component.
| Product | Rating | Best suited for | Limitation | Cost driver |
|---|---|---|---|---|
| n8n | ★★★★★ 5 | Custom processes with control over executions | Self-hosting shifts operation and maintenance to us. | Hosting / executions / model costs |
| Make | ★★★★★ 5 | Visually built integration flows | Error paths, loops and operations have to be planned. | Plan / credits / model costs |
| Zapier Agents | ★★★★★ 5 | Agents using existing SaaS connections | Autonomous actions need clearly limited permissions. | Plan / activities / additional services |
Comparison test: Process an inquiry into a CRM draft; test duplicates, an API outage, retries and human approval.
44 · Testing, Monitoring & Operating Agents
LangSmith for observability; Databricks for AI inside the data platform. Both need target cases and error thresholds of their own.
| Product | Rating | Best suited for | Limitation | Cost driver |
|---|---|---|---|---|
| LangSmith | ★★★★★ 5 | Traces and investigation of agent runs | Logs alone do not prove quality; target cases and sign-offs are additionally required. | Traces / storage / plan |
| Databricks AI | ★★★★½ 4.5 | AI operations within an existing data platform | Data engineering and domain evaluation remain work of their own. | Cloud / compute / platform |
Comparison test: Inject known faults: wrong source, timeout, duplicate job, unauthorized tool; check detection and recovery.
The series has nine parts; the remaining parts carry the numbers 1, 3, 4, 5, 6, 7, 8, 9. Stars stand for documented suitability, not for a winner: in several categories the comparison deliberately recommends two candidates side by side because they solve different tasks.
FAQ
Is this an own product test?
No. It rates documented suitability taken from official vendor documentation. Output quality, speed and value for money in daily use are not measured here.
How are the stars calculated?
From three sub-scores of 1 to 5: feature fit counts 50 percent, controllability and workflow integration 25 percent each. The result is rounded to half stars.
Why are two products sometimes tied?
Because they solve different tasks. In those cases the comparison names both on purpose instead of inventing a winner.