Compare agent surfaces by autonomy, host, approval model, and rollout risk.

Each candidate agent surface links autonomy, host, approval model, rollout risk, and its current evidence-review status.

Agent surface

ChatGPT Work

AutonomyCan change code

For teams, Business closes much of the gap between personal chat subscriptions and a governed workspace by combining admin controls, shared context, and connectors.

Risk: Coding is better than general assistants used to be, but still not as IDE-native as Cursor.

Needs review · low confidence · Aug 14, 2026Evidence to review: ChatGPT pricing
Agent surface

Claude Code

AutonomyCan change code

For teams, Claude fits best where a smaller group of heavy users can justify Premium seats while the rest stay on lower-cost Standard seats.

Risk: Team pricing scales quickly once a subset of users needs Premium seats for heavier Claude Code usage.

Needs review · low confidence · Aug 14, 2026Evidence to review: Introducing Claude Sonnet 5
Agent surface

Unified Agents Window

AutonomyCan change code

For engineering teams, Cursor now has enough orchestration and admin structure to be a deliberate team purchase, but the price only pays back if parallel agents, shared MCP workflows, or local-cloud handoff become a normal part of how the team ships.

Risk: It is still a weak fit for writing, meetings, and general knowledge work outside engineering.

Needs review · low confidence · Aug 13, 2026Evidence to review: Cursor site
Agent surface

Support agent builder

AutonomyPolicy-grounded response

For support teams, CustomGPT.ai is strong because the product, trial path, and pricing all point toward customer-service deployment rather than general AI experimentation.

Risk: It is a specialist buy, so the seat is harder to justify if the same budget also needs broad writing, meetings, and research coverage.

Agent surface

Advanced Capabilities

AutonomyCan change code

For teams, Devin becomes compelling when ticket throughput, migrations, and backlog clearing matter more than just code suggestions inside the editor.

Risk: It is not a cheap default coding seat for every developer.

Needs review · low confidence · Aug 7, 2026Evidence to review: Devin pricing
Agent surface

Agent mode

AutonomyCan change code

For teams, it makes the most sense when Cloud and platform work matter alongside day-to-day IDE assistance.

Risk: It is less GitHub-native than Copilot and less editor-opinionated than Cursor or Windsurf.

Needs review · low confidence · Aug 10, 2026Evidence to review: Gemini Code Assist site
Agent surface

Coding agent

AutonomyCan change code

For teams, Copilot Business is easiest to approve when leaders want AI inside existing GitHub and IDE workflows without paying Cursor-level seat prices.

Risk: Less opinionated and less immersive than Cursor for agent-first IDE work.

Needs review · low confidence · Aug 10, 2026Evidence to review: GitHub Copilot pricing
Agent surface

Assistant and agents

AutonomyWorkspace action execution

For teams, Glean becomes compelling when company knowledge already spans many systems and the team wants one permission-aware layer for search, assistant, and agents.

Risk: It is sales-led and enterprise-oriented, so it is harder to justify for small teams that do not yet have meaningful cross-system knowledge sprawl.

Needs review · low confidence · Aug 10, 2026Evidence to review: Glean product overview
Agent surface

Grok Build

AutonomyCan change code

For teams, Grok becomes interesting when employee demand is real and the team has one bounded pilot: connector-backed automation, a reviewed voice-bot call flow, or a Grok Build workflow. The Business seat is only one cost line; Voice Agent Builder and API usage need their own cap and owner.

Risk: The public business surface is still narrower than ChatGPT or Claude on connector breadth and workflow maturity.

Needs review · low confidence · Aug 13, 2026Evidence to review: Introducing Grok 4.5
Agent surface

Pre-built agents and Copilot Studio

AutonomyWorkspace action execution

For teams, Copilot Business is most compelling when Outlook, Teams, Word, Excel, and SharePoint are already the operating surfaces and meeting follow-through matters every day.

Risk: The paid Copilot layer still requires a qualifying Microsoft 365 base plan, so total seat economics can climb quickly.

Needs review · low confidence · Aug 18, 2026Evidence to review: Microsoft 365 Copilot Business site
Agent surface

Notion Agent

AutonomyWorkspace action execution

For teams, Notion AI becomes compelling when leaders want search, meeting notes, research, and execution to happen inside one shared workspace instead of scattering across chat tools and docs.

Risk: It is weaker than ChatGPT or Claude as a standalone general assistant outside the Notion workspace.

Needs review · low confidence · Aug 13, 2026Evidence to review: Notion AI site
Agent surface

Replit Agent

AutonomyCan change code

For teams, Replit is strongest for prototyping and lightweight product delivery where browser access, built-in deployment, connectors, and Agent-led security work accelerate iteration.

Risk: It is not the best choice for deep local-codebase IDE workflows compared with Cursor or Windsurf.

Needs review · low confidence · Aug 13, 2026Evidence to review: Replit pricing