ModelGrade
The independent AI agent scorecard

What five AI agents do —
and what they don't.

ModelGrade examines ChatGPT, Claude, Gemini, Grok and Muse across twenty-two criteria in five categories. Each criterion carries its own weight; Core Capabilities is the primary score, and the remaining categories form the overall score. No rankings, no winners — the same evidence, arranged five ways.

Scores

Primary: Core Capabilities · Secondary: Overall · /10

The primary score is the weighted Core Capabilities mean. The overall score weights all five categories (Core 40 · Developer 20 · Ecosystem 15 · Privacy 15 · Cost 10). Agents appear alphabetically; the order carries no meaning.

Full support — documented & available
Partial — limited, gated, or in transition
Not observed — no public evidence
†Provisional — still being verified

How scoring works

Every criterion is scored 10 for full support, 5 for partial or limited support, and 0 where no public evidence was found. Criteria are weighted differently — the weights are shown under each criterion in the table and listed here. Hover or tap any mark for its evidence note.

Method & limits

  • Primary score = weighted mean of Core Capabilities (8 criteria).
  • Overall score = Core 40% · Developer & API 20% · Ecosystem 15% · Privacy 15% · Cost & Access 10%.
  • Evidence from official documentation, pricing pages and certification registries, September 2026.
  • Consumer tiers only — enterprise-only capabilities are out of scope.
  • Raw benchmark scores and unverified controversies are excluded by design.
  • † marks facts still being verified with the vendor; treat as provisional.
  • ModelGrade is independent and unaffiliated with any vendor shown.