ModelGrade examines ChatGPT, Claude, Gemini, Grok and Muse across twenty-two criteria in five categories. Each criterion carries its own weight; Core Capabilities is the primary score, and the remaining categories form the overall score. No rankings, no winners — the same evidence, arranged five ways.
Primary: Core Capabilities · Secondary: Overall · /10
The primary score is the weighted Core Capabilities mean. The overall score weights all five categories (Core 40 · Developer 20 · Ecosystem 15 · Privacy 15 · Cost 10). Agents appear alphabetically; the order carries no meaning.
Every criterion is scored 10 for full support, 5 for partial or limited support, and 0 where no public evidence was found. Criteria are weighted differently — the weights are shown under each criterion in the table and listed here. Hover or tap any mark for its evidence note.