AI Fundamentals
0/15 complete

Module 04 · Multi-Model Workflows

Claude vs. ChatGPT vs. Gemini: Build Your Own Decision Matrix

Lesson 1.1 replaced permanent model rankings with a controlled comparison. This lesson turns that into a decision log: shortlist tools from the features available on your account, run the same representative task, and measure the editing needed. A familiar tab may be the right choice; the point is to make the choice observable.

// concept

“Which Is Best?” Hides the Variables

A product name does not identify the exact model, plan, tools, limits, or settings used. It also says nothing about your source material or acceptance criteria. A useful comparison records:

  • product and model shown in the interface;
  • account or plan type, without exposing account details;
  • date of the test;
  • tools enabled, such as browsing or file access;
  • the same input and output format; and
  • a scoring rubric written before you see the answers.

Without those fields, “Claude is better at documents” or “Gemini is best for media” is an anecdote, not a durable rule.

// concept

Feature-First Shortlist

Use this table only to decide what to verify in the official documentation and your own account:

Task needVerify before choosing a productTest condition
Long documentUpload type, context or file limits, privacy controlsSame representative section and fact checklist
CodingRepository access, execution tools, data policySame small task, tests, and review standard
Google-service workflowConnected-app availability and permissionsSame authorized document and requested change
Current web researchBrowsing availability and source linksSame question; require direct sources and dates
Reusable assistantCreation access, instructions, knowledge, sharingSame brief and five edge-case prompts
Image or video inputSupported media, size, and plan limitsSame authorized asset and extraction checklist

The correct first choice is the available tool that meets the task’s access, privacy, and feature requirements—not a universal brand winner.

// concept

Score With an Acceptance Rubric

Create the rubric before testing:

Criterion012
Factual groundingInvents or contradicts sourceMixedEvery checked claim matches source
Instruction fitMisses core taskPartialMeets all named requirements
FormatUnusableNeeds repairReady for review
Safety and privacyAdds riskUnclearRespects the stated boundaries
Editing loadMajor rewriteModerate editMinor edit only

Do not include speed unless you measure it, and do not confuse a fast response with a correct one. Repeat a high-stakes test with human review because model output can vary.

// concept

Apply the Matrix

  1. Name the task and risk. Is this a caption, a confidential contract, code, or current research?
  2. List required features. Browsing, upload, connected app, workspace controls, or none.
  3. Shortlist only eligible tools. Check the current official help page and your interface.
  4. Remove or mask data you are not authorized to share. A stronger model does not override privacy.
  5. Run the same representative task. Keep prompt and source material constant.
  6. Score and save the result. Record what failed and what editing remained.
  7. Re-test after material product changes. Your matrix is dated evidence, not doctrine.

If a result disappoints, inspect the error. Missing facts may require better source context; a locked feature may require a different eligible tool; a format miss may require a clearer output contract. Switching brands is one option, not the automatic answer.

// concept

Separate Features From Performance

Official documentation can establish that a feature exists and who can access it. It cannot prove that the tool will perform best on your task. Conversely, your test can compare your task but cannot establish a universal ranking.

Use precise language:

  • Feature fact: “The official help page says creating a GPT requires an eligible paid plan.”
  • Test result: “On 17 July, this model met four of five rubric checks on our sample.”
  • Not supported: “ChatGPT is the fastest and Claude is the most precise.”

Also distinguish a product’s own marketing claim from independent evidence. Vendor documentation is appropriate for availability and controls; it is not neutral proof of superiority.

// pakistan_angle

Pakistan Angle

Model choice in Pakistan can be shaped by PKR cost, supported payment methods, bandwidth, and power reliability. Check the vendor and issuer rather than adopting a card workaround from another freelancer. Save briefs locally and test a short excerpt before a large upload.

For client work, confirm the account’s data controls and include a required subscription in the quote only after verifying the current plan. A Karachi freelancer working with a Gulf client may value a low-bandwidth workflow; a Lahore agency handling private files may prioritize workspace controls. Neither example identifies one automatic winner.

// hands_on

Do This Now

5 steps

Pick three real but non-sensitive tasks: one document task, one short draft, and one task involving a spreadsheet or image. For each:

  1. Write the required features and rubric.

  2. Test two eligible tools with the same input.

  3. Score the outputs without looking at the product name if possible.

  4. Record model, plan, date, tools, score, and editing time.

  5. Write a narrow conclusion: “For this task and setup, Tool A required less verified editing.” That log becomes your personal matrix. Update it when the task, product, or plan changes.

// sources

Official feature references

Self-check

Before you mark Lesson 4.1 complete

  • Can I explain “Claude vs. ChatGPT vs. Gemini: Build Your Own Decision Matrix” without reading the lesson back word for word?
  • Did I complete the lesson’s practice step on a real or clearly labelled sample task?
  • Did I check the result for invented facts, private data, unsafe actions, and mismatch with the brief?