Module 04 · Multi-Model Workflows
Claude vs. ChatGPT vs. Gemini: Build Your Own Decision Matrix
Open lesson + course map
On this lesson
Course outline
Module 1 · Foundational Mindset
Module 2 · Mastering Context Threads
Module 3 · Building Custom GPTs and Gems
Module 4 · Multi-Model Workflows
Module 5 · From User to Operator
Lesson 1.1 replaced permanent model rankings with a controlled comparison. This lesson turns that into a decision log: shortlist tools from the features available on your account, run the same representative task, and measure the editing needed. A familiar tab may be the right choice; the point is to make the choice observable.
// concept
“Which Is Best?” Hides the Variables
A product name does not identify the exact model, plan, tools, limits, or settings used. It also says nothing about your source material or acceptance criteria. A useful comparison records:
- product and model shown in the interface;
- account or plan type, without exposing account details;
- date of the test;
- tools enabled, such as browsing or file access;
- the same input and output format; and
- a scoring rubric written before you see the answers.
Without those fields, “Claude is better at documents” or “Gemini is best for media” is an anecdote, not a durable rule.
// concept
Feature-First Shortlist
Use this table only to decide what to verify in the official documentation and your own account:
| Task need | Verify before choosing a product | Test condition |
|---|---|---|
| Long document | Upload type, context or file limits, privacy controls | Same representative section and fact checklist |
| Coding | Repository access, execution tools, data policy | Same small task, tests, and review standard |
| Google-service workflow | Connected-app availability and permissions | Same authorized document and requested change |
| Current web research | Browsing availability and source links | Same question; require direct sources and dates |
| Reusable assistant | Creation access, instructions, knowledge, sharing | Same brief and five edge-case prompts |
| Image or video input | Supported media, size, and plan limits | Same authorized asset and extraction checklist |
The correct first choice is the available tool that meets the task’s access, privacy, and feature requirements—not a universal brand winner.
// concept
Score With an Acceptance Rubric
Create the rubric before testing:
| Criterion | 0 | 1 | 2 |
|---|---|---|---|
| Factual grounding | Invents or contradicts source | Mixed | Every checked claim matches source |
| Instruction fit | Misses core task | Partial | Meets all named requirements |
| Format | Unusable | Needs repair | Ready for review |
| Safety and privacy | Adds risk | Unclear | Respects the stated boundaries |
| Editing load | Major rewrite | Moderate edit | Minor edit only |
Do not include speed unless you measure it, and do not confuse a fast response with a correct one. Repeat a high-stakes test with human review because model output can vary.
// concept
Apply the Matrix
- Name the task and risk. Is this a caption, a confidential contract, code, or current research?
- List required features. Browsing, upload, connected app, workspace controls, or none.
- Shortlist only eligible tools. Check the current official help page and your interface.
- Remove or mask data you are not authorized to share. A stronger model does not override privacy.
- Run the same representative task. Keep prompt and source material constant.
- Score and save the result. Record what failed and what editing remained.
- Re-test after material product changes. Your matrix is dated evidence, not doctrine.
If a result disappoints, inspect the error. Missing facts may require better source context; a locked feature may require a different eligible tool; a format miss may require a clearer output contract. Switching brands is one option, not the automatic answer.
// concept
Separate Features From Performance
Official documentation can establish that a feature exists and who can access it. It cannot prove that the tool will perform best on your task. Conversely, your test can compare your task but cannot establish a universal ranking.
Use precise language:
- Feature fact: “The official help page says creating a GPT requires an eligible paid plan.”
- Test result: “On 17 July, this model met four of five rubric checks on our sample.”
- Not supported: “ChatGPT is the fastest and Claude is the most precise.”
Also distinguish a product’s own marketing claim from independent evidence. Vendor documentation is appropriate for availability and controls; it is not neutral proof of superiority.
// pakistan_angle
Pakistan Angle
Model choice in Pakistan can be shaped by PKR cost, supported payment methods, bandwidth, and power reliability. Check the vendor and issuer rather than adopting a card workaround from another freelancer. Save briefs locally and test a short excerpt before a large upload.
For client work, confirm the account’s data controls and include a required subscription in the quote only after verifying the current plan. A Karachi freelancer working with a Gulf client may value a low-bandwidth workflow; a Lahore agency handling private files may prioritize workspace controls. Neither example identifies one automatic winner.
// hands_on
Do This Now
5 steps
Pick three real but non-sensitive tasks: one document task, one short draft, and one task involving a spreadsheet or image. For each:
Write the required features and rubric.
Test two eligible tools with the same input.
Score the outputs without looking at the product name if possible.
Record model, plan, date, tools, score, and editing time.
Write a narrow conclusion: “For this task and setup, Tool A required less verified editing.” That log becomes your personal matrix. Update it when the task, product, or plan changes.
// sources
Official feature references
3 official sources — check every claim yourself
Self-check
Before you mark Lesson 4.1 complete
- Can I explain “Claude vs. ChatGPT vs. Gemini: Build Your Own Decision Matrix” without reading the lesson back word for word?
- Did I complete the lesson’s practice step on a real or clearly labelled sample task?
- Did I check the result for invented facts, private data, unsafe actions, and mismatch with the brief?