AI Content Creation
0/24 complete

Module 03 · Audience Research

Mining Comments and DMs for Content Gold

20 minfocused lesson7practical steps4grounded questions4source links
Open lesson + course map

On this lesson

Course outline

Comments and DMs contain first-party questions, objections, and natural wording. Mining turns those signals into a private, anonymized question bank—without copying inboxes, scraping profiles, or publishing screenshots.

You will build 20 redacted questions, a theme table supported by signal IDs, and a reply, content, permission, or exclusion decision for each.

// concept

Set the Permission Boundary Before Collection

Use only material you are entitled to review: comments on accounts you manage and messages sent to those accounts, subject to your organisation's policy. From another creator's public comments, record only a generalized topic and the post URL—never the commenter's identity or story. Do not automate collection or bypass platform controls.

Visibility is not consent. Ask before showing a username, quote, screenshot, voice note, order story, or identifying detail. DMs stay private; public commenters have not agreed to be featured. Without recorded permission, use only a generalized question.

Choose one handling path:

SignalSafe defaultContent use
Your post's commentRedact identity; keep link separatelyGeneralize; ask before quoting
DM to your accountTranscribe the minimum; never export the chatParaphrase; ask before quoting or showing
Another account's public commentLog topic and post URL, not identityOne signal, not proof of demand
CNIC, phone, address, payment, medical, child detailExclude; handle privatelyNever send to AI or content

// concept

Run a 20-Minute Weekly Harvest

  1. Minutes 0–2 — scope: choose one date window and one or two channels.
  2. Minutes 2–10 — collect: scan your own surfaces manually for questions, objections, and repeated wording. On YouTube, use YouTube Studio > Community > Published; the official help says you can filter for questions. Confirm current menus on the official page.
  3. Minutes 10–15 — redact locally: remove handles, profile links, order numbers, locations, faces, and contact details. Do not upload screenshots.
  4. Minutes 15–18 — route: mark each row reply, content, ask_permission, or exclude.
  5. Minutes 18–20 — cluster: group redacted rows; keep one-off questions as singletons.

Use this schema. Keep the source map separately; never put it in AI input.

FieldExample valueRule
signal_idIG-2026W29-04Working ID, not a user ID
channelInstagram commentNo handle or profile link
collected_on2026-07-18Date observed
redacted_questionDo you deliver outside [AREA]?Preserve meaning, remove identity
safe_phrasedelivery chargesNon-identifying wording
permissionnot requestedQuote/feature permission
sensitivitylowlow, review, or exclude
pathreply + contentHandling path

// concept

Redact First, Then Ask AI to Cluster

Redaction is removal, not blurring. A screenshot can still expose a photo, phone, shop, timestamp, location, or distinctive wording. Transcribe the question, then check whether combined details identify the sender.

Before ChatGPT, review its official Data Controls. Turning off model improvement does not create consent. Manual spreadsheet clustering is a free alternative.

Paste only redacted records into this prompt:

// prompt — copy me22 lines
Role: You are a qualitative research assistant.
Input: JSON lines containing signal_id, channel, redacted_question,
safe_phrase, permission, sensitivity, and path.

Tasks:
1. Flag any row that may still contain a person, phone, CNIC, address,
   order number, account handle, workplace, or uniquely identifying story.
2. Cluster the remaining rows by the need expressed, not by keywords alone.
3. Create a theme only when at least two distinct signal_ids support it.
4. Keep unsupported rows under "Singletons"; do not force a cluster.
5. For each theme, return: theme, supporting IDs, safe audience phrases,
   neutral question, suggested reply, and one content angle.

Rules:
- Do not infer age, gender, city, income, intent, or sentiment.
- Do not claim that "the audience" wants something.
- Never reconstruct removed details or produce a quote from a DM.
- Describe frequency only as "N of N collected signals in this date window."

<signals>
[PASTE REDACTED JSON LINES]
</signals>

// concept

Turn Themes Into Replies and Content Paths

A cluster is not a market conclusion. Write “delivery cost appeared in 3 of 20 collected signals from 12–18 July,” not “Pakistanis care most about delivery.” One post or group may dominate.

Preserve non-identifying phrases such as “delivery charges” for hooks, but paraphrase distinctive DM wording. Produce both a direct reply and a generalized content brief. If the value depends on a personal story, choose ask_permission.

// worked_example

Worked Example

This fictional sample for a Rawalpindi home-baking creator uses six invented signals, not real customers.

// json6 lines
{"signal_id":"IG-01","channel":"Instagram comment","redacted_question":"Do you deliver outside [AREA]?","safe_phrase":"deliver outside","permission":"not requested","sensitivity":"low","path":"reply + content"}
{"signal_id":"WA-02","channel":"WhatsApp DM","redacted_question":"What are delivery charges to [AREA]?","safe_phrase":"delivery charges","permission":"not requested","sensitivity":"low","path":"reply + content"}
{"signal_id":"IG-03","channel":"Instagram comment","redacted_question":"How early should I order for an event?","safe_phrase":"how early","permission":"not requested","sensitivity":"low","path":"reply + content"}
{"signal_id":"WA-04","channel":"WhatsApp DM","redacted_question":"Can I book for [EVENT_DATE]?","safe_phrase":"book for","permission":"not requested","sensitivity":"low","path":"reply + content"}
{"signal_id":"IG-05","channel":"Instagram comment","redacted_question":"Is an egg-free option available?","safe_phrase":"egg-free option","permission":"not requested","sensitivity":"review","path":"reply + content"}
{"signal_id":"WA-06","channel":"WhatsApp DM","redacted_question":"My child has [MEDICAL_DETAIL]; which item is safe?","safe_phrase":"","permission":"not requested","sensitivity":"exclude","path":"exclude"}

The prompt produces this sample output, checked manually:

ThemeSupportWhat the evidence permitsReply/content path
Delivery coverage and costIG-01, WA-022 of 6 ask about deliveryReply with current details; make a delivery FAQ
Order lead timeIG-03, WA-042 of 6 ask when to orderCheck capacity; make a lead-time explainer
Ingredient suitabilityIG-05 onlySingletonAnswer availability; make no safety claim
Sensitive health detailWA-06ExcludedHandle privately; make no recommendation

Draft one failed because its DM screenshot exposed a profile photo, phone, date, and landmark despite a blurred name. The creator removed the screenshot, deleted that AI conversation using the tool's current controls, transcribed only “Can I book for [EVENT_DATE]?” locally, and reran the cluster. The resulting lead-time post used no sender quote or story.

// failure_cases

Failure Cases to Diagnose

6 cases to diagnose

  • Theme from one loud comment

    one ID is labelled “common.” Move it to Singletons.

  • False anonymity

    a row removes the name but keeps a phone, CNIC, workplace, landmark, or rare story. Exclude or generalize it.

  • Public-means-permitted assumption

    a post draft displays a public comment screenshot. Replace it with a generalized question or request explicit permission to feature the person.

  • Keyword-only grouping

    “price” merges delivery, product price, and budget. Re-cluster by the underlying decision.

  • Audience-wide claim

    the caption says “everyone wants COD” when the bank only records a few questions. State the collection window and sample count, or present it simply as an FAQ.

  • Reply forgotten after mining

    the creator turns a DM into an idea but leaves the sender unanswered. Write the direct reply first; content is the second path.

// pakistan_angle

Pakistan Angle

Pakistani inboxes often contain phone numbers, WhatsApp names, easypaisa or JazzCash details, delivery landmarks, and CNIC images. Keep them private. Never paste payment screenshots, CNICs, receipts, rider locations, or full WhatsApp chats into AI. “Can I pay through [PAYMENT_METHOD]?” is enough for research.

Language can identify people too. Paraphrase a distinctive Roman-Urdu voice note, neighbourhood reference, or rare Urdu phrase. Keep ordinary wording such as “delivery charges kitne hain?” only if it cannot identify anyone. Cluster “COD hai?” with “cash on delivery available?” by meaning while retaining the language label for hook writing. Treat Karachi, Lahore, Peshawar, and smaller-city signals separately until your own bank supports combining them.

// hands_on

Hands-On Exercise

7 steps

Build your first mined-and-clustered question bank.

  1. Choose a seven-day window from an account you manage.

  2. Manually collect at least 20 questions or objections; never scrape or bulk-export an inbox.

  3. Apply the schema locally. Remove identities and mark sensitive rows exclude before AI use.

  4. Set permission and path. Public comments default to not requested for quotes; DMs stay private.

  5. Run the prompt on redacted rows, or cluster manually.

  6. Check every cluster against its IDs; move unsupported themes to Singletons.

  7. Add one direct reply and one generalized content angle per supported theme. You are done when the bank contains 20 or more anonymized entries, no private screenshot or direct identifier, a theme table tied to signal IDs, and a documented reply/content/permission/exclusion path for every row.

// completion_rubric

Completion Rubric

6 checks — tick as you verify

0/6

// sources

Sources

// check_yourself

Check yourself

4 questions · answers and options are taken word-for-word from this course

0/4
  1. 1 / 4 · diagnose

    Your work shows this failure mode: “Public-means-permitted assumption.” The lesson describes it like this: “A post draft displays a public comment screenshot.” What does the lesson tell you to do about it?