Self-reported attribution means asking every new lead or customer how they found you, storing the answer on their record and reading it next to your click-based data. Done well, it’s the cheapest way to see the podcasts, word of mouth, communities and AI answers that tracking never records. Done badly, it produces a pile of “Google” and “social media” answers nobody can act on.
What software attribution can’t see
Click-based attribution, whether GA4, CRM source fields or an ad platform, records visits it can tie to a person or device. It’s blind to the moment someone decided you were worth a look, unless that moment was a trackable click:
- A podcast episode heard on a run
- A colleague saying “we use them, they’re good” in a meeting
- A link shared in a Slack group, WhatsApp chat or DM, which lands as direct traffic (dark social)
- A LinkedIn post or YouTube video watched without clicking
- A ChatGPT or Perplexity answer that named you, followed by a direct visit days later
- A conference talk, a newsletter mention or a review site comparison
Each time, demand gets created somewhere invisible and captured somewhere visible, usually brand search or direct. Click-based tools credit the capture; self-reported attribution asks about the creation. The MMM vs multi-touch attribution post places it in the wider measurement stack; this post is the build.
Open text vs dropdown
The field type decides what you’ll be able to learn, so choose it on purpose.
| Open text | Dropdown | Dropdown + follow-up | |
|---|---|---|---|
| Finds sources you didn’t list | Yes | No | Partly, via “Other” |
| Effort for the respondent | Higher | Lowest | Low |
| Analysis effort | Needs tagging | None | Light |
| Detail (which show, which person) | Often included | Lost | Captured in the follow-up |
| Best for | B2B lead forms, low to mid volume | High-volume stores, if kept short | Most ecommerce and self-serve SaaS |
My default is open text wherever volume is low enough for detail to matter, which covers most B2B demo and contact forms. One answer like “heard your founder on a RevOps podcast” beats a hundred clicks on “Podcast.” AI tagging has removed most of the cost of reading open text, the old argument against it.
If you use a list:
- Keep it to 6-10 options, written the way buyers think. “Instagram,” “TikTok,” “YouTube” and “Friend or colleague,” not “Paid social” and “Referral.”
- Split vague buckets. “Search engine” and “Social media” tell you nothing. Give “ChatGPT or another AI assistant” its own option.
- Add conditional follow-ups. “Podcast” opens “Which show?” and “Search engine” opens “What did you search for?”
- Always include “Other (please specify)” and read those answers. New channels show up there first.
- Randomize option order if your tool supports it, with “Other” pinned last, so the top option doesn’t win by position.
For wording, “How did you first hear about us?” nudges people toward the original source rather than the last thing they clicked. Whatever you pick, keep it fixed; changing it mid-stream breaks every trend line.
Where to ask: form, post-purchase or onboarding
Ask at the moment of highest intent, once per buyer, in a spot where the question can’t block a conversion.
| Placement | Fits | Setup | Watch out for |
|---|---|---|---|
| Lead or demo form | B2B, sales-led | Required open-text field | Check form conversion for a few weeks after adding it |
| Post-purchase page | Ecommerce and DTC | Optional survey on the order confirmation page | Never put it inside checkout |
| Onboarding | Self-serve and product-led SaaS | One question in the welcome flow, after the account exists | Don’t add it to the signup form itself |
| Sales discovery call | B2B with buying committees | Rep asks and logs it in the CRM the same day | Answers typed in from memory weeks later |
For B2B, ask twice where you can: on the form, and again in discovery with “How did your team first come across us?” The person filling in the form is often not the one who found you. Long, multi-person journeys get their own treatment in B2B attribution for long sales cycles.
Store the answer on the record that carries revenue: the contact and deal in HubSpot or Salesforce, the order or customer in Shopify. Create dedicated fields for the raw answer and its category. Don’t overwrite HubSpot’s original source property or Salesforce’s Lead Source field; you need both values side by side. Most post-purchase survey apps can write answers to Shopify orders or customers or export them; confirm that before picking one.
Categorizing responses at scale
Open text only pays off once it’s tagged into a fixed taxonomy. Use two levels: a category you report on, and a detail field that keeps the specific show, person or community.
| Category | Raw answers it catches | Detail to keep |
|---|---|---|
| Podcast | “heard you on a marketing podcast” | Show name |
| Word of mouth | “my cofounder,” “a friend at another agency” | Relationship, if given |
| Community / dark social | “Slack group,” “Reddit thread,” “WhatsApp” | Community name |
| Organic social | “your LinkedIn posts,” “saw a video” | Platform, person |
| Paid social | “Instagram ad,” “sponsored post” | Platform |
| Search | “Google,” “searched for X” | Query, brand vs non-brand |
| AI assistant | “ChatGPT recommended you” | Assistant |
| Events and press | “your talk,” “a newsletter” | Event or outlet |
| Unknown | “don’t remember,” “n/a” | None |
A tagging workflow that holds up:
- Pull new answers weekly from the CRM or store: raw text plus record ID.
- Send each answer to an LLM with the taxonomy, a one-line definition per category and 15-20 real labeled examples. Ask for category, detail and a confidence level as structured output.
- Route low-confidence answers and anything tagged “Other” to a human review queue.
- Write the tags back to the record, next to the untouched raw answer.
- Spot-check a random sample monthly, and add misfiled answers to the prompt as new examples.
Set a rule for multi-source answers. “Heard you on a podcast, then googled you” gets tagged Podcast, because the earliest source mentioned is the one click-based data can’t see. Version the taxonomy, and when you split a category, re-run history so trends stay comparable. An n8n or Zapier workflow calling an LLM API can run all of it.
Reconciling self-reported with click-based data
The insight comes from the comparison. Once a month, build a crosstab: self-reported category down the side, click-based original source across the top, counting deals or new customers.
A hypothetical example with made-up numbers, 190 new B2B deals with a usable answer:
| Self-reported ↓ / Click source → | Direct | Organic search | Paid search | Paid social |
|---|---|---|---|---|
| Podcast | 22 | 14 | 6 | 1 |
| Word of mouth | 30 | 9 | 3 | 0 |
| Search | 4 | 38 | 21 | 0 |
| Paid social | 3 | 2 | 1 | 18 |
| AI assistant | 12 | 5 | 1 | 0 |
How to read it:
- Podcast, word of mouth and AI assistant land mostly in Direct and Organic search. Demand was created off the grid and captured by a brand search or typed URL, so click-based reports give those channels no credit.
- Search answers that arrived through search agree with the click data. Check the query detail: if most are brand searches, the real discovery source sits upstream.
- Paid social mostly matches paid social clicks. Where both lenses agree, you can act on the click data with more confidence.
Two adjustments before you use the shares. Respondents may differ from non-respondents, so compare the click-based source mix of both groups; if they look similar, extrapolating is reasonable. And leave “Unknown” out of share calculations rather than spreading it across channels.
Using it for budget decisions
Self-reported data tells you whether a channel deserves budget and a proper test, not how to set bids. For channels you pay for but can’t click-track, estimate a self-reported cost per customer. A hypothetical example: a podcast sponsorship costs $8,000 a month. Of 800 new customers, 400 answered the survey and 40 named the podcast. That’s 10% of respondents; extrapolated to all 800, about 80 customers, or roughly $100 each. Compare that with blended CAC, not with the podcast’s trackable clicks, which will be close to zero.
| Pattern | Likely meaning | Action |
|---|---|---|
| Self-reported high, click-based near zero (podcasts, communities) | Demand creation tracking can’t see | Protect the budget; confirm with an on/off period or holdout |
| Click-based high, self-reported low (retargeting, brand search) | Capturing demand created elsewhere | Test incrementality before scaling |
| Word of mouth rising steadily | Product and customer experience doing the work | Invest in referrals and reviews; don’t credit paid for it |
| AI assistant mentions rising | Buyers shortlisting inside AI answers | Fund AI search visibility work |
| Both lenses high | The channel works | Scale in steps |
Two guardrails: judge on a trailing 90 days, and don’t act on a category with only a handful of answers. My working rule, not a standard, is at least 30 responses before moving money.
Setting up the fields, the tagging and this monthly comparison is a standard part of my marketing attribution work.
Common mistakes
- Changing the question without versioning. New wording or options make old and new data incomparable.
- Counting responses instead of revenue. Join answers to deals and orders, so a channel bringing small buyers doesn’t look as good as one bringing large ones.
- Overwriting click-based fields. Keep both values; the gap between them is the insight.
- Filling the field after the deal closes. Reps guessing from memory adds noise.
- Never reading “Other.” That’s where unknown podcasts and communities show up.
- Treating shares as ROI. Memory favors recent and memorable touches. Use the data for direction and confirm big moves with tests.
Get it built
If your reports credit direct traffic and brand search for demand you suspect started somewhere else, I can add self-reported attribution, automate the tagging and build the monthly comparison. Start with a Growth Audit, $1,500 fixed and credited if we continue. See pricing or get in touch.