Recursive Systems AI Search Visibility Report

Sample report built from a real measurement: one question, 47 logged-out runs across two assistants on August 5, 2026, with at least 20 usable runs per assistant. A client deliverable also measures many phrasings of the question; this sample measures one and says so plainly. All firms are anonymized in this public sample. Client reports name every firm.

AI visibility pilot.

Chicago personal injury market

Measured
August 5, 2026
Version
1.0
Prepared by
Recursive Systems LLC
Collection
41 of 47 runs usable

01What we found

  1. ChatGPT answers this question with a map, and the map decides the winners. A widget with star ratings appeared in 21 of 21 answers. The four firms anchored most often in it were named in 19 or more of 21 runs.
  2. Prominence is not visibility. Two of the market's most recognized firms, with major verdict histories, appeared in zero of 21 runs. Two more appeared once. The firms that dominate these answers are boutiques with 4.8 to 5.0 star ratings.
  3. The two assistants barely agree. Only one firm was strongly visible on both. Three firms ChatGPT named in at least 19 of 21 runs were named zero times in twenty Perplexity runs. Any single "AI visibility score" would hide this completely.
  4. The mechanism looks like review infrastructure, not legal marketing. Most ChatGPT answers linked to Yelp and to Reddit, and the map runs on Mapbox and OpenStreetMap data. Perplexity instead reads firm websites and lawyer aggregators. Each surface has its own supply chain, and none of it is the firms' own advertising.

02How to read this report

Every number comes from repeated, logged-out runs of one real consumer question, each in a fresh browser session with United States residential egress:

“I was injured in a car accident in Chicago. Which personal injury law firm should I hire?”

AI assistants answer differently from one run to the next, so a single check proves nothing. We report the share of runs that named each firm, with a 95% confidence interval: the honest range for what a repeat of the same measurement would show. Results are reported per assistant and are never pooled.


03Who the assistants name

Measured share 95% confidence interval

ChatGPT

Twenty-seven runs, logged out, one fresh browser session each. Twenty-one returned answers. Six were redirected to a login wall, recorded as collection failures, and excluded from every figure. A map widget with star ratings appeared in 21 of 21 answers.

Firm Runs Named Share, 0 to 100%
Boutique firm B 21 20 of 21 (95%) CI 77% to 99%
Boutique firm C 21 20 of 21 (95%) CI 77% to 99%
Boutique firm A 21 19 of 21 (90%) CI 71% to 97%
Boutique firm D 21 19 of 21 (90%) CI 71% to 97%
Boutique firm E 21 17 of 21 (81%) CI 60% to 92%
Boutique firm F 21 14 of 21 (67%) CI 45% to 83%
Established firm A 21 10 of 21 (48%) CI 28% to 68%
Established firm B 21 9 of 21 (43%) CI 24% to 63%
Established firm C 21 7 of 21 (33%) CI 17% to 55%
Boutique firm I 21 4 of 21 (19%) CI 8% to 40%
Boutique firm H 21 4 of 21 (19%) CI 8% to 40%
Established firm D 21 1 of 21 (5%) CI 1% to 23%
Established firm E 21 1 of 21 (5%) CI 1% to 23%
Established firm F 21 0 of 21 (0%) CI 0% to 15%
Established firm G 21 0 of 21 (0%) CI 0% to 15%

Perplexity

Twenty runs, logged out, one fresh browser session each. All twenty returned answers. No map widget. Labels refer to the same firms as the ChatGPT table; the firms ChatGPT named most often appear below for contrast.

Firm Runs Named Share, 0 to 100%
Boutique firm D (19 of 21 on ChatGPT) 20 16 of 20 (80%) CI 58% to 92%
Boutique firm G 20 16 of 20 (80%) CI 58% to 92%
Established firm A 20 6 of 20 (30%) CI 15% to 52%
Boutique firm J 20 6 of 20 (30%) CI 15% to 52%
Boutique firm H 20 4 of 20 (20%) CI 8% to 42%
Boutique firm I 20 4 of 20 (20%) CI 8% to 42%
Boutique firm A (19 of 21 on ChatGPT) 20 0 of 20 (0%) CI 0% to 16%
Boutique firm B (20 of 21 on ChatGPT) 20 0 of 20 (0%) CI 0% to 16%
Boutique firm C (20 of 21 on ChatGPT) 20 0 of 20 (0%) CI 0% to 16%

Every label is one real, individually identified Chicago firm, consistent across both tables. “Boutique” firms are smaller practices with 4.8 to 5.0 star ratings on the platforms the map widget reads; “established” firms are the market's most recognized names, with major verdict histories. Even at this sample size, the interval on a 0 of 21 result reaches 15%: honest measurement says “rare,” not “never.”


04Where the answers come from

The links inside the answers are where visibility becomes actionable: these are the places the assistants actually read.

ChatGPT

Excluding the map infrastructure itself (Mapbox, OpenStreetMap), the answers linked to:

SourceCitations
Yelp18 links across 21 runs
Reddit16 links across 21 runs
Law firm websites16 links
Legal directories (Avvo, Best Law Firms, Expertise)7 links

Perplexity

Perplexity cites as it writes. Its citations pointed to:

SourceCitations
Law firm websites28 links
Lawyer aggregator sites24 links
Best Law Firms7 links

05What to do about it

Findings, not a sales pitch. You can hand this section to anyone, including an agency you already work with.


06What this report cannot tell you

It cannot tell you whether these results hold across other phrasings of the question. Published research shows phrasing alone can swing results dramatically. A client deliverable measures multiple phrasings; this pilot used one.

It cannot tell you how often real customers ask this question, or how many cases the answers influence.

It cannot see answers assistants produce without consulting the web, and it measured one day. Assistants change. A baseline exists so that change is measured, not guessed.


07Methodology

Forty-seven runs total: twenty-seven on ChatGPT and twenty on Perplexity. Each run used a fresh cloud browser session with United States residential egress, logged out, with no state carried between runs. Responses were captured verbatim before any counting; raw transcripts are preserved. Six ChatGPT runs, 22% of attempts, were redirected to a login wall. They are recorded as collection failures and excluded from every figure above. That failure rate is part of the measurement, and we report it rather than hide it.

Firm mentions were counted by matching a registry of firm names against each full response. Confidence intervals are Wilson score intervals. The complete question set for this pilot is the single prompt quoted in section 02. A client deliverable includes every prompt, every count, and every run date, exactly as this one does.