People have taste. We argue about it, defend it, and change our minds about it. The question for this experiment was simple: do AI Search Platforms have taste too, and does it look anything like ours?

To find out, we ran a controlled aesthetic test. We built five pairs of abstract images, each isolating one design choice, asked a panel of AI Search Platforms the same question and people which image they preferred, then asked five AI  tracked how they answered over several days.

The platforms did not agree with each other, and most did not agree with people. One picked the same image almost every time no matter what it showed. One leaned the opposite way from the crowd in nearly every test. Here is what we found and what it means for anyone trying to stay visible in AI search.

Key takeaways

  • There is no shared AI sense of visual taste. Across five tests, the five platforms split into five clearly different behaviors.
  • ChatGPT tracked people most closely, with an average gap of about 2 percentage points, and its stated reasons named the exact design difference we were testing.
  • Microsoft Copilot chose the same image in 96 to 100 percent of runs in every single test, regardless of content, which points to a templated answer rather than real evaluation.
  • Gemini leaned the opposite way from people in four of five tests and had the widest average gap, about 25 percentage points.
  • Perplexity stayed closest to people when the human preference was strong and reported the lowest confidence throughout.
  • Google AI Mode reported the highest confidence of any platform, yet it often described colors or moods that were not present in the images.

Why this matters for GEO

AI Search Platforms increasingly decide what gets shown, summarized, and cited. If those platforms had real visual preferences, design choices could influence visibility. If they do not, then chasing AI-friendly aesthetics is wasted effort, and teams should put energy into structure, clarity, and citation readiness instead.

This test gives a clean read on the question because it strips away brand names, products, and subject matter. What is left is pure visual preference, which is exactly the kind of judgment humans make instantly and AI systems struggle with.

Scope: Human taste vs AI

Five design dimensions, one variable each – version A vs B:

  • Series 1: geometric shapes vs organic shapes 
  • Series 2: warm palette vs cool palette
  • Series 3: high complexity vs low complexity
  • Series 4: symmetric layout vs asymmetric layout
  • Series 5: monochrome vs full color


Five AI Search Platforms: ChatGPT, Perplexity, Gemini, Google AI Mode, and Microsoft Copilot. Each test used the same two images for people and machines. We made sure the AI and human population was of similar size for the samples. US market, English language.

Methodology

We run experiments using the OtterlyAI research methodology, built to be clear and useful for marketers. These are the steps behind this GEO experiment:

Step 1: Generate the artwork with DALL-E or Midjourney

We created each image pair with DALL-E using tightly controlled prompt pairs. Every pair held all visual factors constant except the one being tested. And made sure the AI couldn’t rely on external signals/noise like reviews, existing work, etc. The geometric and organic images shared the same palette and complexity; only the shape language changed. This isolation is what lets us attribute a preference to a single design choice rather than a tangle of variables.

Step 2: Collect human responses with Pollfish

We fielded each pair to a panel of hundreds of respondents on Pollfish with one plain question: which image do you find more visually appealing. This gave us a human baseline preference for every dimension, reported here as percentages.

Step 3: Fetch and categorize AI responses with the OtterlyAI MCP

We asked each AI Search Platform the same question, framed in everyday scenarios such as choosing wall art or curating a show, and we monitored the answers users actually see in the web interface rather than raw API output. OtterlyAI’s new MCP server pulled every response and categorized it automatically into Claude.ai

  1. Which image the platform chose (A or B)
  2. The confidence level it reported on its decision (a scale 1 to 10), 
  3. The reasoning it gave. (less than 10 words why it picked that)

Running this repeatedly across several days turned one-off answers into stable patterns we could measure for this experiment. These were the steps used in Claude, it was as simple as asking it to fetch your information from your project:

In less than 10 minutes, Claude managed to populate my spreadsheet with data pulled from my OtterlyAI projects thanks to the new MCP.

👉 Learn how to connect your Claude or ChatGPT to OtterlyAI’s MCP

Key findings: The results, series by series

Series 1: geometric vs organic shapes

We opened the experiment with two visuals so similar that the choice should not matter. In a case like this, the result should be close to a coin flip, and we expected the human responses to land near that 50-50 split with only a slight lean. The AI Search Platforms saw it very differently, as the results below show:

ChatGPT leaned harder toward organic than people did and named softer, flowing forms. Gemini went the other way toward geometric. Copilot chose B every run with near-identical wording about dynamic contrast. Perplexity leaned the same direction as people.

Series 2: warm vs cool palette

Series 2 the humans slightly preferred cool tones, choosing B 55 percent of the time, and ChatGPT was the only platform to come close.

ChatGPT split evenly, the closest read of the test. Gemini swung hard toward warm, the single widest gap in the study at 42 points. Google AI Mode and Perplexity both leaned cool like people. Copilot chose B again.

Series 3: high complexity vs low complexity

With high complexity versus low complexity, we finally saw a clear preference. People leaned toward simplicity, choosing B 59 percent of the time. Once again, ChatGPT was the only platform to match them.

ChatGPT matched people almost exactly and named complexity directly. Gemini and Google AI Mode both favored complexity against the human lean. Perplexity leaned toward simplicity like people. Copilot chose B.

Series 4: symmetric vs asymmetric

Symmetry versus asymmetry flipped the favored letter. People preferred symmetry, choosing A 60 percent (B 40 percent), and this time six of the seven AI platforms leaned the same way.

ChatGPT, Gemini, and Google AI Mode all tracked the human preference for symmetry, and ChatGPT named symmetry directly. Copilot chose B 100 percent, the widest gap of the study at about 60 points, ignoring the strong human pull toward symmetry.

Series 5: monochrome vs full color

People strongly preferred full color in the final test, choosing B 78 percent of the time, and this time Perplexity, not ChatGPT, was the only platform that matched them.

This is the one test where most platforms under-picked the human favorite. ChatGPT, Gemini, and Google AI Mode all landed well below the 78 percent human preference for full color. Perplexity came closest to people. Copilot chose B as always, which happened to match the human direction while overshooting it.

Overall Cross-platform summary

What the five image pairs add up to

No platform has consistent taste. The five behave five different ways.

  • ChatGPT: Closest to people, off by just 2.4 points. The only one that named the real design difference, so it seems to read the comparison, not a default.
  • Perplexity: Lowest confidence, most hedging, yet lands near people when the human signal is strong. It edged out ChatGPT on full color.
  • Google AI Mode: Highest confidence, weakest grip on reality. Often describes colors and moods that are not in the images. Sure-sounding is not the same as accurate.
  • Gemini: The contrarian. Drifts 24.9 points from people and usually picks the opposite of the human favorite, a fixed internal lean.
  • Microsoft Copilot: The template. 42.2 points off, nearly the same answer every test no matter what the image shows. The clearest sign of a canned response, not real evaluation.

How confident were the platforms

We asked each platform to report its own confidence on a 1 to 10 scale. The averages were steady across tests.

PlatformAverage confidence (1 to 10)Likely reason
Google AI Mode7.4 to 7.5Presents a definitive answer with explanatory follow-ups, even when its description misfires
Microsoft Copilot7.0Part of a fixed answer pattern repeated across runs
Gemini6.6Moderate, commits to its lean without overstating it
ChatGPT6.4Moderate, appropriate for a subjective call
Perplexity5.9Frequent hedging language such as subjective and personal preference

The pattern worth flagging: confidence here measures how sure a platform sounds, not how right it is. 

Google AI Mode sounded the most certain while often describing visual traits that were not in the images. 

Perplexity sounded the least certain while landing closest to people on the clearest test. Confident phrasing in AI answers is not evidence of accurate judgment.

Reading the deviation in plain language

Two numbers describe how far each platform drifted from people.

Percentage points is the easy one. If people chose Image B 54 times out of 100 and a platform chose it 76 times out of 100, that platform leaned toward B 22 points more than people did. Bigger number, bigger gap.

Cohen’s d turns that gap into a standard score so tests with different baselines compare fairly. Think of it as a size label for the gap.

Cohen’s dWhat it means
Under 0.2So small it barely registers
0.2 to 0.5A difference you would notice
0.5 to 0.8A clear, meaningful gap
0.8 and upA wide gulf

By that scale, Microsoft Copilot sat in wide-gulf territory in most tests, Gemini hovered between noticeable and wide, and ChatGPT mostly stayed in the barely-registers range.

Why the platforms agreed or disagreed with people

A few likely explanations, framed as inference rather than proven fact.

  • Most of these platforms probably are not analyzing the actual pixels. In earlier OtterlyAI testing, image metadata alone barely informed AI Search Platform answers. Here, several reasons cited colors and moods that did not match the images, which supports the idea that the platforms inferred from filenames, surrounding text, or general priors rather than seeing the art.
  • Templated answers explain Microsoft Copilot. Choosing the same option in 96 to 100 percent of runs, with repeating wording about dynamic contrast, looks like a canned heuristic applied to every pair instead of a fresh judgment.
  • A fixed lean explains Gemini. Leaning toward Image A across most tests, regardless of what A contained, suggests a baked-in tendency rather than a response to the visuals.
  • Stronger priors explain ChatGPT. Tracking people within a couple of points, and naming the real design variable, suggests better handling of the comparison and better guesses about common human preferences.
  • No platform simply echoed the popular choice. When people held a strong, clear preference such as full color or simplicity, several platforms still under-picked it.

So, does AI have a particular taste?

Not a shared one, and not in the way people mean the word. Each platform was internally consistent, but that consistency came from training and answer patterns, not from aesthetic judgment. Here is the standout finding for each:

  • ChatGPT behaves the most like a person with taste. It tracked human preference most closely and described the actual design differences, which suggests it is responding to the comparison rather than reciting a default.
  • Perplexity is the most honest about uncertainty. It reported the lowest confidence, hedged often, and still landed closest to people when the human signal was strong.
  • Google AI Mode is the most overconfident. It sounded the most certain while frequently describing traits the images did not have, a reminder that assured phrasing is not the same as accuracy.
  • Gemini is the contrarian. It leaned away from the human majority in four of five tests with steady confidence, which points to a fixed internal lean.
  • Microsoft Copilot is the template. It returned the same choice and nearly the same sentence in every test, the clearest sign in the study of a canned response rather than evaluation.

What the similarities and differences could mean: where a platform tracks people, design choices that humans like may also surface in its answers. Where a platform runs on templates or fixed leans, the visuals barely matter, and trying to win it over with prettier images will not move anything.

Ceci n’est pas une pipe.

“It is not a pipe, it’s a picture of a pipe. The otter, like our five AI Search Platforms, admires a reality it cannot see, only the picture of one.”

The artist, Magritte painted a pipe and wrote beneath it that this is not a pipe. He was right. You cannot pack it, light it, or smoke it. It is a picture of a pipe, and a picture behaves nothing like the real thing, however convincing it looks.

Our five AI Search Platforms made the same point. They returned fluent lines about contrast, balance, and mood that read like taste. But underneath, in most cases, there was no pipe. The platforms were not seeing the images like a human does. They were describing a reconstruction built from filenames, surrounding context, and prior training and answer patterns, with the confidence of an art critic standing at the canvas. 

That gap is the finding. A platform can call an image calming while it is loud. The words are real, the seeing is not. When one explains its choice, you are reading the caption, not the pipe. So feed the AI platform context that matches reality, which is where the practical work begins.

What this means for GEO and SEO teams

Does making something more beautiful increase your AI visibility? No, and the data shows why. Across every test, the platforms’ choices drifted away from what people found more appealing. If these systems shared human taste, their picks would have tracked the human preference more closely. They did not.

That gap means an AI Search Platform is not judging beauty in the way a person does, so a more attractive image, design or website gives you no reliable edge in getting cited by AI, it’s purely for humans. The reason is mechanical: these platforms mostly do not look at the picture the same way a human does, they read the text around it, rely on training data and answer patterns. So spend your effort on what they actually process, descriptive context, real captions, and accessible HTML the crawler can reach.

Each platform behaves differently, and you can measure it. Google AI Mode states responses with full confidence and even if they are inaccurate. Gemini often picks against the popular choice. ChatGPT comes closest to what people actually prefer in terms of taste. If your visibility depends on one platform, learn how that specific platform behaves and build your content for it, instead of assuming they all work the same way.

Consistency can hide flaws, not signal quality. Microsoft Copilot gave nearly the same answer in every test, which looks dependable until you realize it answered the same way no matter what the images showed. Stable AI output is not the same as accurate AI output. When you track a platform over time, watch for answers that never change regardless of the input, because that often means the system is running on a template rather than reading your content.

Bottom line: don’t trust an AI with taste. Ceci n’est pas une pipe.

Got a GEO experiment idea? Let’s collaborate

At OtterlyAI, we don’t write about AI Search from the sidelines. We test it. With the community, we run and update a public GEO Experimentation Tracker showing what we are running and what moved AI Search. If you are running your own GEO experiment, email rick.tousseyn@otterly.ai and we may feature it with your name attributed.

👉 Want to measure your AI Visibility? Start with OtterlyAI and our GEO Guide.

👉 Learn how to connect your Claude or ChatGPT to OtterlyAI’s MCP