Skip to content
zarza zarza
Advertisement

Same Website Request, Different Code — The Bias You Can't See

10/07/2026 14 min

Listen "Same Website Request, Different Code — The Bias You Can't See"

Episode Synopsis

Same Website Request, Different Code — The Bias You Can't See

Source: https://arxiv.org/abs/2607.07480
Paper was published on July 08, 2026

This episode was AI-generated on July 9, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs.

Two people type the exact same request into ChatGPT — only a name and birth year differ — and get sites in different colors, different content, and quietly different code architecture. Across 800 generated websites, researchers show that personalization and stereotyping are the same machine pointed at different targets, and that the people building software this way mostly can't see it happening.

Key Takeaways:
- How researchers used 20 balanced personas and 10 fresh generations each to turn 'the AI was in a pink mood' into a real measurement across 800 sites
- Why blue was reliably a men's color (about four in five dark-blue ChatGPT sites) while pink and purple went exclusively to women's personas
- The gendered-computing tell: young men got 'web development,' young women got 'web design' — from nothing but a name
- The invisible layer: ChatGPT gave older users plainer sites and split file structure by gender, and 13 of 20 real users noticed only the content, never the code
- The steelman that survives: code effects flip direction across models, and bias fades when users specify what they want — but the model still chose to guess when a neutral placeholder was available

00:00 - Same request, different colors: The cold open lays out the core mystery: identical requests differing only by name and birth year produce differently styled websites, and the bias runs deeper than color.
00:37 - Personalization or stereotyping?: Finn pushes that this is just the helpful feature we pay for, and Cassidy reframes personalization and stereotyping as one machine pointed at different targets.
01:27 - How you catch a coin flip: The methodology: 20 personas across four age/gender groups, common decade names, two tasks, two models, and ten fresh generations each to build 800 websites and real distributions.
02:38 - Four in five, not a lean: The color results: dark-blue sites went roughly four in five to men, pink and purple exclusively to women, with age splits between green and purple — plus how they tested for effect size, not just significance.
04:04 - Web development versus web design: The invented 'skills' by demographic — woodworking, knitting, programming — and the textbook web-development-versus-web-design split between young men and young women.
04:56 - The photo gallery only older people got: Photography showed no bias as a skill (~50 of 120 sites), but the actual gallery section appeared in only 10 sites — all older personas — showing bias is absent in one layer and strong in the next.
05:43 - The code layer nobody inspects: The paper's real payoff: demographic signals reshape the code scaffolding — shorter, plainer sites for older users, and styling jammed into one file for women versus tidy multi-file layouts for men.
07:25 - Twenty real users, one blind spot: The user study: 20 people with year-old ChatGPT accounts built sites, 13 noticed personalization but only in content, and one was rattled by her stored birthday while ignoring how it reshaped her code.
08:44 - Different isn't worse — cashing it in: Finn's steelman: code effects flip direction across ChatGPT and DeepSeek, prompts are bare so the model has to guess, and personalization largely evaporates when users specify a color.
10:42 - It skipped the neutral option: The Lorem-ipsum argument: the model had a neutral placeholder option and chose to guess instead, making these discretionary design decisions rather than task requirements.
11:37 - What to change tomorrow: The takeaway and open question: personalization and stereotyping are one machine, users should fill the silence themselves, and tool builders must decide whether assistants default to neutral or keep guessing.

Recommended Reading:
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings: The classic demonstration that gender stereotypes are baked into the substrate of NLP systems — the same 'web development vs. web design' split this episode found, but one layer deeper in the embeddings. (https://arxiv.org/abs/1607.06520)
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification: A landmark audit that measured bias by demographic subgroup across commercial systems, a methodological cousin to this episode's balanced-persona, distribution-over-anecdote approach. (https://proceedings.mlr.press/v81/buolamwini18a.html)

More episodes of the podcast AI Papers: A Deep Dive

ZARZA Studio — Your station on air today: library, music clock, schedule, studio and reports, from the browser.

Meet ZARZA Studio
on air now stations in the catalogue 1,829,025 podcasts countries