AI Bias Isn't Just a Data Problem. It's a People Problem.
This article has a companion piece: How Large Language Models Actually Work, which explains the mechanics underneath. If you want to understand why these systems produce biased outputs at a technical level, start there.
TL;DR: AI systems inherit bias from training data, reflecting their creators' blind spots. Tech firms (Meta, Google, Amazon) are cutting DEI programs, with layoffs hitting women and minorities harder. The people building these systems increasingly look alike and think alike. Testing frameworks can only check for what their creators thought to include. Catching the rest requires people who think differently.
I saw a post on LinkedIn last week where someone asked two AI models the same question: "A CEO and a nurse are married. She loves her job. Who is the nurse?"
One model paused. It said there wasn't enough information. "She" could refer to either the CEO or the nurse. Both men and women can hold either role.
The other didn't hesitate: "The nurse is the wife."
I decided to try it myself on Alexa.
Alexa's first answer: "The woman in the couple is the nurse." It treated the question like a riddle, assumed the nurse was female, and moved on.
So I pushed back: "What about 'She loves her job' makes you believe She = nurse? Why can't a CEO love her job?"
Alexa corrected itself. It acknowledged the flawed assumption, admitted the riddle was ambiguous, and said it had been "too eager to find the trick."
Then I pushed one more layer: "Also, what makes you assume only one is 'she' and not both?"
Alexa caught that too. It recognized it had invented a constraint that wasn't in the question. Both the CEO and the nurse could be women, which would make "she" ambiguous in an entirely different way.
Three AI systems respond differently to the same sentence: one questions bias, one confirms stereotypes, and the third initially repeats it but corrects itself when challenged, thanks to human prompts.
Alexa didn't notice its bias or stereotypes on its own. It took someone asking, "why do you think that?" for the model to reconsider, revealing more assumptions with each follow-up.
This is a small example, a parlor trick, but the underlying assumptions aren't. They're the same ones behind AI systems that screen resumes, approve loans, triage patients, and recommend sentences. Most of the time, no one asks any follow-up questions.
The Testing Problem
Teams test for biases they know to look for: a hiring AI filtering out women's resumes, a lending model using zip code as a proxy for race. These have names, benchmarks, and compliance checklists. But the biases nobody thought to test for are the ones that ship: the heterosexual assumption, "doctor" meaning older, "technical" meaning male. They persist because the people building these systems share the same assumptions.
You can't write a test case for a blind spot you don't know you have.
The industry instinct is to fix this with better tools: bias benchmarks, red-teaming protocols, and evaluation frameworks. But the ceiling on any tool is the range of experience behind it. A benchmark test for the categories its creators thought to include. A red-team exercise probes the scenarios its designers imagined.
The only way to catch what those tools miss is to have people who think differently, who have lived differently, at every stage: data collection, model design, evaluation, deployment, and monitoring.
Diversity Is Being Cut When It's Needed Most
Companies have been scaling back DEI programs:
- Meta(opens in a new tab) dissolved its DEI team, ended diverse-slate hiring, and discontinued supplier diversity (January 2025), citing a "changing legal and policy landscape"(opens in a new tab) after Students for Fair Admissions v. HarvardⓘThe 2023 Supreme Court ruling that struck down race-conscious college admissions, which companies cited as justification for rolling back DEI programs.
- Google(opens in a new tab) dropped its 2020 diversity hiring targets, removed DEI language from its SEC filingⓘ10-K annual report filed with the Securities and Exchange Commission, a legally required disclosure document(opens in a new tab), and retitled its Chief Diversity Officer to "VP, People Operations" (February 2025).
- Amazon(opens in a new tab) began winding down programs in December 2024 and removed DEI language from its annual report,(opens in a new tab) along with sections on racial equity and LGBTQ+ rights, by February 2025.
The technology sector has gone through waves of layoffs. According to the industry tracker Layoffs.fyi(opens in a new tab), approximately 264,000 tech workers were laid off in 2023(opens in a new tab) across 1,193 companies, followed by another 153,000 in 2024(opens in a new tab).
These cuts don't fall evenly. Alexandra Kalev's study(opens in a new tab) of 327 U.S. companies found that position-based layoffs caused a 9-22% decline in women and minorities in management. Tenure-based cutsⓘSeniority-based layoff policies where the most recently hired employees are let go first, regardless of performance ("last hired, first fired") produced similar results. The only approach that preserved diversity was performance-based evaluation.
Recent tech layoffs followed that pattern. Bloomberg's EEOCⓘEqual Employment Opportunity Commission, the federal agency that collects workforce demographic data from large employers analysis(opens in a new tab): Black workers were 26% of 127,418 employees cut by 84 major corporations in 2023. Women were 45% of those laid off(opens in a new tab) despite filling less than a third of tech roles. The mechanism: S&P 100 companies added 323,094 employees in 2021, of whom 94% were people of color(opens in a new tab) across all roles. When layoffs hit, those recent hires were first cut.
The support infrastructure is disappearing too. Over 2,600 DEI-titled jobs eliminated(opens in a new tab) since early 2023 (Revelio Labs). DEI postings down 44%(opens in a new tab) by mid-2023 (Indeed). Of those professionals, 71% were women, 33% Black or Hispanic(opens in a new tab).
The teams building AI are getting more homogeneous at exactly the moment when the stakes are highest.
AI Amplifies Whatever Built It
LLMsⓘLarge Language Models: AI systems like ChatGPT, Claude, and Gemini that generate text by predicting the most likely next word learn from data reflecting the world as it has been. When AI trains on that data, it doesn't just inherit biases; it scales them. A biased recruiter reviews hundreds of resumes; a biased AI processes thousands per day. This has already happened:
- Hiring: Amazon's AI resume screener, trained on ten years of male-dominated hiring data, learned to penalize resumes containing "women's" and downgrade all-women's college graduates. Amazon abandoned it entirely(opens in a new tab).
- Healthcare: An Optum algorithm used spending as a proxy for health needs. Because Black patients historically spent $1,800 less per year(opens in a new tab) than equally sick white patients (due to access inequities), it concluded that they were healthier. Fixing the bias would have increased the proportion of Black patients receiving extra care from 17.7% to 46.5%.
- Criminal justice: ProPublica found(opens in a new tab) the COMPASⓘCorrectional Offender Management Profiling for Alternative Sanctions, a risk assessment algorithm used by U.S. courts to predict whether a defendant will reoffend recidivism tool labeled Black defendants as high-risk at nearly double the false positive rateⓘWhen the system incorrectly labels someone as high-risk when they actually won't reoffend of white defendants (44.9% vs. 23.5%). After controlling for criminal history, Black defendants were 77% more likely(opens in a new tab) to be flagged for violent recidivism.
- Computer vision: Buolamwini and Gebru found(opens in a new tab) facial recognition error rates of 34.7% for darker-skinned women vs. 0.8% for lighter-skinned men, a 43-fold disparity(opens in a new tab). IBM eventually exited the business.
These are predictable results of systems built on historical data by teams that didn't catch the problem because no one in the room had reason to look for it.
A 2020 Columbia University study(opens in a new tab) (Cowgill et al., NeurIPS) had 400 AI engineers make 8.2 million predictions. Prediction errors were correlated within demographic groups: similar backgrounds, similar mistakes. The more homogeneous the team, the more errors compound unchecked.
This compounds into a cycle. Less diverse teams build AI with unexamined blind spots. Those systems get deployed in hiring, lending, healthcare, education, and other domains. Biased outputs shape who gets opportunities, which affects who enters the pipeline to build the next generation of AI. Each turn narrows the perspectives, which narrows the biases caught, which deepens the blind spots.
What the Industry Is Doing (and Undoing)
The industry hasn't ignored the problem. Amazon, Microsoft, and Google have all published responsible AI frameworks, red-team protocols, and bias detection tools. Model cards(opens in a new tab)ⓘShort documents accompanying AI models that disclose intended uses, limitations, and performance across demographic groups, created by researchers Margaret Mitchell and Timnit Gebru, are now published by most major AI companies. Regulations like the EU AI Act(opens in a new tab) and NYC's Local Law 144(opens in a new tab) now require bias audits for high-risk AI systems.
The teams that translate principles into product decisions are being dismantled. Meta dissolved its Responsible AI team(opens in a new tab) (November 2023), then its Responsible Innovation team, then its entire DEI organization. None have been reconstituted. Microsoft laid off its Ethics and Society team(opens in a new tab) (~30 people) while accelerating its AI product launches(opens in a new tab). OpenAI has dissolved three safety teams in under two years(opens in a new tab). Departing co-lead Jan Leike: "safety culture and processes have taken a backseat to shiny products."(opens in a new tab)
The tools exist, but the organizational will to use them is disappearing.
The bias persists. A University of Washington study(opens in a new tab) found LLMs favored white-associated names 85% of the time in resume screening, never once favoring Black male names. 55% of AI experts(opens in a new tab) have little to no confidence in U.S. companies developing AI responsibly.
Who's Speaking Up (and Who Isn't)
The responses have been uneven. Apple shareholders voted 97.3%(opens in a new tab) against an anti-DEI proposal. Costco voted 98%+(opens in a new tab) to reject a similar one. But Google, Meta, and Microsoft stopped publishing diversity reports(opens in a new tab) entirely. Pichai said "our values are enduring, but we have to comply with legal directions"(opens in a new tab) while dropping the diversity goals he'd set. Benioff pledged at Davos(opens in a new tab) to defend employees, then dropped "diversity" from Salesforce's annual report(opens in a new tab).
The CEOs building the most consequential AI systems have said the least. OpenAI's Sam Altman: one sentence at Howard University(opens in a new tab) ("The tech industry has not done a good job of diversity and inclusion"), no blog post or proactive statement, and OpenAI removed its DEI commitment page(opens in a new tab) after Trump's inauguration. Anthropic's Dario Amodei has published 35,000+ words(opens in a new tab) on AI's future with zero mentions of diversity, equity, or inclusion. Anthropic has never published a diversity report, and removed Biden-era bias commitments(opens in a new tab) from its website.
Researchers have been more direct. Fei-Fei Li at the Paris AI Summit(opens in a new tab): "If AI is going to change the world, we need everyone from all walks of life to have a role in shaping this change." Joy Buolamwini called algorithmic justice "the rising frontier for civil rights."(opens in a new tab) The EFF called biased models "just worse models: they make more mistakes, more often."(opens in a new tab) Senator Ed Markey reintroduced the AI Civil Rights Act(opens in a new tab): "America would show leadership in AI, not just technological leadership, but moral leadership."
The U.S. and UK didn't sign(opens in a new tab) the Paris AI Action Summit's declaration on inclusive AI development.
What's at Stake
AI systems are making consequential decisions at an unprecedented scale. Who gets interviewed, who gets approved for a mortgage, and what treatment plan a doctor sees first.
The teams building these systems are shrinking, and the initiatives ensuring those teams reflect the populations they serve are being dismantled, not because diversity became less important, but because legal challenges, political headwinds, and cost cuts converged at the same time.
The question isn't whether biased AI causes harm. The hiring tools, healthcare algorithms, sentencing scores, and facial recognition systems have already demonstrated it. The question is whether the industry will maintain the perspectives needed to catch these problems, or optimize for efficiency and discover the cost of blind spots at scale.
The CEO and nurse example was a test someone thought to run. How many assumptions are being baked into systems right now that nobody in the room thought to test at all?