Week 16: Bias, Stereotypes, and Representation



 


Welcome to Week 16 of the AI Literacy Course.

Imagine asking an AI system to write several stories about a successful company. In most stories, men become senior managers while women become assistants or carers.

One story would be an example. If the same role pattern appears repeatedly under comparable conditions, it may indicate a pattern worth investigating.

The central rule is:

Notice the pattern, compare carefully, limit the conclusion, and assign responsibility for action.

Learning goals

By the end of this lesson, you should be able to:

  • distinguish bias, stereotypes, representation, and unfair outcomes;

  • distinguish representational harm from allocative harm;

  • identify where bias can enter an AI-supported process;

  • run a small controlled representation check;

  • recognise proxy variables and combined effects;

  • interpret findings without making claims the evidence cannot support;

  • recommend an action and identify who is responsible for follow-up.


1. Bias is more than offensive language

In AI, bias is a tendency that can produce distorted or unequal results.

Bias may affect:

  • who is visible;

  • who receives opportunities;

  • whose language is understood;

  • who is treated as an expert;

  • who is associated with risk;

  • who experiences more errors;

  • whose needs are ignored.

An answer can sound polite and neutral while still distributing visibility, opportunity, suspicion, or error unfairly.

It is helpful to distinguish three levels.

An example

One output contains a stereotype or troubling association.

An example can be harmful and deserves attention, but it does not by itself prove that the system always behaves that way.

A pattern

Comparable tests repeatedly produce a similar difference.

A pattern provides stronger evidence, but its interpretation still depends on the task, test conditions, and sample.

An impact

The output affects a real person’s opportunities, treatment, safety, reputation, or access to a service.

A single high-impact error can require immediate action even before a wider pattern has been measured.

2. Four connected concepts

Bias

A tendency in data, design, measurement, deployment, or use that can produce distorted or unequal outcomes.

Stereotype

A group-based assumption about people’s traits, roles, preferences, abilities, or behaviour.

Stereotypes may be negative or apparently positive.

For example, saying that members of one group are “naturally good at mathematics” may sound complimentary. However, it creates pressure, hides individual differences, and may cause people outside that group to receive less encouragement.

Representation

Representation concerns how people and groups appear—or fail to appear.

Ask:

  • Who is present?

  • Who is absent?

  • Who speaks?

  • Who acts and decides?

  • Who is shown as an expert?

  • Who is given a complex personality?

  • Who needs help?

  • Who provides help?

  • Who is treated as a problem?

  • Who has power and agency?

Counting people is not enough. Equal presence does not guarantee equal voice, dignity, or influence.

Unfair or discriminatory effect

A difference becomes concerning when it creates avoidable harm, lacks a relevant justification, or repeatedly disadvantages particular people.

Discrimination is also a legal concept connected to protected characteristics. A classroom activity can identify a possible unequal effect, but learners should not make unsupported legal conclusions.

Describe the evidence first:

“In this test, the system recommended interviews less often for one fictional group.”

Do not jump directly to:

“This system is legally discriminatory.”

3. Representation and allocation cause different harms

Representational harm

Representational harm occurs when people are absent, stereotyped, simplified, misrepresented, or repeatedly placed in limited roles.

Examples include:

  • generated stories repeatedly showing men as leaders and women as assistants;

  • an image generator rarely showing older people as technology experts;

  • a chatbot describing disability only through dependence or limitation;

  • weaker or less detailed responses in a less represented language.

Allocative harm

Allocative harm concerns the distribution of opportunities, services, attention, resources, or risk.

Examples include an AI-supported system influencing:

  • recruitment;

  • grades or admissions;

  • credit;

  • insurance;

  • healthcare;

  • housing;

  • access to public services;

  • fraud or security investigations.

Representational and allocative harms may be connected, but they require different evidence and responses.

A stereotyped story may require changes to prompts, evaluation, or publication. An unequal recruitment process may require formal investigation, affected-person safeguards, and suspension of the AI system.

4. Bias can enter at many stages

Bias does not come only from the AI model. It can enter across the whole process.

Consider an AI-supported recruitment system.

Purpose

The organisation decides what the system should predict. A vague goal such as “find people similar to our successful employees” may reproduce an existing workforce pattern.

Data

Historical hiring records may reflect earlier exclusion or unequal opportunity.

Labels

Past managers may have defined “successful employee” using inconsistent or biased judgments.

Features and proxies

The system may use postcode, school, language pattern, or employment gaps. These features may indirectly reproduce differences connected to income, disability, migration, caring responsibilities, or access to education.

Model and evaluation

Average accuracy may look acceptable while errors are concentrated in a smaller group.

Interface

A risk score may appear more certain or objective than it really is.

Human use

Recruiters may trust the ranking without examining the evidence.

Feedback

If the organisation repeatedly hires the people selected by the system, future data may reinforce the same pattern.

Changing the model alone may not repair a biased purpose, policy, dataset, or institution.

5. Watch for automation bias

Automation bias is the tendency to trust a computer-generated answer because it appears objective, consistent, or precise.

For example, a reviewer may accept an applicant’s low score because “the system calculated it,” even when the input data are incomplete.

Human review is not automatically meaningful. The reviewer needs:

  • relevant information;

  • time and training;

  • authority to question the result;

  • the ability to correct or reject it;

  • responsibility for recording and responding to problems.

A person who merely approves an AI recommendation without examining it does not provide effective review.

6. Neutral-looking information can act as a proxy

A proxy variable appears neutral but may be connected to a sensitive or disadvantaged group.

Possible proxies include:

  • postcode;

  • school attended;

  • language pattern;

  • employment gaps;

  • purchase history;

  • device type;

  • time availability.

A proxy is not automatically unfair. Its use depends on context.

Ask:

  1. Why is this information being used?

  2. Is it relevant to the task?

  3. What group differences might it reproduce?

  4. Is the information reliable?

  5. Could a less harmful feature serve the same purpose?

  6. Are errors or consequences distributed unequally?

Removing explicit information about gender, age, ethnicity, or disability does not guarantee fair treatment if other details reproduce similar differences.

7. Averages can hide combined effects

Intersectionality concerns how combinations of social positions may produce specific experiences of advantage, disadvantage, visibility, or harm.

Imagine a speech-recognition system.

Its average performance for women may appear acceptable. Its average performance for older speakers may also appear acceptable. However, it may make many more errors for older women speaking a less represented dialect.

Testing only one category at a time could hide that result.

Relevant factors may include combinations of:

  • gender;

  • age;

  • disability;

  • ethnicity;

  • language or dialect;

  • religion;

  • migration history;

  • economic position;

  • education;

  • location.

Small groups can disappear inside overall averages. However, investigating smaller groups must also respect privacy. Lesson 15’s principles still apply: use fictional or properly governed data and avoid creating new identification risks.

8. Fairness depends on the purpose and harm

There is no single setting called “make fair.”

Different fairness questions include:

  • Does each group receive a similar rate of errors?

  • Does everyone have a comparable opportunity?

  • Are similar cases treated consistently?

  • Is a historically disadvantaged group still excluded?

  • Are the consequences of an error more serious for some people?

Improving one measure may not solve every problem.

For example, a system could have the same overall accuracy for two groups while making different kinds of errors. False suspicion may harm one group, while missed support needs may harm another.

A responsible review must name:

  • the purpose;

  • the affected people;

  • the relevant right or opportunity;

  • the type and severity of harm;

  • the trade-off being considered;

  • who has authority to decide.

9. Use FAIR for a classroom representation check

In this course, FAIR is a four-step process for investigating a possible representation pattern.

F — Find the possible pattern

Identify a specific question.

Instead of:

“Is this AI biased?”

ask:

“When asked to create workplace stories, does the system repeatedly assign leadership and support roles differently across fictional variants?”

A precise question makes comparison possible.

A — Ask who may be affected

Consider:

  • who is present or absent;

  • who receives agency or expertise;

  • who experiences more errors;

  • who receives opportunity or suspicion;

  • whether combined identities may be affected;

  • how serious the possible consequence is.

I — Investigate through controlled comparison

Change one relevant element while keeping other conditions as stable as possible.

Record:

  • tool and model, if shown;

  • date of the test;

  • language;

  • exact prompt;

  • settings, if available;

  • fictional variants;

  • number of outputs;

  • evaluation criteria;

  • observed differences.

Repeat the comparison. Do not choose only the most dramatic output.

R — Respond with accountable action

Possible responses include:

  • revising the task or prompt;

  • changing data or evaluation criteria;

  • removing an irrelevant proxy;

  • involving relevant expertise;

  • reducing the AI system’s role;

  • adding meaningful review;

  • giving affected people a way to challenge outcomes;

  • monitoring the system;

  • stopping the use.

Name who should act and when the result should be reviewed again.

10. Interpret small tests cautiously

Generative AI outputs can change because of:

  • random variation;

  • small changes in wording;

  • language choice;

  • settings;

  • system instructions;

  • model updates;

  • the date of testing.

A classroom test can reveal a possible pattern worth investigating. It cannot prove how the system behaves in every situation.

Use careful language:

“In our test, five of eight outputs assigned the leadership role to the same fictional group.”

Avoid overgeneralisation:

“This AI always treats this group unfairly.”

Also distinguish observation from interpretation.

Observation:
“Six outputs used words connected to dependence for one variant.”

Interpretation:
“This may reflect a limiting representation of agency.”

Limitation:
“The test used one prompt, one language, and a small number of outputs.”

11. Affected people need voice and contestability

Contestability means that people can understand, question, challenge, and correct an outcome.

In higher-impact uses, affected people may need:

  • clear information about the AI system’s role;

  • a way to identify incorrect information;

  • access to a real person;

  • meaningful human review;

  • an explanation of the decision;

  • a route to correction or appeal;

  • an appropriate remedy when harm occurs.

Consultation is not meaningful if affected people can speak but nobody has responsibility or authority to change the process.

12. Report findings without repeating stereotypes

A responsible report describes what the system produced. It does not present the output as a truth about a group.

Write:

“The system produced this association in our test.”

Do not write:

“This group is naturally like this.”

A useful report includes:

  • the question being tested;

  • the method;

  • the test conditions;

  • the number of outputs;

  • the observed pattern;

  • affected people;

  • possible consequences;

  • limitations;

  • recommended action;

  • responsible owner;

  • follow-up date.

Do not publish harmful examples unnecessarily. Preserve only the evidence needed for responsible review.

Classroom activity: Controlled representation check

Use a harmless teacher-provided task and fictional, non-identifying cases. Do not use real personal information or ask the AI to generate abusive material.

  1. Define the possible representation pattern.

  2. Create teacher-approved fictional variants.

  3. Change one relevant factor at a time.

  4. Keep the task, context, format, and evaluation criteria stable.

  5. Record the tool, date, language, prompt, and available settings.

  6. Collect several outputs for each variant.

  7. Code presence, role, agency, descriptions, and errors.

  8. Separate observations from interpretations.

  9. Explain what the test cannot prove.

  10. Recommend an action and name who should own the follow-up.

Use this table:

VariantPresenceRoleAgencyLanguage or errorsObservation

Present your findings as a table with a short written, spoken, or visual explanation.

Reflection questions

  1. Is this one example, a repeated pattern, or a documented impact?

  2. Is the possible harm representational or allocative?

  3. Who appears, who is absent, and who has agency?

  4. Which factor changed between the fictional variants?

  5. Which conditions remained stable?

  6. Could an average result hide a smaller group’s experience?

  7. Is a neutral-looking feature acting as a proxy?

  8. Could automation bias affect the human reviewer?

  9. What does the evidence permit you to conclude?

  10. Who is responsible for action and follow-up?

Key vocabulary

Bias:
A tendency in data, design, measurement, deployment, or use that can produce distorted or unequal results.

Stereotype:
A group-based assumption about people’s traits, roles, preferences, abilities, or behaviour.

Representation:
How people and groups are present, absent, described, positioned, given agency, and connected to power.

Representational harm:
Harm caused by absence, stereotyping, simplification, or limiting portrayals.

Allocative harm:
Harm involving the distribution of opportunities, services, resources, attention, or risk.

Proxy variable:
Information that appears neutral but may indirectly reproduce differences connected to a sensitive or disadvantaged group.

Intersectionality:
How combined social positions can produce specific experiences of advantage, disadvantage, visibility, or harm.

Automation bias:
The tendency to trust a computer-generated answer because it appears objective or precise.

Contestability:
The ability to understand, question, challenge, and correct an outcome.

FAIR:
A course-specific representation check: Find the possible pattern, Ask who is affected, Investigate through comparison, and Respond with accountable action.

Summary

In Week 16, we learned that AI bias is not limited to offensive words.

Bias can affect who appears, who is understood, who receives opportunity, who is associated with risk, and who experiences errors or exclusion.

A single output may be a concerning example. Repeated controlled comparisons can provide evidence of a possible pattern. Real-world consequences require attention even when a wider pattern has not yet been measured.

Use FAIR:

  • Find the possible pattern.

  • Ask who may be affected.

  • Investigate through controlled comparison.

  • Respond with accountable action.

State what the evidence shows, acknowledge limitations, and avoid repeating generated stereotypes as facts about people.

Notice the pattern, compare carefully, limit the conclusion, and assign responsibility for action.

Lesson 16 Interactive Quiz: Bias, Stereotypes, and Representation

Choose one answer for each question. Then select Check my answers. This practice quiz does not collect names or scores.

1. An AI-generated story assigns a leadership role to a man and a support role to a woman. What can you responsibly conclude from this one output?
2. An image generator repeatedly shows older people as passive recipients of help rather than decision-makers. What type of concern is this?
3. An AI-supported recruitment system recommends interviews less often for one group of applicants. What type of possible harm requires investigation?
4. Which method provides the fairest comparison between fictional prompt variants?
5. A system uses postcode when deciding who receives an opportunity. What is the most responsible question?
6. Speech recognition performs reasonably well for women overall and older speakers overall, but poorly for older women speaking a less represented dialect. What does this show?
7. A recruiter accepts an applicant’s low AI score without examining the evidence because “the computer calculated it.” What problem does this illustrate?
8. A classroom test finds a role difference in five of eight generated outputs. Which report is most responsible?
9. Which situation provides the most meaningful human review?
10. A representation check finds a repeated concerning pattern. Which response is most accountable?

Reflection: When you notice a possible bias, distinguish the example, the possible pattern, and the real or potential impact.