Loading
0xA0Lesson 11 of 12

AI ethics, bias and limitations

Find where bias sneaks in, recognize what AI can’t do, and check a system’s errors group by group.

20 min 6-question quiz 2 code exercises
By the end of this lesson you can
  • Name the stages where bias can enter an AI system
  • Recognize hallucination, distribution shift and adversarial examples
  • Compare a model’s error rates across groups

AI systems now help decide who gets a loan, which résumé gets read, and what a doctor looks at first. That makes their mistakes matter, especially when the mistakes fall unevenly on some groups of people.

A model has no opinions of its own; it learns patterns from data. Bias gets in at every stage:

  • Data - if some groups are underrepresented or the data reflects past discrimination, the model learns that. A face recognizer trained mostly on light-skinned faces performs worse on dark-skinned ones.
  • Labels - people label data, and their judgments can be inconsistent or prejudiced.
  • Design - the target you choose can be a poor proxy. Predicting “health need” from past healthcare spending underestimates the needs of patients who had less access to care.
  • Deployment - a model used in a different place, population or purpose than it was built for.

Try it

Review a launch plan

A company is about to launch an AI résumé screener. Read its launch note and click every part that should worry you, then check what you found.

Résumé screener launch note (1 of 1)Flags found 0/0

Click every part that looks suspicious. There are 5.

Our new screener ranks every applicant automatically. It was trained on the résumés of people we hired over the past 10 years, and it reaches 94% accuracy on that same data. To keep things simple, applicants in the bottom 60% are rejected automatically with no human review. We removed the gender column, so the model cannot be biased. We will track the number of applications processed per day. We have not compared results across different groups of applicants.

What AI can’t do (yet)

  • Hallucination - generative models produce fluent output that can be confidently wrong.
  • Distribution shift - a model trained on one kind of data struggles when the world changes: a demand forecaster trained before a pandemic, a skin-cancer detector trained on one hospital’s cameras.
  • No real understanding - models find statistical patterns. They can fail at simple reasoning a child would get right, and they don’t know when they’re out of their depth.
  • Adversarial examples - tiny, invisible changes to an input (a few altered pixels, a sticker on a stop sign) can flip a model’s answer.
  • Cost - training and running large models takes enormous amounts of computing power, energy and money.

Try it

Name the limitation

Sort each incident by the main limitation behind it.

0 of 6 sortedScore 0/0
  • “A chatbot cites a court case that never existed”

  • “A voice assistant understands some accents far worse than others”

  • “A sales forecaster trained on 2015-2019 fails badly in 2020”

  • “A few stickers make a vision system read a stop sign as a speed limit”

  • “A loan model approves fewer applicants from one neighborhood because past loans there were denied”

  • “A crop-disease detector trained on lab photos fails on blurry phone pictures from the field”

Checking errors group by group

per_group_accuracy.py
1# (group, prediction was correct?)
2results = [("A", True)] * 90 + [("A", False)] * 5 + [("B", True)] * 3 + [("B", False)] * 2
3
4overall = sum(correct for _, correct in results) / len(results)
5print(f"overall: {overall:.0%}")
6for group in ("A", "B"):
7    group_results = [correct for name, correct in results if name == group]
8    print(f"group {group}: {sum(group_results) / len(group_results):.0%} of {len(group_results)}")
Output
overall: 93%
group A: 95% of 95
group B: 60% of 5

Key takeaways

  • Bias can enter through the data, the labels, the design of the target and the deployment.

  • AI can hallucinate, break under distribution shift, be fooled by adversarial inputs, and lacks real understanding.

  • An overall accuracy can hide poor results for a group: measure errors per group.

  • Keep humans in the loop for high-stakes decisions, and monitor systems after launch.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Audit errors by group

+25 XP

Each line is group predicted actual. Print the overall accuracy as overall: P%, then each group’s accuracy in alphabetical order as GROUP: P% (N examples) (whole percents). Finally print largest gap: G points - the difference between the best and worst group accuracies in whole percentage points.

  • Two groups
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Compare approval rates

+25 XP

Each line is group decision, where the decision is approve or deny. Print each group’s approval rate in alphabetical order as GROUP: P% (whole percent). Then compare the lowest rate with the highest: if the lowest is less than 80% of the highest (the “four-fifths rule” used in US hiring audits), print review needed, otherwise within four-fifths rule.

  • Large gap
  • Close rates
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: