KnowledgeInSight
AI Literacy
0% of Course 3 complete

Course 3 · Module 1

Safety and Alignment

This module covers the argument that advanced AI could be hard to control, and the evidence offered for and against it. You'll be able to explain the alignment problem, judge what laboratory safety tests do and don't show, and compare forecasts that run from catastrophe to abundance to ordinary technological change. The module gives you the terms of a debate in which well-informed people disagree sharply.

Module objectives

  • Explain the alignment problem and why it is hard to give an AI system the goals you intend.
  • Assess what safety evaluations and interpretability research do and don't show about current models.
  • Compare the main forecasts about advanced AI and identify the assumptions on which they differ.

3 lessons · 28 items · 3h 39m

Lessons

  1. Lesson 1 Safety and Alignment: The alignment problem You'll look at why it's hard to give a machine exactly the goal you mean, a problem identified in 1960 and still unsolved. You'll be able to explain the alignment problem in plain… 61111h 3m
  2. Lesson 2 Safety and Alignment: What safety evaluations and interpretability show You'll examine the laboratory tests behind headlines about AI systems that deceive, scheme, or resist shutdown, and the criticisms of those tests. You'll be able to say what such… 61111h 3m
  3. Lesson 3 Safety and Alignment: Forecasts in conflict: catastrophe, abundance, and normal technology You'll compare the main forecasts about where advanced AI leads, each in the form its advocates would accept. You'll be able to name the assumptions on which they differ and tell… 611111h 33m