AI Safety Research Collaborations
Research work on self-preservation propensity, emergent misalignment, and character training.
Research on safe, trustworthy, and verifiable AI systems.
Research work on self-preservation propensity, emergent misalignment, and character training.
Selected reading on narrow finetuning and broad behavioral change.
Long-form references on benchmarks, measurement, and what evaluations actually test.