Multi-Turn
Papers on multi-turn degradation, recovery, and interaction-centric failure analysis.
Research on safe, trustworthy, and verifiable AI systems.
Papers on multi-turn degradation, recovery, and interaction-centric failure analysis.
Research work on self-preservation propensity, emergent misalignment, and character training.
Selected reading on narrow finetuning and broad behavioral change.
Long-form references on benchmarks, measurement, and what evaluations actually test.
Contributed to an open-source benchmark suite for medical LLM capabilities.