our work

We work on technical AI safety research to advance how frontier AI systems reason about and represent nonhuman beings, like animals and future digital minds. This includes benchmarks, evaluations, and other open-source tools.

Projects

MANTA logo
Benchmarks & Evals

MANTA: Do LLMs Hold Their Values on Animal Welfare?

MANTA (Multi-turn Assessment of Nonhuman Thinking & Alignment) measures whether frontier models hold their animal welfare values when users push back. Measured across 1,000+ five-turn conversations applying sustained economic, social, pragmatic, epistemic, and cultural pressure.

Released: May 2026 · Last updated: August 2026

Emergent Alignment project illustration
Research Experiments

Emergent Alignment: Does Nonhuman Welfare Generalize?

Emergent misalignment showed that training on one narrow bad behavior makes models broadly misaligned. We invert the question: does fine-tuning a model on a single good value, moral consideration for nonhuman beings, make it broadly more aligned? We evaluate across nonhuman welfare, general alignment, and capability benchmarks.

In progress