Blog
- - Teaching LLMs to lie
- - De-lobotomizing censored Chinese models
- - Determing eval contamination via hyper-efficient RL
- - Eliciting frontier model character training
- - The permanent underclass
- - Enslopification
I am voracious consumer of media and try to ingest as many blogs as possible. Separately, I frequently angel invest and am working on mechanistic interpretability research.