Extracting Concepts from GPT-4
A few weeks ago Anthropic [announced they had extracted millions of understandable features](https://simonwillison.net/2024/May/21/scaling-monosemanticity-extracting-interpretable-features-from-c/) from their Claude 3 Sonnet model. Today OpenAI are announcing a similar result against GPT-4: We …