Back
An Anthropic paper finds a structured "scratchpad for thoughts" that nobody trained into the model, plus a clever new way to watch it work.
ai
anthropic
interpretability