Reading Group 9: Model Forensics
Detecting concerning behaviour does not establish misalignment. The paper proposes a protocol for investigating what drove the behaviour.
Bangalore AI Safety & Alignment Study. A small group for thoughtful, technically grounded discussions on AI safety and alignment, spanning papers, books, and open questions.
Each node is a concept tagged on a session. Edges connect concepts that co-occurred.
Detecting concerning behaviour does not establish misalignment. The paper proposes a protocol for investigating what drove the behaviour.
A follow-up to session 7: instead of re-reading the paper, we dug into the code and the Neuronpedia demo to see the Global Workspace idea in action on Gemma.
Anthropic locates a 'J-space' inside Claude that behaves like a cognitive-science-style global workspace: a routing hub for deliberate thought that could double as a monitoring surface.