1 result
for mechanistic-interpretability
-
Mechanistic interpretability ("mech interp") is the branch of explainable AI that tries to understand a neural network the way you'd reverse-engineer conventional software: identify the concrete strucfield/mechanistic-interpretability · interpretability, mechanistic-interpretability, sae, circuits, llm, safety