Open Issues Need Help
View All on GitHub Stress-test gating and detector calibration under distribution shift about 2 hours ago
help wanted question
Transformer interpretability research framework for circuit discovery, sparse features, and controlled residual-stream interventions.
Python
#activation-steering#adversarial-robustness#gpt2#interpretability#machine-learning-research#mechanistic-interpretability#sparse-autoencoders#transformers
help wanted question
Transformer interpretability research framework for circuit discovery, sparse features, and controlled residual-stream interventions.
Python
#activation-steering#adversarial-robustness#gpt2#interpretability#machine-learning-research#mechanistic-interpretability#sparse-autoencoders#transformers
Replicate residual-steering and gating results across seeds about 2 hours ago
enhancement help wanted
Transformer interpretability research framework for circuit discovery, sparse features, and controlled residual-stream interventions.
Python
#activation-steering#adversarial-robustness#gpt2#interpretability#machine-learning-research#mechanistic-interpretability#sparse-autoencoders#transformers