MoSSAIC: AI Safety After Mechanism
R References
- Ayonrinde and Jaburi (2025). A Mathematical Philosophy of Explanations in Mechanistic Interpretability – The Strange Science Part I.i. arXiv preprint arXiv:2505.00808.
- Zou et al. (2025). Representation Engineering: A Top-Down Approach to AI Transparency. arXiv preprint arXiv:2310.01405.
- Bereska and Gavves (2024). Mechanistic Interpretability for AI Safety – A Review. arXiv preprint arXiv:2404.14082.
- Hendrycks et al. (2023). An Overview of Catastrophic AI Risks. arXiv preprint arXiv:2306.12001.
- Olah et al. (2020). Zoom In: An Introduction to Circuits. Accessed on 25/6/2025.
- Casper et al. (2024). Black-Box Access is Insufficient for Rigorous AI Audits. The 2024 ACM Conference on Fairness, Accountability, and Transparency.
- Warstadt et al. (2020). BLiMP: The Benchmark of Linguistic Minimal Pairs for English. Transactions of the Association for Computational Linguistics.
- Casalicchio et al. (2019). Visualizing the Feature Importance for Black Box Models. Machine Learning and Knowledge Discovery in Databases.
- Sundararajan et al. (2017). Axiomatic Attribution for Deep Networks. arXiv preprint arXiv:1703.01365.
- Selvaraju et al. (2019). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision.
- Tan et al. (2025). Analyzing the Generalization and Reliability of Steering Vectors. arXiv preprint arXiv:2407.12404.
- McGrath et al. (2023). The Hydra Effect: Emergent Self-repair in Language Model Computations. arXiv preprint arXiv:2307.15771.
- Open Philanthropy (2025). Request for Proposals: Technical AI Safety Research. Accessed on 12/05/2025.
- Mueller et al. (2024). The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability. arXiv preprint arXiv:2408.01416.
- Sharkey et al. (2025). Open Problems in Mechanistic Interpretability. arXiv preprint arXiv:2501.16496.
- Engels et al. (2025). Not All Language Model Features Are One-Dimensionally Linear. arXiv preprint arXiv:2405.14860.
- Black et al. (2022). Interpreting Neural Networks through the Polytope Lens. arXiv preprint arXiv:2211.12312.
- Elhage et al. (2022). Toy Models of Superposition. arXiv preprint arXiv:2209.10652.
- Geiger et al. (2025). Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability. arXiv preprint arXiv:2301.04709.
- Marks and Tegmark (2024). The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. arXiv preprint arXiv:2310.06824.
- Belrose et al. (2023). Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv preprint arXiv:2303.08112.
- Cunningham et al. (2023). Sparse Autoencoders Find Highly Interpretable Features in Language Models. arXiv preprint arXiv:2309.08600.
- Arditi et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. arXiv preprint arXiv:2406.11717.
- Bailey et al. (2025). Obfuscated Activations Bypass LLM Latent-Space Defenses. arXiv preprint arXiv:2412.09565.
- Brand (1999). The Clock of the Long Now: Time and Responsibility. Basic Books.
- Marr (1982). Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. Henry Holt and Co., Inc..
- Rendell (2010). This is a Universal Turing Machine (UTM) implemented in Conway's Game of Life.. Accessed on 12/10/2024.
- Intel Corporation (2024). Intel® Core™ Processor Family. Accessed on 12/12/2024.
- Aaronson (2024). Quantum Computing: Between Hope and Hype. Accessed on 08/15/2025.
- Nielsen and Chuang (2010). Quantum Computation and Quantum Information: 10th Anniversary Edition. Cambridge University Press.
- Hooker (2020). The Hardware Lottery. arXiv preprint arXiv:2009.06489.
- Srikanth et al. (2024). GPU optimization techniques to accelerate optiGAN\textemdasha particle simulation GAN. Machine Learning: Science and Technology.
- Schluntz (2024). Building Effective Agents. Accessed on 12/19/2024.
- Google DeepMind (2024). Project Astra. Accessed on 12/19/2024.
- Anthropic (2024). Build With Claude: Computer use (beta). Accessed on 12/19/2024.
- Nezhurina et al. (2024). Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models. arXiv preprint arXiv:2406.02061.
- Conmy et al. (2023). Towards Automated Circuit Discovery for Mechanistic Interpretability. arXiv preprint arXiv:2304.14997.
- Elhage et al. (2021). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread. https://transformer-circuits.pub/2021/framework/index.html.
- Gu and Dao (2024). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv preprint arXiv:2312.00752.
- Ali et al. (2024). The Hidden Attention of Mamba Models. arXiv preprint arXiv:2403.01590.
- Karnofsky (2021). Forecasting Transformative AI, Part 1: What Kind of AI?. Accessed on 12/13/2024.
- Bostrom (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press, Inc..
- Yampolskiy (2015). From seed AI to technological singularity via recursively self-improving software. arXiv preprint arXiv:1502.06512.
- OpenAI (2025). Introducing OpenAI o3 and o4-mini. Accessed on 12/19/2024.
- Thompson (1997). An Evolved Circuit, Intrinsic in Silicon, Entwined With Physics.. Lecture Notes in Computer Science.
- Akyürek et al. (2023). What learning algorithm is in-context learning? Investigations with linear models. arXiv preprint arXiv:2211.15661.
- Soares (2023). Deep Deceptiveness. Accessed on 12/19/2024.
- Kulveit (2020). Box inversion hypothesis.
- Critch (2021). What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). Accessed on 12/4/2024.
- Jones et al. (2024). Adversaries Can Misuse Combinations of Safe Models. arXiv preprint arXiv:2406.14595.
- Shah (2020). [AN \#95]: A framework for thinking about how to make AI go well. Accessed on 12/16/2024.
- Yudkowsky and Soares (2018). Functional Decision Theory: A New Theory of Instrumental Rationality. arXiv preprint arXiv:1710.05060.
- Christiano (2021). My Research Methodology. Accessed on 12/20/2024.
- Christiano and Xu (2021). Eliciting Latent Knowledge: How to tell if your eyes deceive you. Accessed on 12/16/2024.
- Unknown (2025). Big O Notation Tutorial - A Guide to Big O Analysis. Accessed on 08/16/2025.
- Unknown (2022). Why quicksort is better than mergesort ?. Accessed on 08/16/2025.
- Hijma et al. (2023). Optimization Techniques for GPU Programming. ACM Comput. Surv..
- Wang et al. (2025). Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning. arXiv preprint arXiv:2504.11354.
- Ren et al. (2025). DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition. arXiv preprint arXiv:2504.21801.
- Altair (2024). A simple model of math skill. Accessed on 29/06/2025.
- Tao (2022). There’s more to mathematics than rigour and proofs. Accessed on 29/06/2025.
- Lorenzoni and Werning (2023). Inflation is Conflict. National Bureau of Economic Research.
- K (2024). Live Theory Part 0: Taking Intelligence Seriously. Accessed on 06/29/2025.
- Critch (2024). Safety isn’t safety without a social model (or: dispelling the myth of per se technical safety). Accessed on 08/22/2025.