MoSSAIC: AI Safety After Mechanism

R References

  1. Ayonrinde and Jaburi (2025). A Mathematical Philosophy of Explanations in Mechanistic Interpretability – The Strange Science Part I.i. arXiv preprint arXiv:2505.00808.
  2. Zou et al. (2025). Representation Engineering: A Top-Down Approach to AI Transparency. arXiv preprint arXiv:2310.01405.
  3. Bereska and Gavves (2024). Mechanistic Interpretability for AI Safety – A Review. arXiv preprint arXiv:2404.14082.
  4. Hendrycks et al. (2023). An Overview of Catastrophic AI Risks. arXiv preprint arXiv:2306.12001.
  5. Olah et al. (2020). Zoom In: An Introduction to Circuits. Accessed on 25/6/2025.
  6. Casper et al. (2024). Black-Box Access is Insufficient for Rigorous AI Audits. The 2024 ACM Conference on Fairness, Accountability, and Transparency.
  7. Warstadt et al. (2020). BLiMP: The Benchmark of Linguistic Minimal Pairs for English. Transactions of the Association for Computational Linguistics.
  8. Casalicchio et al. (2019). Visualizing the Feature Importance for Black Box Models. Machine Learning and Knowledge Discovery in Databases.
  9. Sundararajan et al. (2017). Axiomatic Attribution for Deep Networks. arXiv preprint arXiv:1703.01365.
  10. Selvaraju et al. (2019). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision.
  11. Tan et al. (2025). Analyzing the Generalization and Reliability of Steering Vectors. arXiv preprint arXiv:2407.12404.
  12. McGrath et al. (2023). The Hydra Effect: Emergent Self-repair in Language Model Computations. arXiv preprint arXiv:2307.15771.
  13. Open Philanthropy (2025). Request for Proposals: Technical AI Safety Research. Accessed on 12/05/2025.
  14. Mueller et al. (2024). The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability. arXiv preprint arXiv:2408.01416.
  15. Sharkey et al. (2025). Open Problems in Mechanistic Interpretability. arXiv preprint arXiv:2501.16496.
  16. Engels et al. (2025). Not All Language Model Features Are One-Dimensionally Linear. arXiv preprint arXiv:2405.14860.
  17. Black et al. (2022). Interpreting Neural Networks through the Polytope Lens. arXiv preprint arXiv:2211.12312.
  18. Elhage et al. (2022). Toy Models of Superposition. arXiv preprint arXiv:2209.10652.
  19. Geiger et al. (2025). Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability. arXiv preprint arXiv:2301.04709.
  20. Marks and Tegmark (2024). The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. arXiv preprint arXiv:2310.06824.
  21. Belrose et al. (2023). Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv preprint arXiv:2303.08112.
  22. Cunningham et al. (2023). Sparse Autoencoders Find Highly Interpretable Features in Language Models. arXiv preprint arXiv:2309.08600.
  23. Arditi et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. arXiv preprint arXiv:2406.11717.
  24. Bailey et al. (2025). Obfuscated Activations Bypass LLM Latent-Space Defenses. arXiv preprint arXiv:2412.09565.
  25. Brand (1999). The Clock of the Long Now: Time and Responsibility. Basic Books.
  26. Marr (1982). Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. Henry Holt and Co., Inc..
  27. Rendell (2010). This is a Universal Turing Machine (UTM) implemented in Conway's Game of Life.. Accessed on 12/10/2024.
  28. Intel Corporation (2024). Intel® Core™ Processor Family. Accessed on 12/12/2024.
  29. Aaronson (2024). Quantum Computing: Between Hope and Hype. Accessed on 08/15/2025.
  30. Nielsen and Chuang (2010). Quantum Computation and Quantum Information: 10th Anniversary Edition. Cambridge University Press.
  31. Hooker (2020). The Hardware Lottery. arXiv preprint arXiv:2009.06489.
  32. Srikanth et al. (2024). GPU optimization techniques to accelerate optiGAN\textemdasha particle simulation GAN. Machine Learning: Science and Technology.
  33. Schluntz (2024). Building Effective Agents. Accessed on 12/19/2024.
  34. Google DeepMind (2024). Project Astra. Accessed on 12/19/2024.
  35. Anthropic (2024). Build With Claude: Computer use (beta). Accessed on 12/19/2024.
  36. Nezhurina et al. (2024). Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models. arXiv preprint arXiv:2406.02061.
  37. Conmy et al. (2023). Towards Automated Circuit Discovery for Mechanistic Interpretability. arXiv preprint arXiv:2304.14997.
  38. Elhage et al. (2021). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread. https://transformer-circuits.pub/2021/framework/index.html.
  39. Gu and Dao (2024). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv preprint arXiv:2312.00752.
  40. Ali et al. (2024). The Hidden Attention of Mamba Models. arXiv preprint arXiv:2403.01590.
  41. Karnofsky (2021). Forecasting Transformative AI, Part 1: What Kind of AI?. Accessed on 12/13/2024.
  42. Bostrom (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press, Inc..
  43. Yampolskiy (2015). From seed AI to technological singularity via recursively self-improving software. arXiv preprint arXiv:1502.06512.
  44. OpenAI (2025). Introducing OpenAI o3 and o4-mini. Accessed on 12/19/2024.
  45. Thompson (1997). An Evolved Circuit, Intrinsic in Silicon, Entwined With Physics.. Lecture Notes in Computer Science.
  46. Akyürek et al. (2023). What learning algorithm is in-context learning? Investigations with linear models. arXiv preprint arXiv:2211.15661.
  47. Soares (2023). Deep Deceptiveness. Accessed on 12/19/2024.
  48. Kulveit (2020). Box inversion hypothesis.
  49. Critch (2021). What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs). Accessed on 12/4/2024.
  50. Jones et al. (2024). Adversaries Can Misuse Combinations of Safe Models. arXiv preprint arXiv:2406.14595.
  51. Shah (2020). [AN \#95]: A framework for thinking about how to make AI go well. Accessed on 12/16/2024.
  52. Yudkowsky and Soares (2018). Functional Decision Theory: A New Theory of Instrumental Rationality. arXiv preprint arXiv:1710.05060.
  53. Christiano (2021). My Research Methodology. Accessed on 12/20/2024.
  54. Christiano and Xu (2021). Eliciting Latent Knowledge: How to tell if your eyes deceive you. Accessed on 12/16/2024.
  55. Unknown (2025). Big O Notation Tutorial - A Guide to Big O Analysis. Accessed on 08/16/2025.
  56. Unknown (2022). Why quicksort is better than mergesort ?. Accessed on 08/16/2025.
  57. Hijma et al. (2023). Optimization Techniques for GPU Programming. ACM Comput. Surv..
  58. Wang et al. (2025). Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning. arXiv preprint arXiv:2504.11354.
  59. Ren et al. (2025). DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition. arXiv preprint arXiv:2504.21801.
  60. Altair (2024). A simple model of math skill. Accessed on 29/06/2025.
  61. Tao (2022). There’s more to mathematics than rigour and proofs. Accessed on 29/06/2025.
  62. Lorenzoni and Werning (2023). Inflation is Conflict. National Bureau of Economic Research.
  63. K (2024). Live Theory Part 0: Taking Intelligence Seriously. Accessed on 06/29/2025.
  64. Critch (2024). Safety isn’t safety without a social model (or: dispelling the myth of per se technical safety). Accessed on 08/22/2025.
Stay close to the work
© 2026 Groundless