Position paper

MoSSAIC: AI Safety After Mechanism

Management of Substrate-Sensitive AI Capabilities

Matt Farr ¹ *, Aditya Arpitha Prasad ¹ *, Chris Pang ², Aditya Adiga ¹, Jayson Amati ¹, Sahil K ¹

¹ Groundless AI  ·  ² Independent

* Correspondence to Matt Farr, Aditya Arpitha Prasad

Abstract

This is a position paper. In it, we identify a causal–mechanistic paradigm in AI safety, using mechanistic interpretability as our motivating example. We cite recent results that suggest limits to the paradigm's utility in answering questions about the safety of neural networks. We argue further that those results give a taste of what is to come, by proposing a sequence of scenarios in which safety affordances based upon the causal–mechanistic paradigm break down.

Through this, we connect current empirical evidence with several persistent threat models from the agent-foundational literature (e.g., deep deceptiveness, robust agent-agnostic processes). We suggest how we might unify these threat models under a common framework, centered around our provisionally defined concept of substrate. We then present an initial, high-level sketch of a supplementary framework, MoSSAIC (Management of Substrate-Sensitive AI Capabilities), that addresses some of the core assumptions underlying the causal–mechanistic paradigm. We further present the complementary research infrastructure we are currently designing to allow us to keep pace with substrate-flexible intelligence.

1 Introduction

Neural networks (NNs) are famously described as black boxes. Their inner workings resist reduction to human-understandable concepts [1][2]. Their decision-making processes are therefore difficult to properly audit to ensure safety [2].

Neural networks are also increasingly being deployed to make decisions on behalf of humans across high-stakes domains. Our lack of understanding of how trained NNs process information to arrive at decisions poses challenges to their safe deployment [2][3][4].

The sub-field of AI safety known as interpretability seeks to produce human-understandable explanations of NN behaviors [3][5].

Bereska & Gavves (2024) [3] classify interpretability approaches into four main paradigms, which we'll take a quick look at:

Behavioral Interpretability treats models as black boxes, analyzing input-output relations without examining internal processes. They are model-agnostic and practical for complex systems but lack real insight into internal processes [3][6]. For example, minimal pair testing compares model outputs on almost-identical inputs to test for specific linguistic capabilities (e.g., "The cat sat on the mat" vs. "The cats sat on the mat" to test pluralization) [7], and perturbation analysis systematically alters inputs to see how the output changes (e.g., testing robustness to adversarial examples) [8].

Attributional Interpretability examines how individual features of the input affect the output, using gradient-based methods. These approaches offer more transparency over black-box methods, but still do not provide any information on the internal structures of models [3]. The simplest version of this is vanilla gradients, which computes the gradient of the output with respect to a change in some input feature. Subsequent versions offer more refined techniques on this basic premise [9][10].

Concept-based Interpretability seeks high-level concepts governing network behavior [3]. For instance, it might classify model outputs into honest and dishonest categories, take averages over the intermediate activations for each class and work out the difference between these averages as a vector in latent space, representing the concept “honesty” [11]. This paradigm allows for "representation engineering"—manipulating these internal representations to upregulate desirable concepts [2].

Mechanistic Interpretability starts from the bottom, identifying clusters of neurons (called "circuits") that together perform a function in the decision-making process, from there seeking to understand the relations between these circuits and how these give rise to system behavior [2][3][12]. This field treats human-understandable "features" as the fundamental unit of analysis, trying to isolate these via a number of techniques [3][5].

In this paper, we focus on mechanistic interpretability.1 We argue that mechanistic interpretability exemplifies what we term a “causal–mechanistic paradigm” in AI safety.2 In Section Section 2, we briefly overview mainstream mechanistic interpretability and specify more clearly what we mean by the causal–mechanistic paradigm. We then present extant work that suggests some underlying problems for the paradigm. In Section Section 3, we offer a provisional and pre-formal characterization of “substrate” and suggest several scenarios in which we can expect safety assurances based upon the causal–mechanistic paradigm to falter as capabilities advance, framing these in terms of substrate-flexibility. Finally, in Section Section 4, we provide a speculative frame, that of live theory, which aims to engineer general risk-mitigation methods by scaling specificity directly instead of generalizing via substrate-independent abstractions. This is our attempt to rethink the tools we use in AI safety research and orient towards more flexible, intelligent approaches to generalization, to more effectively tackle risks that can neither be tied to a specific substrate nor be defined in substrate-agnostic ways. We refer to this problem–solution pipeline as MoSSAIC, “Management of Substrate-Sensitive AI Capabilities.” We present it not as a set of fixed conclusions, but as a developing hypothesis and research bet, and we invite feedback from the research community.

Notes

  1. We acknowledge that MI encompasses diverse approaches, and our critique targets specific assumptions that become load-bearing in safety applications, not the field as pursued for pure research purposes.
  2. We choose mechanistic interpretability as a motivating example in our work for the following reasons: (1) It is the clearest current instantiation of the causal–mechanistic paradigm at work, and concentrates the specific extrapolation/fixity bets our paper investigates. (2) It is heavily resourced and highly visible; as a rough indication of current investment, we note that approximately a third of the topics listed in Open Philanthropy’s recent request for technical AI safety proposals are on mechanistic interpretability or are closely related [13].

The rest of the paper is structured as follows:

  • Section 2 briefly overviews mainstream mechanistic interpretability and specifies what we mean by the causal–mechanistic paradigm, then presents extant work suggesting underlying problems for it.
  • Section 3 offers a provisional and pre-formal characterization of substrate and suggests six scenarios in which we can expect safety assurances based upon the causal–mechanistic paradigm to falter as capabilities advance.
  • Section 4 provides a speculative frame, that of live theory, which aims to engineer general risk-mitigation methods by scaling specificity directly instead of generalizing via substrate-independent abstractions.
  • Section 5 records who did what.
  • Appendix A gives the live theory infrastructure diagram: how theory-producers, theory-consumers and AI math agents feed one another.
  • Appendix B describes the research tools we are building on live theory: live conversational threads, formalism generation, and live discernment.
  • Appendix B describes the research tools we are building on live theory: live conversational threads, formalism generation, and live discernment.
  • References lists the 64 works cited, in order of first citation.
Stay close to the work
© 2026 Groundless