One way of viewing AI Safety is through the question: how do we make future generally intelligent systems work for the betterment of humanity and society?
Put that way, it is plainly not only a computer science question. It asks what a general intelligence is. It asks what it would mean for a system to be moral. And it asks what betterment amounts to for a species that does not agree with itself about it. Those are questions about life, about mind, and about value, and they are much older than machine learning.
Meanwhile the work is urgent in a way philosophy usually is not. We are building systems that plan, model themselves, and act in the world, and we are deploying them faster than we can say clearly what they are.
Serious effort goes into keeping them safe: evaluations, oversight, interpretability. Almost all of it inherits a vocabulary of agents, goals, rewards, and optimization, assembled from economics, control theory, and machine learning. Those concepts were built for other purposes, and they strain here. Problems and unpredictabilities arise from this perspective: how do you continuously evaluate a continual learning agent? How do we find the right primitives for describing general agency in the first place?
The research tradition closest to this concern is agent foundations, which treats agency, goals, and reasoning as open problems rather than settled primitives. AFFINE takes its name from the tradition: the Agent Foundations Field Network.
The current safety work might be built on the wrong concepts, and careful work aimed through an unexamined concept can be precise and still miss. We believe that the field is pre-paradigmatic in Kuhn’s sense: it has not yet settled the concepts that would let it tell good questions from dead ends.
We would rather find that out early. If the frame is sound, this upstream work costs a few careful people some years. If it is not, this upstream work is the only thing that helps.