“how do we ensure that generally intelligent systems, now and in the future, are beneficial to the moral ambitions of humanity?” A simple, obviously relevant question with no present formalization that does not lose most of the important nuance.
For one thing: it is plainly not a question about computer science, though computer science will be relevant for a number of implementation details. The answers needed before one gets to coding are about the nature of intelligence, about the attributes of complex systems which constitute morality and about what it means to empower a species that disagrees with itself. These are problems regarding life, mind, value and all those things philosophy has sought to grasp since long before machine learning has emerged onto the scene. The stage has not been fully built and we do not have a language in which the script might be written.
Troublingly and simultaneously, this frame-finding work is far more urgent than philosophy is by habit. The artificial systems planning, modelling and acting within the world are already here, regardless of our comprehension, and the speed of their deployment is only increasing.
Serious effort is required to ensure their safety. Economics, control theory and machine learning are brought to bear in the construction of oversight-, evaluation- and interpretability regimes, though these concepts are borrowed and stretched beyond their rightful domains. They creak and strain and the insufficient fit of their assumptions are a thing for intelligence to exploit. Unpredictability and trouble follows the mismatch: how to continuously evaluate a continual learning agent? How to identify the right primitives to describe general agency in the first place?
The research tradition most in touch with these concerns is Agent Foundations, to which reasoning, goals and the titular agency are open problems rather than settled primitives. Though our interests range beyond it, this is where AFFINE, the Agent Foundations Field Network, gets its name.
We allege that the safety work being conducted right now might be built atop the wrong concepts and that at least a fraction of the effort exerted to push the frontier should be directed at careful, rigorous thought about the generator and those unexamined assumptions which allow one to be precise while nonetheless missing the target. This is what it means to be pre-paradigmatic in Kuhn’s sense: to lack the tools with which to discern a correct question from a complicated dead end.
It is imperative to find out early. In a sound frame, the downstream work may still take some competent researchers a number of years but in an unsound frame, the downstream work will waste the time of geniuses for eternity or, more likely, until something catastrophic happens.