Agent Foundations Field Network · Analytic Philosophy

Interrogate the assumptions of AI Safety

AFFINE is a network of researchers practicing analytic philosophy in a context where most current AI safety scholarship operates within a machine learning frame. We re-examine the concepts that undergird the field from agents to goals, optimization to values. Through residential seminars, research fellowships and retreats, we seek to advance new questions, new perspectives and ultimately a new paradigm which can serve as better, more idiomatic grounding to our discipline.

Why we exist

The core question of AI Safety is far larger than what is being stated

“how do we ensure that generally intelligent systems, now and in the future, are beneficial to the moral ambitions of humanity?” A simple, obviously relevant question with no present formalization that does not lose most of the important nuance.

For one thing: it is plainly not a question about computer science, though computer science will be relevant for a number of implementation details. The answers needed before one gets to coding are about the nature of intelligence, about the attributes of complex systems which constitute morality and about what it means to empower a species that disagrees with itself. These are problems regarding life, mind, value and all those things philosophy has sought to grasp since long before machine learning has emerged onto the scene. The stage has not been fully built and we do not have a language in which the script might be written.

Troublingly and simultaneously, this frame-finding work is far more urgent than philosophy is by habit. The artificial systems planning, modelling and acting within the world are already here, regardless of our comprehension, and the speed of their deployment is only increasing.

Serious effort is required to ensure their safety. Economics, control theory and machine learning are brought to bear in the construction of oversight-, evaluation- and interpretability regimes, though these concepts are borrowed and stretched beyond their rightful domains. They creak and strain and the insufficient fit of their assumptions are a thing for intelligence to exploit. Unpredictability and trouble follows the mismatch: how to continuously evaluate a continual learning agent? How to identify the right primitives to describe general agency in the first place?

The research tradition most in touch with these concerns is Agent Foundations, to which reasoning, goals and the titular agency are open problems rather than settled primitives. Though our interests range beyond it, this is where AFFINE, the Agent Foundations Field Network, gets its name.

We allege that the safety work being conducted right now might be built atop the wrong concepts and that at least a fraction of the effort exerted to push the frontier should be directed at careful, rigorous thought about the generator and those unexamined assumptions which allow one to be precise while nonetheless missing the target. This is what it means to be pre-paradigmatic in Kuhn’s sense: to lack the tools with which to discern a correct question from a complicated dead end.

It is imperative to find out early. In a sound frame, the downstream work may still take some competent researchers a number of years but in an unsound frame, the downstream work will waste the time of geniuses for eternity or, more likely, until something catastrophic happens.

Open problems

We are not sure these are the right questions yet

Before one finds answers one finds questions which are worth answering, and a field before its paradigm is still working to discover these. The following is a small but representative sample of the kinds of questions that we would like to pose, which is to say that we want answers to the kind of thing we should have been asking instead of these. To find those truer ways of interrogating reality by dissolving, merging or expanding our confusions is something we would consider to be tremendous progress.

What is an agent, and where does it end?

How to specify an agent’s objective is an ongoing area of research, but every formal answer we have seen thus far achieves its crispness by assuming a boundary somewhere. Should we discover that the boundary itself is the hard part, objectives might prove to be an entirely hopeless place to start our inquiry.

Background: what counts as one individual · Critch’s «Boundaries» sequence →

Where does wanting originate?

We don’t understand physics to desire or prefer certain outcomes over others and yet organisms do. So long as we lack clarity on how want arises from ordinary matter, we cannot hope to predict whether it might crystallize in a machine by accident.

Background: how ends enter a physical world · outer alignment →

When is a collective (or anything else) a mind?

Markets, institutions, and insect-colonies compute. Some appear aimed at outcomes no individual member desires, some sacrifice their parts where no part would consent to it. Both entities within the phrase “align AI with humanity” may be collectives, and we know neither how to model the former nor how to scrupulously speak for the latter.

Background: why scale changes the rules · Ngo’s scale-free theory of agency · ACS on hierarchical agency →

What is optimization?

It would be nice to capture evolution, markets, gradient descent, and desire under the umbrella of a single term, but it is not obvious that the underlying shape-movement is identical rather than similar, nor that the plausible subtleties of their difference may be readily and harmlessly ignored.

Background: regulation and requisite variety · The ground of optimization →

How to maintain a shape under vast optimization pressure?

Superintelligence is, among other things, an enormous amount of optimization pressure applied to the world it inhabits. All those things we seek to preserve, be they values, boundaries or structures must be able to hold their shape under those conditions. Working backwards from possible required fix-points of this type may shift the landscape of which problems appear central.

Background: reflective stability · Yudkowsky on coherence and the VNM axioms · Thornley on shutdown and decision theory →

Precedent

Hard problems are often dissolved rather than solved

The philosophy of science has already learned many of the lessons needed to make progress on our young science of agents. It would by no means be the first field to discover that it was chasing the wrong questions and that new concepts and “what if”s are necessary in order to make foundational progress.

1770s · Chemistry

By what process is phlogiston expelled from burning matter?

What combines with what when matter burns?

Lavoisier’s discovery of oxygen dissolved any question about phlogiston and paved the way for modern chemistry.

1820s · Heat

How and why does caloric fluid flow from hot to cold?

What if heat is the motion of particles?

There did not prove to be a fluid. Instead we got thermodynamics and the reframing eventually led to a physical account of information.

1900s · Light

How fast are we moving through the ether?

What if time itself bends?

Einstein had no need for the assumed medium and in turn was able to more precisely predict measurements.

1890s–1930s · Life · The closest case

What is the vital force that animates all living matter?

What structures make matter behave as though it wants something?

Driesch’s entelechy was a real answer to a real puzzle, which by no means made it correct. The appearance of purpose came down to a way of organizing matter, not to a particular substance.

We cannot, nor do we claim to know which among the concepts of our current toolbox will dissolve like the above examples, though both “all” and “none” seem like unlikely answers in the light of scientific history. What we are saying is that the question deserves being asked, that there should be some people working full time on inspecting those vital foundations atop which everything else is built and that almost nothing in the present funding landscape actually incentivizes them to do so.

Which of our terms will prove to be as obviously wrongheaded as phlogiston? What words must be invented for us to speak clearly? What concepts must be split or merged, abandoned or rescued?

Method

Thinking near the core

Normal research proceeds where the central concepts are locked in place as new ideas and novel data are funneled through them, parcelled, analyzed and incorporated. Where the foundation is sound this approach is entirely unobjectionable and will result in steady, efficient progress.

Occasionally a research trajectory stumbles upon underdefined concepts without noticing and begins to dig fiercely for an insight that does not ultimately exist. The science degrades until someone eventually realizes the mistake. If this is the case in AI Safety it would be wise to survey a few more dig-sites and distribute our eggs across a wider array of baskets.

This, of course, is difficult. How does one evaluate questions for their ability to frame a concept that does not yet exist? Rigorous conceptual thinking is not only slow and thankless, but almost impossible to measure and easy to get wrong. How can we produce it reliably?

The two components we need are simple, but difficult to fund: Space for unhurried thought and people willing to tell you that your framing is wrong instead of running with the prompt. What we want, in short, is a research monastery, so that is what we are trying to build. A setting where the conceptual scaffold can be taken apart carefully by people who have read enough to know what it is intended to do so that the parts may be weighed, rotated and when necessary replaced.

What AFFINE is

A residential research network

AFFINE grew out of a month-long residential seminar held in Czechia in May 2026, and is now undergoing the same process of refactoring described above. We’re looking through the parts we used at what we were intending to use them for and building something larger and more principled out of those insights: recurring seminars, fellowships, and retreats that bring researchers from across our unsettled field into the same room for long enough to properly disagree about the central problem.

Founded2026, by the organisers of the first AFFINE seminar.
Pilot ProgramMay 2026: 30 participants for one month in Hostačov. Read our public retrospective to see what worked and what we changed.

The Seminar

A month-long residency empowering a few dozen researchers with mathematical training and philosophical range to be guided by curiosity. Our Seminars have roughly one core mentor per five participants, reading, workshops, and a scaffold for peer teaching, aimed at open problems and foundational methods of inquiry.

The Network

Retreats, collaborations and other programs to connect researchers across sub-disciplines and career stages. In the longer term we hope to be a permanent home for foundational scholarship on intelligent agency.

Get involved

Come argue with us

There are many useful things that scholars from adjacent disciplines can bring to a young field of research, from methods to framings to object level results, though chief among them is pointing out where it is confused: The places where it isn’t answering the questions it thinks it is.

  • If you work on these problems: visit, give a talk or better yet a workshop, and spend some time poking our participants with interesting questions. The most impactful version of that visit is the one where you explain what our framing-attempts get wrong and which considerations or literature we have missed.
  • If you are an early-career researcher: apply to our next seminar. We choose participants based on philosophical aptitude and a basic foothold in mathematical formalism, not on the degree to which they share our opinions.
  • If you are curious: read along. Much of our thinking will be public, starting with the retrospective of our first seminar.